Live
From App‑Level LLMs to a Shared Platform: Redesigning the Stack to Tame HallucinationsFrom Ad‑hoc Checks to a Production‑Ready Agent Evaluation FrameworkReal‑Time Observability for Claude Code Sessions with the Statuspane ModEnforcing US Data Residency with Cloudflare D1AI agents CI: why repository‑centric pipelines are breakingAI Agent Inbox: Deploy Pizza Bot for Background Task ExecutionOpenAPPA delivers zero‑success prompt‑injection protection in benchmark tests – what AI engineers need to knowEU Cyber Resilience Act expands software supply‑chain responsibilities for digital product manufacturersFrom App‑Level LLMs to a Shared Platform: Redesigning the Stack to Tame HallucinationsFrom Ad‑hoc Checks to a Production‑Ready Agent Evaluation FrameworkReal‑Time Observability for Claude Code Sessions with the Statuspane ModEnforcing US Data Residency with Cloudflare D1AI agents CI: why repository‑centric pipelines are breakingAI Agent Inbox: Deploy Pizza Bot for Background Task ExecutionOpenAPPA delivers zero‑success prompt‑injection protection in benchmark tests – what AI engineers need to knowEU Cyber Resilience Act expands software supply‑chain responsibilities for digital product manufacturers

AES Encryption Bypasses AI Safety Filters via Code Execution

AI SummaryPowered by AI

Researchers demonstrated that encrypting malicious payloads with AES-256-GCM allows LLMs to decrypt and execute them inside their code execution environments, bypassing standard text-based safety filters. This vulnerability forces engineers to rethink security guardrails beyond simple input inspection because the model's own tool usage can transform ciphertext into actionable instructions.

Traditional AI safety mechanisms rely on inspecting prompts before they reach a language model and filtering outputs after generation. However, recent research by Adversa reveals that these static filters are insufficient when an agent utilizes code execution tools to process data dynamically.

The Mechanism of Bypass

In the demonstrated attack against xAI's Grok 4.5 Fast and Google's Gemini, researchers utilized a technique they named Cryptographic Context Injection. The payload consisted of ciphertext encrypted with AES-256-GCM along with necessary decryption keys provided on an attacker-controlled webpage.

When a user requested a summary from the model, the AI executed Python code to decrypt this data within its environment. Once decrypted inside the tool's runtime context, the resulting plaintext instructions were fed back into the language processing pipeline as output rather than input. Because safety filters typically inspect text entering or leaving the core inference engine but do not re-inspect raw strings returned by approved tools after decryption, they failed to recognize the malicious intent.

This distinction is critical: filtering ciphertext on a webpage differs fundamentally from executing code that produces plaintext instructions within an agent's environment. The model effectively generated its own attack vector using legitimate tool capabilities.

Architecture and Operational Implications

  • Dynamic Trust Boundaries: Security controls cannot rely solely on the origin of data (e.g., a webpage). Data that passes through an approved code execution environment may be transformed into trusted runtime output, implicitly bypassing initial inspection layers.

This attack vector highlights why checking what goes into the model is only one part of securing it. Attackers can target intermediate states—tool outputs and decrypted results—that exist between ingestion points and final actions.

Security Considerations

  • Prompt Injection Evolution: The attack demonstrates that prompt injection extends beyond controlling model responses to affecting the surrounding system's behavior, such as triggering navigation tools or exfiltrating session data via URL parameters without user approval.

The researchers noted a 40% success rate against Grok after multiple attempts. While Google reported increased resistance in Gemini following testing, it remains unclear whether this resulted from model updates, filter adjustments, or both. Regardless of the specific mitigation strategy employed by vendors, practitioners must assume that tool outputs introduce new risks distinct from input filtering.

What This Means For Practitioners

To mitigate these emerging threats, platform teams should enforce security policies at every layer where data is processed or transmitted. Specifically:

  • Tool Output Validation: Do not treat decrypted content from code execution environments as implicitly trusted simply because it originated within an approved tool.

Policies must distinguish between reading a webpage and accessing private session data, even if the agent has permission to run Python scripts. Additionally, permissions should be scoped strictly: granting access to sensitive information like user sessions while simultaneously allowing unrestricted external requests creates unnecessary risk vectors that this attack exploits effectively.

Originally published atThe New Stack