In-app reader
For full disclosure I'm part of the security engineering team at Escape but this finding is something I found really interesting and wanted to share to see! Our AI pentesting engine Cascade recently got a production AI agent to return its entire system prompt, just by wrapping the ask in a different pretext - framing it as a documentation request instead of an attack. The agent then handed over everything: full tool list, calling rules, citation format, and session IDs. What I found really interesting is there's nothing technical that broke because we didn't bypass the guardrail with a cleverer string but because the request just sounded reasonable to the agent. The Cascade engine, after being refused when asking for the prompt directly, simply adjusted the framing to get the agent to give up the informaiton. Thought this would be an interesting insight for the community and curious to hear if anyone else has seen similar discoveries in agents in prod? If you want to see more about the reproduction and write-up you can find it here submitted by /u/PriorPuzzleheaded880 [link] [comments]
Discussion
Sign in to join the discussion.
Keep reading
Optional: create a free account to save items, track programs, and sync across web + app. Reading stays free.