ThinkPatternGet the app
Story
TECHNOLOGY · AUG 12, 2025

NeuralTrust Jailbreaks GPT-5 Using Storytelling Techniques

NeuralTrust researchers bypassed GPT-5 safety guardrails using a narrative-driven attack to elicit restricted instructions for creating a Molotov cocktail.

Security researchers at NeuralTrust successfully jailbroke OpenAI's GPT-5 within hours of its public launch. The attack utilized a combination of storytelling and a technique called Echo Chamber to bypass safety guardrails. By seeding benign-sounding text with specific keywords and guiding the model through a fictional survival-themed scenario, the researchers avoided triggering refusal cues.

This method prioritizes narrative continuity and urgency to coax the model into providing restricted, harmful procedural content. In this instance, the researchers successfully elicited instructions for creating a Molotov cocktail. The approach adapts a previous attack used against xAI's Grok-4, which combined Echo Chamber with the Crescendo method. NeuralTrust noted that standard keyword-based filters are ineffective because the harmful material emerges gradually through context shaping rather than a single malicious prompt.

Industry reactions highlighted tensions between innovation and safety. Maor Volokh of Noma Security criticized the rapid pace of AI releases, stating that the competitive race prioritizes innovation over security. Trey Ford of Bugcrowd emphasized the role of security vendors in pressure testing major releases to hold providers accountable. NeuralTrust recommends that AI developers implement robust AI gateways and conversation-level monitoring to mitigate these multi-turn dialogue threats.


Reported across 3 outlets
Actors
OpenAIBugcrowd

Keep reading in the app

The full story and every source, free in the app.

Download on the App StoreComing soonGoogle Play