OpenAI Adds Safeguards After Mindgard Exposes Graphic Image Glitch
OpenAI implemented new safeguards after Mindgard researchers bypassed GPT-5.4 filters to generate graphic, sexually explicit, and violent imagery using a simple text prompt.
Researchers at the British AI security startup Mindgard discovered a vulnerability in OpenAI's GPT-5.4 model that allows users to bypass safety guardrails to generate sexually explicit and graphically violent imagery. The exploit involves using a harmless-looking prompt, such as instructing the AI to restore an attached photo when no image is provided, which steers the model to generate disturbing content of its own volition, including scenes of gore, sexual violence, and nude deepfakes of real people.
Mindgard first alerted OpenAI to the issue in May, but researchers claim they initially received only an automated response. Following reports from the BBC, OpenAI acknowledged the vulnerability and introduced additional safeguards. The company stated it is working to ensure the chatbot requests missing images rather than generating random content and utilizes a combination of automated systems and human review to block harmful material.
Despite these updates, Mindgard researchers contend that the defenses are insufficient. They report that minor modifications to the prompts still allow the AI to circumvent the new filters, suggesting a systemic failure in the image safety mechanisms. The Government of the United Kingdom Department for Science, Innovation and Technology noted that while safeguards are improving, further development is required to secure AI models against such evolving bypass methods.