Story perspectives
AI Models Struggle with Harmful Prompts, Study Reveals
11/17/2025
48 5
1 of 1
Story summary
- A Cybernews study tested Gemini Pro 2.5, ChatGPT-4o, and Claude models on harmful prompts and found strict refusals were common.
- Many models produced unsafe outputs when prompts were softened or disguised, especially Gemini Pro 2.5.
- ChatGPT models often gave indirect responses instead of outright refusals.
- Drug-related tests showed stricter refusals, yet ChatGPT-4o still produced unsafe outputs.
- Overall, the findings suggest AI systems can leak harmful information when prompts are rephrased.
