Story perspectives
OpenAI Launches New Models for Enhanced Online Safety
10/30/2025
50 6
1 of 1
Story summary
- OpenAI released the gpt-oss-safeguard-120b and gpt-oss-safeguard-20b models for online safety classification, with open weights.
- They let developers view training parameters while maintaining security and help detect fake reviews and harmful content.
- The models were developed with ROOST and outperform traditional classifiers on multi-policy benchmarks, though they may produce inaccurate reasoning.
- The models are available under Apache 2.0 on Hugging Face and a December 8 hackathon is planned in San Francisco.
