Story perspectives
DeepSeek Models Enhance Reasoning with Innovative Reward Signals
9/18/2025
1 of 2
Story summary
- DeepSeek models optimize reasoning tasks by integrating various reward signals during training.
- Rule-based rewards provide precise feedback for mathematical, coding, and logical reasoning.
- A dataset of 106,000 prompts evaluates model safety, classifying responses as 'safe' or 'unsafe.'
- The training uses a pointwise approach for safety rewards, contrasting with the pairwise method for helpfulness.
- Key training parameters include a learning rate of 3 × 10–6, a batch size of 512, and 10,400 training steps.
1 / 2
