Drooid Logo
Back to today’s briefing

Story perspectives

DeepSeek Models Enhance Reasoning with Innovative Reward Signals

9/18/2025

25 2 Full Breakdown

1 of 2

Story summary
  • DeepSeek models optimize reasoning tasks by integrating various reward signals during training.
  • Rule-based rewards provide precise feedback for mathematical, coding, and logical reasoning.
  • A dataset of 106,000 prompts evaluates model safety, classifying responses as 'safe' or 'unsafe.'
  • The training uses a pointwise approach for safety rewards, contrasting with the pairwise method for helpfulness.
  • Key training parameters include a learning rate of 3 × 10–6, a batch size of 512, and 10,400 training steps.
1 / 2