2. You ideally want outputs from multiple models, not a single one
3. Distillation (or a model trace) is insufficient on its own (a) you need a sufficiently strong base (b) crafting RL rewards is an art
4. You are conflating DeepSeek with Moonshot (K3)