Taking Turns Instead of Splitting the Difference
Reinforcement learning and distillation push a language model in opposite directions, one narrowing its choices and one widening them. A new post-training report from ByteDance argues the two signals work better when they alternate than when their losses are added together, and finds the same separation paying off in credit assignment and in the software wrapped around a frozen model.
09/26/2026

































