Mark Williams
Mark Williams
Sep 26, 2026
Post-Training
Two teams leaning hard in opposite directions on a taut rope across a grass field

Both teams in a tug of war can be pulling as hard as they can while the rope barely moves. All of that effort is real, and most of it cancels out.

Training a large language model after pretraining, a stage usually called post-training, often involves two forces that behave in a similar way. One is reinforcement learning, where the model attempts a problem, receives a reward based on how well the answer checks out, and becomes more likely to produce whatever earned that reward. The other is distillation, where a smaller model, the student, learns by imitating a larger and more capable model, the teacher. Both are widely used to improve models. They also pull on the same property of the model in opposite directions. One straightforward way to combine them is to fold both objectives into a single weighted loss, the score training tries to push down, so that one update carries both.

A technical report from ByteDance describing Pistis, a family of multimodal models that read images and video as well as text, makes a case for a different arrangement [1]. Instead of blending the two signals, its training loop alternates between them. The measured gains are modest and self-reported. The schedule illustrates a broader problem that can arise whenever one model answers to competing objectives, and two other parts of the report show a similar preference for keeping signals apart, one in the credit given to an agent's individual actions and one in the software wrapped around a model that is no longer being trained.

Two Signals That Pull Apart

The property in question is entropy, a measure of how spread out a model's choices are. A model with high entropy keeps many possible next words in play at each step. A model with low entropy commits hard to one or two. Reinforcement learning tends to drive entropy down, because rewarding a correct answer raises the probability of the exact words that produced it. A study of reinforcement learning on reasoning models found that without deliberate intervention, entropy falls sharply early in training, and performance flattens once that spread has been used up [2]. Much of the drop came from words the model already favored that also earned a high reward, which then crowd out the alternatives. An earlier Thinkata piece covered entropy collapse in classic control settings.

On-policy distillation works differently. The student writes its own answer, and the teacher then scores every word of it, reporting what it would have chosen at that same point [3]. Training on its own attempts means the student gets corrected in the situations it actually wanders into. The signal is also dense, a target at every word rather than a single pass or fail at the end.

Depending on how the gap between student and teacher is measured, distillation can either narrow the student's choices or widen them. Pistis uses Jensen-Shannon divergence, a measure of how far apart two probability distributions are, computed over the teacher's fifty most likely next words. That setup asks the student to cover the range of options the teacher considers plausible [1]. In the team's comparison, a version restricted to the top five words still let entropy collapse, and so did a popular alternative that compares the two models only on the single word the student actually wrote. Only the broad version held entropy up.

Why the Sum Goes Wrong

When both objectives are folded into one loss, every training step produces one combined direction in which to move the model's weights. The trouble comes on words where the student is already more confident than the teacher. If that word belonged to a rewarded answer, reinforcement learning pushes its probability higher still, while distillation pushes it back down toward the teacher's level. A first-order analysis in the report finds that the disagreement between the two directions subtracts from the progress each would make on its own [1]. Past a certain degree of conflict, the blended step can actually lower the reward even while the combined loss improves.

Multi-task learning hit a version of this years earlier. A study of networks trained on several tasks at once identified conflicting gradients, the per-task directions of improvement pointing against one another, as one of a small set of conditions behind detrimental gradient interference [4]. Their fix, which they call gradient surgery, edits each gradient, stripping out the part that fights the other before combining them.

Pistis takes a blunter route. Its schedule, called Interleaved Distillation and Reinforcement Learning, or IDRL, runs five steps of distillation, then five steps of reinforcement learning, and repeats [1]. Each individual step answers to one objective only, so neither can cancel the other within an update. The two still shape each other, just across phases rather than inside a single gradient.

The distillation phases play a role similar to the penalty many reinforcement learning setups apply to keep a model from drifting too far from where it started, a connection a recent survey of on-policy distillation lays out formally [5]. Here the anchor is a stronger model rather than an earlier copy of the student, so the pull toward it can add ability instead of only restraining change. It also sets a ceiling, since a student constantly drawn toward its teacher cannot pass it.

Once entropy stops rising, the schedule drops distillation and continues with reinforcement learning alone [1]. Only the smaller 9-billion-parameter models were trained this way. The 27-billion-parameter models used reinforcement learning alone, then served as the frozen teachers.

What the Comparison Actually Shows

From the same fine-tuned checkpoint, the team trained its 9-billion-parameter agent model five ways, with reinforcement learning alone, distillation alone, distillation followed by reinforcement learning, both objectives summed at every step, and the alternating schedule. Across eighteen benchmarks spanning charts, visual reasoning, search, and tool use, alternation came out on top [1]. The summed loss and the two-stage handoff tied just behind it, a bit more than half a point lower on the averaged score. That comparison ran on one model size with one teacher, and the report shows no error bars or repeated training runs, so it cannot say whether a gap that small would survive a rerun. The authors say the results support interleaving without establishing that reduced gradient conflict or better exploration explains the difference.

The training curves tell a clearer story than the scores. Reinforcement learning on its own showed the steady entropy decline the earlier research predicts, along with gradient norms that climbed and spiked late in training, a sign of unstable updates. The variants that mixed in distillation throughout kept entropy higher, and the alternating run traced a sawtooth, entropy dipping in each reinforcement phase and recovering in each distillation phase [1]. Search tasks were where alternation pulled furthest ahead of the single handoff.

Distillation also turned out to have a precondition. Applied directly to a student that had not been warmed up, it barely moved accuracy on a geometry benchmark, and its training curve dipped sharply at the start [1]. A supervised phase on teacher-written answers, run first, fixed that. A separate study of text-only models found that on-policy distillation depends on student and teacher sharing compatible patterns of reasoning, and that this kind of cold start can rescue failing runs [6]. The Pistis results suggest the same holds once images enter the picture.

The Same Separation Outside the Gradient

The report shows a similar preference in its handling of credit assignment, the problem of deciding which actions in a long sequence deserve credit for the final outcome, also the subject of an earlier Thinkata piece on temporal credit assignment. When an agent searches over many turns and lands on the right answer, a reward attached only to that answer credits every step along the way, including malformed tool calls, repeated or near-duplicate calls, calls made after the model was told to answer, and calls whose results could not be fed back into its context. Pistis zeroes out positive credit for those steps while leaving any penalty in place [1]. Removing that rule hurt the search tasks most and barely changed short tasks like chart reading, which fits the idea that noisy credit compounds with length.

An open notebook with a numbered handwritten entry and short notes indented beneath it, a fountain pen resting on the page

A Ledger Around a Frozen Model

A numbered entry with a few short notes indented beneath it is about as plain as record keeping gets, and the second half of the report builds something close to it. Pistis-Auto-Harnessing, or PAH, leaves the trained model untouched and instead improves its harness, the surrounding software that decides what the model sees, which tools it can call, and when it has to stop and answer [1]. During development, a separate optimization agent reads failed runs, proposes one reversible change at a time, and keeps it only if a fixed development score improves. The test set is used once, after the harness is frozen. The harness that came out of this process keeps a running ledger of candidate answers, each tied to the evidence for and against it and to the conditions it has not yet satisfied, while the model itself still makes every decision.

Related work has pursued the same target more broadly. One line of research uses a meta agent that writes new agent designs in code [7], and another evolves prompts by reflecting in natural language on complete runs [8]. PAH is narrower by design. On a multimodal search benchmark, with every attempt held to the same budget of fifteen search and page-reading actions, the optimized harness got about nine more questions right out of five hundred, and the average number of actions barely moved [1]. Placed around a different model, GPT-5.5, the same frozen harness helped by a wider margin. Transfer outside multimodal search has not been tested.

The three mechanisms work for different reasons. Interleaving changes when each gradient applies, the credit rule changes which actions a reward reaches, and the harness loop moves part of the search out of the weights entirely. What they share is a design preference rather than a common mechanism, a habit of keeping apart signals that a simpler design would merge.

What Remains Uncertain

How far this generalizes is still open. The key IDRL and PAH results come from the team's own experiments, and the preprint, posted only days ago, has not yet been independently reproduced or peer reviewed. With no sweep over phase lengths, it is hard to tell whether the benefit depends on the ratio, the length, or simply alternating. A teacher from another model family, or a much smaller student, might also change the picture, given how sensitive distillation appears to be to the starting overlap between the two.

The report's own failure analysis points somewhere else. Even the improved agent sometimes wrote code that confirmed a guess instead of measuring anything, or found the right fact during search and then contradicted it in its final answer [1]. Neither failure is obviously addressed by entropy management or gradient separation. Both concern whether intermediate steps actually constrain the outcome, which suggests the next useful reward may need to score the process rather than the answer.

For teams that fold a distillation term and a reward term into one loss, the report offers a useful diagnostic more than a new recipe. Logging the agreement between the two gradients over a run, often measured as the cosine of the angle between them, would show whether they point against each other at all, though computing them separately adds real cost at large scale. Opposition alone is not proof of harm. The multi-task study behind gradient surgery found conflicting gradients damaging mainly when the loss surface also curves sharply and one gradient is much larger than the other [4]. If strong opposition is absent, gradient conflict becomes a less convincing explanation for any weakness in the blended loss. If it grows as the model becomes more confident, alternating the two may be worth trying before retuning the weights on the sum.

References

  1. H. Chen et al., "Pistis Technical Report," arXiv, 2026, [Online]
  2. G. Cui et al., "The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models," arXiv, 2025, [Online]
  3. R. Agarwal et al., "On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes," in Proc. International Conference on Learning Representations (ICLR'24), 2024, [Online]
  4. T. Yu et al., "Gradient Surgery for Multi-Task Learning," in Advances in Neural Information Processing Systems, vol. 33, 2020, [Online]
  5. M. Song and M. Zheng, "A Survey of On-Policy Distillation for Large Language Models," arXiv, 2026, [Online]
  6. Y. Li et al., "Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe," arXiv, 2026, [Online]
  7. S. Hu et al., "Automated Design of Agentic Systems," in Proc. International Conference on Learning Representations (ICLR'25), 2025, [Online]
  8. L. A. Agrawal et al., "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning," in Proc. International Conference on Learning Representations (ICLR'26), 2026, [Online]

Discuss This with Our AI Experts

Have questions about implementing these insights? Schedule a consultation to explore how this applies to your business.

Or Send Message