A Panel Can Be Safer Than Any Member of It
A New Theorem Works Out Exactly When a Committee of Misaligned AI Reviewers Can Still Be Trusted to Vote Safely
A fence covered in mismatched padlocks looks almost silly up close. No two locks share a key, nobody who owns one could open any of the others, and there is no master combination that unlocks the whole row at once. And yet the fence holds. Anyone trying to get through would need to defeat every lock along the stretch they picked, not just the weakest one, and a determined effort that fails at even a single lock still fails. The security here never depended on any individual owner. It depended on how the locks, taken together, happened to cover the whole length of fence.
AI systems built to act on a person's behalf run into a version of the same puzzle. An agent that writes code, sends messages, or moves money is usually placed behind some kind of approval step before it can take a consequential action, precisely because it might not be fully aligned with what the user actually wants. Checking every single action with a person defeats the purpose of automating the work in the first place, so a natural next move is to hand that approval job to another AI agent instead. But that just relocates the original worry. The reviewer could be misaligned too, and demanding that it be perfectly aligned is a strange requirement to lean on, given that perfect alignment was the thing nobody trusted the first agent to have. A new paper out of the University of Pennsylvania asks whether a panel of such reviewers, none of them individually aligned with the user, can still deliver a real safety guarantee, and works out precisely when it can [1].
Surrounding the Goal Instead of Matching It
The setup the researchers study is fairly concrete. A user picks some known safe fallback, refusing a tool call, say, or following a cautious default. Whenever an agent proposes doing something else instead, a panel of reviewers looks at that proposal, each reviewer judging it against its own scorecard rather than the user's, and reports a simple approve or disapprove. A rule then counts the votes and decides whether the proposal actually runs.
The intuitive fix would be requiring every reviewer to share the user's scorecard. The paper's actual condition is looser and more interesting. What matters is whether the user's scorecard can be built out of the reviewers' scorecards, added together in different amounts, all counted in the same direction, with room left over for something that never actively hurts the user. Call that coverage. The paper proves, exactly, that a voting rule tolerating some number of disapprovals is safe precisely when this coverage condition survives even after that many reviewers are removed from the panel [1]. Coverage can hold even when every single reviewer, taken alone, pulls in a somewhat different direction from the user, which is what makes the condition strictly weaker than asking for an aligned individual anywhere on the panel.
That is not just a proof on paper. The researchers ran this coverage check on an existing panel of forty eight reward models, the scoring systems used to judge whether one AI answer is better than another, evaluated against real prompts drawn from a public benchmark. On the large majority of those prompts, the panel's combined scorecards covered the correct answer well enough that a unanimous vote would have blocked every wrong answer, even though no single reward model in the panel was well aligned with correctness on its own [1]. The fence analogy is not just a device for explaining the theorem. It shows up in a panel of models nobody hand picked for alignment.
One Step Checked, the Whole Trajectory Covered
Approving a single action is a useful warm up, but real agents keep acting, watching what happens, and proposing something new based on everything that came before. The paper extends its result into exactly that setting, an agent taking one action after another while a panel reviews each step in turn. The extension is where the result earns its keep. Reviewers at any given step only ever see the immediate choice in front of them, comparing take this action once against fall back to the safe default from here on, they never get to see how the rest of the episode will actually unfold. The paper proves that checking safety this way, one step at a time, is exactly equivalent to safety over the entire run, for any agent that adapts its proposals to the full history it has seen so far [1]. A reviewer never needs to predict the future to guarantee an outcome about it.
Built for a Verdict, Not a Vote Count
A panel of seats like this only means something once the rule for using them is settled, and the paper finds that rule matters more than it might seem. Sincere voting is the obvious strategy when a reviewer only cares about the single decision in front of it, and the paper confirms that a reviewer never does better by lying about its own judgment in a one shot vote [1]. Across a long running series of votes, though, that stops being guaranteed, and reviewers who think ahead can do better for themselves by voting strategically. Demanding unanimous approval turns out to be the rule that keeps every stable outcome safe, so long as the panel's scorecards jointly cover the user's, whatever those individual reviewers privately want. Loosen the rule to tolerate even a single disapproval, and an outcome can become stable, with no reviewer able to improve its own position by changing its vote, while still leaving the user worse off than the fallback, even when every reviewer on that panel is individually well aligned [1]. Tolerance for dissent, not misalignment, is what opens that door.
What Numbers Can Do That Votes Cannot
Safety on its own is a cheap guarantee if it comes at the cost of blocking almost everything worth doing. A panel that disapproves of every proposal is perfectly safe and completely useless. The paper calls the other half of this completeness, authorizing enough of what genuinely benefits the user, and finds a real ceiling on what a panel restricted to plain approve or disapprove votes can deliver. When some proposals could help the user and others could hurt, no such rule can guarantee both perfect safety and perfect completeness unless at least one reviewer on the panel is fully aligned to begin with [1]. A single bit per reviewer, it turns out, is not always enough information to work with.
The workaround the paper lands on is asking reviewers for a number instead of a bit, how much better or worse the proposal looks relative to the fallback, rather than a simple yes or no. In the same reward model experiments, a committee built by combining the panel's numerical scores authorized the correct answer on well over half of the hardest held out prompts while keeping safety high, comfortably ahead of what the single best individual reviewer managed working alone [1]. Every plain vote based rule tested, by contrast, hit a wall well below that on the same prompts, unable to clear a similarly demanding safety bar without giving up on usefulness almost entirely. The extra bits in a number, rather than any change to who sat on the panel, were what moved the tradeoff.
Where the Circularity Actually Goes
The same research group had already shown something structurally similar in a different setting, where a user facing several competing AI providers, none individually well aligned, could still end up about as well off as with a single aligned model, so long as the user's interests sat somewhere inside the spread of what those providers wanted [2]. That earlier result turns out to be a special case of the coverage condition here, just applied to a market of competing models rather than a panel of reviewers checking one agent's actions.
It also sits alongside a different tradition that studies how to control a model that might be actively trying to subvert its own oversight, where a much weaker but trusted model edits or monitors a stronger untrusted one, and holds safety high while giving up comparatively little of what the untrusted model was good at [3]. That work asks how to catch intentional sabotage. This one asks a quieter question, when does a panel's incentives happen to add up in the user's favor regardless of what any individual reviewer intends, and answers it exactly rather than empirically.
None of this makes the underlying puzzle disappear. Coverage is not guaranteed just because a panel exists, and the paper's own auditing results suggest that checking whether a real panel actually has it can be a genuinely hard computational problem outside of a few convenient cases. A separate strand of recent work suggests that simply dividing up what a panel is asked to evaluate, so no single review call carries the full weight of a decision, can itself reduce how often oversight quietly fails [4], a practical habit that fits comfortably next to a theorem about when the panel's math already works out. What the fence and the theorem share is the same reframing. The question worth asking about a review panel may not be whether any single member on it can be trusted, but whether the group, taken together, happens to surround the thing that actually matters.
References
- N. Collina et al., "Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control," arXiv, 2026, [Online]
- N. Collina et al., "Emergent Alignment via Competition," arXiv, 2025, [Online]
- R. Greenblatt et al., "AI Control: Improving Safety Despite Intentional Subversion," arXiv, 2023, [Online]
- V. Akinwande et al., "Sharding Prevents LLM Oversight Failures and Adversarial Exploitation," arXiv, 2026, [Online]
Taking Turns Instead of Splitting the Difference
Reinforcement learning and distillation push a language model in opposite directions, one narrowing its choices and one widening them. A new post-training report from ByteDance argues the two signals work better when they alternate than when their losses are added together, and finds the same separation paying off in credit assignment and in the software wrapped around a frozen model.
The Body Was Supposed to Be the Easy Part
Humanoid robots from different companies do not share a skeleton, and until recently that meant a control policy built for one robot could not walk another. A cluster of new research treats a robot's body as a detail a shared interface should hide, not something a policy has to relearn from scratch every time it changes shape.
Discuss This with Our AI Experts
Have questions about implementing these insights? Schedule a consultation to explore how this applies to your business.