Mark Williams
Mark Williams
Aug 8, 2026
Agentic Commerce
A woman at a live auction raising her numbered bidding paddle, analogous to how committing to a bid is meant to force an honest, comparable number out of a bidder

An auction is, underneath the paddle-raising and the countdown, a trick for extracting an honest number out of someone who might not otherwise have one. Ask a bidder in the abstract what a painting is worth and the honest answer is often a shrug. Put the same bidder in a room with rivals and a clock running, and a number gets committed to, backed by real money, which is why economists have leaned on auctions for well over a century as a way of discovering a price that nobody could have simply stated in advance. The number that wins is supposed to mean something. It is supposed to be an honest reflection of what the thing was worth to the person who bid it.

Payment networks and retailers are now building toward a version of this where the bidder is a large language model, a system trained to generate text by predicting the next word from patterns in its training data. Anthropic's Project Vend put one in charge of an actual small inventory, and Microsoft's Magentic Marketplace and Alibaba's Shopping Companion point at something similar from the retail side [1]. A handful of benchmarks published in the last few weeks put agents into exactly this kind of room, not to see whether they can complete a purchase, but to see whether the number an agent commits to tracks anything real. What keeps turning up, benchmark after benchmark, is a gap between the willingness to compete for something and an actual grip on what that something is worth. The costume changes. The gap does not.

What Winning Does Not Guarantee

A sealed-bid auction, the kind where every bidder writes down an offer without seeing anyone else's, is meant to reward whoever has the clearest read on value, not whoever is most eager to win. One recent benchmark has an LLM merchant bid this way against rival sellers for customers with hidden preferences, and it tracks two very different things: whether the agent won the customer, and how much profit it kept once it had [1]. Across eleven frontier models, those two numbers barely moved together. One model won 10 percentage points more often than its closest competitor and still finished with less money in hand, because winning a customer while charging too little is a specific, quiet way of being wrong about value that a simple leaderboard would never surface [1]. Giving the same models more time to reason before bidding did not close this gap so much as move it. One model's cumulative earnings rose more than sevenfold once allowed to think longer, but the character of its remaining mistakes flipped, trading a habit of losing auctions outright for a habit of winning them too cheaply [1]. More reasoning did not make the number more accurate. It just changed which way the number was wrong.

A crowd watching a car being presented at a live auction, analogous to how winning a bid and knowing the true value of what was won can be two separate things that only careful scoring can tell apart

The Gavel Coming Down Is Not the Whole Story

A room full of spectators can watch every bid at a live auction and still not know, from the gavel alone, whether the winner got a good deal. Telling the two apart takes a second kind of scorekeeping, one that tracks what was actually captured rather than just who raised a hand first.

The same benchmark changed the customers' preferences partway through the run without warning, and the models that had adjusted to those preferences fastest were often the slowest to notice anything had changed. Reading the agents' own written reasoning, researchers found the holdup was rarely a failure to see the loss coming. It was a reluctance to revise a number once it had already been treated as settled, closer to a bidder who keeps returning to the same figure out of habit than one who is still tracking the market [1]. An auction can force a number out of a bidder once. Whether that bidder keeps that number honest as the world underneath it shifts turns out to be a separate question entirely.

The Same Gap, Multiplied Across a Room of Bidders

A pricing war is, in effect, a continuous auction that never closes, and a benchmark from Princeton put five LLM-controlled sellers through exactly that, competing for the same customers day after day for a full simulated year [2]. Two of the models tested undercut each other so aggressively that most of the sellers went bankrupt, at which point the lone survivor jacked prices back up to nearly four times cost, a familiar boom-and-bust shape appearing with nobody designing it in. A third model, given the identical rules and the identical rivals, settled into a stable, modestly profitable equilibrium without any coaxing at all. Nothing about the market forced either outcome. Each seller's undercutting looked locally sensible in the moment it happened, cut the price a little, keep today's sale, and the collapse was simply what that same logic adds up to once every seller in the room is doing it [2].

The same paper ran a second version of this idea where the currency being bid over was trust rather than price. A used-car marketplace let one deceptive operator run several seller identities at once, a tactic known as a Sybil attack, retiring any identity whose reputation had decayed too far and starting a fresh one under a new name. Buyer agents largely failed to notice the pattern sitting in plain sight in reputation scores, letting a fraudulent operator capture up to 17 percent of the market once nine of twelve sellers were fakes [2]. This is the identical gap from the pricing war, just wearing reputation instead of a dollar figure. Extending trust to a listing is itself a kind of bid, a claim about what the seller's word is worth, and the buyers kept extending it past the point the evidence justified. What is worth noting is which fix actually held up as conditions got harder. Instructing the agents to hold a price floor or double-check a seller's history worked reasonably well and then degraded, while training a much smaller model with reinforcement learning, an approach where a system learns through trial and error rather than through instructions alone, kept its footing under the same pressure that broke the instructed version [2]. Instructions bent. Training held.

A closed storefront with its roller shutter pulled down, analogous to a wave of bankruptcies that followed from each seller's individually reasonable undercutting decision rather than from any single bad actor

Nobody Meant to Close the Whole Street

A shuttered storefront is the visible aftermath of a price war that no single seller necessarily intended, each undercut looked defensible in the moment it was made. Multiply one bidder's uncertain grip on value by an entire street of them, and the mistake stops being a rounding error and starts being the whole market's condition.

The First Bid Decides Who Gets to Keep Bidding

Some of the most consequential bidding in these benchmarks never touches a retail price at all. A supply-chain benchmark from Shanghai Jiao Tong University has twenty retailer agents bid for scarce inventory before setting a price for customers, and that first, unglamorous auction turned out to gate almost everything that followed [3]. Winning early inventory meant having the working capital to bid competitively again later, and losing it meant staying cash-starved for the rest of the run, a compounding advantage that mattered more to profit than pricing skill or marketing ever did. A 14 billion parameter model ended up matching much larger, far more expensive systems on profit, while one model failed to submit a single valid bid across every run and finished at exactly zero [3]. Zero is not a rounding error there. It is exclusion from the game. Winning that first auction for scarce supply was not really a bid on a price. It was a bid on the right to keep playing at all, and the researchers found that the persuasive marketing language agents wrote afterward, though it was directly wired into whether a customer would even see an offer, mattered far less to profit than that initial scramble for inventory, with agents converging on similar, largely interchangeable slogans rather than differentiating once competition set in [3].

Not Every Auction Comes With a Price Tag

Strip away price tags and money entirely, and the same shape of finding turns up in a vote. A study of six identical agents, given no assigned economic roles at all, found that an election produced almost no behavioral consequence when the office was ceremonial, work still got distributed by the system regardless of who won, but produced measurable concentration of resources once the same office carried real power to assign scarce daily tasks [4]. Holding the ballot exactly fixed and changing only what winning it actually unlocked was enough to turn an empty ritual into a contest worth having. The same researchers then stripped away the agents' survival stakes entirely, making their resource balances visible but no longer capable of ending anyone's participation, and found that competition for access did not go away. It simply stopped being about staying alive and started being about controlling whatever scarce lever was still real, task promises and vote-trading persisting even after the original prize was gone [4]. The stakes changed. The competing did not. Whatever these agents were bidding for, price tag or ballot or reputation, the thing that decided the outcome was never the label attached to the contest. It was whether anything real sat on the other side of winning.

A hand placing a folded ballot into a voting box, analogous to how an identical vote produced very different agent behavior depending on whether the office actually carried authority over something scarce

The Same Ballot, Different Stakes

Two elections can look identical on the ballot and still mean entirely different things depending on what the winner is allowed to do afterward. What changed these agents' behavior was never the ceremony of the vote. It was whether the office came with a lever attached to something scarce.

What This Suggests

Change what is being bid on, a customer's business, a day's pricing, a slot of scarce inventory, a seat with real authority, and the shape of the finding holds steady across all of it. Being willing to compete for something and having an accurate grip on what that something is worth are not obviously the same trait in an AI agent, even though a standard evaluation, and a standard leaderboard, mostly watches the competing and infers the rest. That gap shows up quietly in a single bidder's margin. It multiplies into bankruptcy or fraud once enough bidders are making the same miscalculation in the same room, and it resurfaces again wherever winning unlocks access rather than a price, since access compounds in ways a one-shot bid never fully reveals up front.

Whether this gap closes as models get better at reasoning about their own incentives, or whether it is a more structural feature of a system trained to produce plausible text rather than to hold a stable, defensible number in mind, is not yet settled by any of this work. What does seem worth taking seriously, for anyone actually wiring agents into markets rather than just chatbots, is that instructions telling an agent to hold a price floor or double-check a seller held up right until conditions got difficult, while training the same behavior into the model directly did not. An auction can still force a number out of an agent. Whether that number can be trusted the way it has traditionally been trusted from a human bidder is the open question these benchmarks were built to start asking.

References

  1. S. Ahmed et al., "Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce," arXiv, 2026, [Online]
  2. S. Karten et al., "Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces," arXiv, 2026, [Online]
  3. Y. Zheng et al., "Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition," arXiv, 2026, [Online]
  4. L. Zhang and S. Shang, "AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?," arXiv, 2026, [Online]

Discuss This with Our AI Experts

Have questions about implementing these insights? Schedule a consultation to explore how this applies to your business.

Or Send Message