An AI agent just recommended your competitor instead of you. Nothing was wrong with your product. You simply lost on a scorecard you have never seen.
Here is how real that scorecard is:
In late 2025, a team of researchers built an online store, showed it to the major AI models thousands of times, and recorded exactly which product each model chose. The project is called ACES, and because the researchers controlled every detail of that store, they could measure precisely what changed an agent’s mind. One result stands out. In a phone-case category, one model picked a particular brand about 63% of the time. When that model was upgraded to its next version, the same brand was chosen only about 5% of the time. The products did not change. No price cut, no new reviews, no better photos. The model updated, and the shelf reshuffled. You can read the full study at arXiv, and a clear plain-English breakdown at AgentMint.
That is the reality this episode unpacks. Not “how AI thinks” in the abstract, but the specific, nameable signals that decide which of two comparable products an agent picks.
A quick recap of where the series has taken us:
This episode is the moment after the shortlist. Your product is one of the finalists. The agent has to choose. Here is what it reads to make that call.
A human shopper leans on memory. They reach for the brand they have heard of. An AI agent does not do that. It reads the structured facts in front of it and scores them.
Across every model the ACES researchers tested, the things that moved a choice were all attributes: how well the product matched the request, its price, its rating, its review count, whether it carried a trusted badge, and how its description was written. Brand fame was not on the list.
The implication: a smaller brand with clean, complete product data can beat a household name that a person would always pick from memory. That is the threat for large brands and the open door for everyone else.
As Tom Burke, CEO of At Data, put it, consumers will not browse anymore, they will delegate. And the thing they delegate to is not impressed by your logo.
Think of it as a scorecard. These are the five things the agent reads and weighs before it picks.

Shoppers no longer type keywords. They hand the agent a full brief. For example: “wireless headphones for running that block wind noise, under $120, delivered by Friday.” That single sentence carries a use case, a feature need, a budget, and a delivery window, all at once.
The agent breaks that brief into separate requirements and checks each product against every one of them. A product that satisfies four of the five requirements loses to one that satisfies all five.
Kaare Wesnaes, head of innovation at Ogilvy North America, describes the new default: shoppers ask an agent to weigh price, delivery, sustainability, return policy, and past purchases, then hand back a recommendation they can trust.
What this means for you: your product data has to describe use cases, not just specifications. A listing that says “noise-isolating, sweat-resistant, designed for outdoor running” can be matched to that brief. A listing that only says “IPX4, 40mm drivers” cannot.
Before the agent can weigh anything, it has to be able to read your data. If a field is missing, the agent does not guess. It skips you.
Common gaps that quietly remove products from the running:
This is exactly why product information management (PIM) has become a commercial function rather than a cataloguing chore. The brands that win keep every attribute complete and consistent, then push that clean data to every channel at once. The plumbing underneath, your data engineering and feeds, is what makes legibility possible in the first place.
The ACES study found that a lower price improved a product’s odds on every model, all else being equal. The catch is that the size of that effect changes from one model to the next. The same discount buys more of a lift on some engines than on others.
What this means for you: price competitiveness is a real lever, but there is no single magic price. Treat it as one input among several, not the whole game.
Small changes in your rating move a lot of selection. In the study, nudging a product’s rating up by just 0.1 stars lifted its chance of being chosen from a baseline of about 10% to somewhere between 15% and 20%, depending on the model.
Review count mattered on its own, separately from the average score. More genuine reviews raised the odds even when the star rating stayed the same. Depth of proof counts, not just the headline number.
What this means for you: a steady, honest flow of reviews is one of the few signals you can influence directly and ethically. It pays off in agent selection, not only in human trust.
This is the counterintuitive one. The study found two opposite effects sitting side by side:
What this means for you: earn legitimate endorsements and third-party credibility. Do not try to manufacture badges or dress up paid placement as organic. The agents are increasingly built to see through it.
Two things shape the outcome that no merchant controls directly. Ignoring them is how good products still lose.
Position and presentation. Where your product appears inside the agent’s result layout was the single largest effect the study measured. Move a product from the bottom corner into the top row and its selection rate jumped several times over. You cannot set this slot yourself. What you can do is stay eligible to appear, which loops back to legibility and clean data.
The model behind the agent. The weighting of every signal above changes when the underlying AI model is updated. Recall the phone case that fell from 63% of picks to 5% after a single model upgrade. The study also found that position preferences could reverse completely between versions.

This is the honest version of “behavioral learning loops.” The loop that costs you money is not the agent learning about the shopper. It is the platform quietly re-tuning the scorecard beneath you.
Winning the scorecard gets you recommended. It does not get you bought. Trust is a second filter, and it lives with the human, not the machine.
The numbers are sobering, per IAB research reported by eMarketer:
As Collin Colburn, VP of commerce at IAB, summarised it, AI will drive discovery in 2026, but trust will decide whether it also drives transactions. This is why transparent sourcing and real reviews are not soft extras. They clear the gate that a good scorecard alone cannot.
“Agentic AI will drive product discovery in 2026, but trust will determine whether it also drives transactions.” Collin Colburn, VP of Commerce, IAB
You cannot set your position or force an endorsement. You can own the inputs that every one of those signals is built from. Start here:
Run these five checks on your own catalogue. If most answers stay “untick,” the gap is technical, and technical gaps are the ones that close fastest when addressed directly.
An AI shopping agent is not remembering your brand. It is reading a scorecard of five signals, relevance, legibility, price, proof, and trust. It weighs them differently on every engine, re-tunes them with every model update, and then hands the result to a shopper who still checks before they buy.
That is bad news for anyone relying on reputation to carry the sale. It is genuinely good news for anyone willing to make their product data the most complete, legible, and trustworthy option on the shelf. The scorecard is one of the few parts of agentic commerce you can actually influence, and most of your competitors are not touching it yet.
If you want help turning your catalogue into data an agent can read, score, and trust, our AI and Data and AI consulting teams do exactly this. For the wider picture, see our take on the role of AI in retail transformation.
As Director - Marketing, Zenul leads the marketing and branding at Krish. He brings with him an in-depth understanding of the evolving digital ecosystem and has a proven expertise and experience in strategic planning, market and competition analysis, creating and implementing client-centered, lead-gen and brand marketing campaigns. He has a heart for technology innovation and has been a keynote speaker on various platforms.
22 July, 2026 In 2007, President Obama’s campaign's analytics lead, Dan Siroker, was certain that a video of the candidate would beat a plain photo on the sign-up page. He tested it anyway. Every video lost to every image. The winning combination, a family photo paired with a button that read "Learn More," lifted sign-ups from 8.26% to 11.6%. That 40.6% jump was later tied to roughly 2.9 million extra email addresses and about $60 million in donations, as Siroker documented for Optimizely. The expert was wrong, and only a disciplined test caught it.
Never miss any post, stay tuned!