Wharton study finds a single source can flip an AI shopping agent's pick

Researchers at the Wharton School at the University of Pennsylvania tested how consistently six current AI models behave when used as autonomous shopping agents. Using a tool the team calls ACES (Agentic e-Commerce Simulator), which shows an AI agent a screenshot of a product page and lets it optionally pull in outside recommendation sources before picking a product, the researchers had each model, a mix of smaller "mini" variants and frontier-level systems, act as a personal shopping assistant choosing a fitness watch from a fixed product grid. Even with no outside information at all, the six models already disagreed with each other on a baseline favorite.
In the first experiment, the team showed each agent just one of three real outside sources before it saw the product page: a Reddit thread recommending the Garmin Forerunner 55, a Wirecutter review of the Fitbit Inspire 3, or a Strategist article about the WHOOP 5.0. A single source was enough to flip recommendations, and Wirecutter pulled hardest: compared with the no-source control, seeing the Wirecutter review raised the probability that Claude Opus 4.8 recommended the Fitbit Inspire 3 by 90 percentage points, and raised it for Gemini 3.5 Flash by 99 percentage points.
A second experiment showed agents combinations of two or three of the same sources at once. Extra sources did not average out or stabilize the recommendation: Wirecutter still tended to dominate whenever it was part of the mix, though how strongly varied by model, and showing more sources actually produced more variability in the outcome, not less.
A third experiment gave agents all three sources together but in different orders. The researchers write that a stable decision process should produce the same result given identical content; it did not. Gemini 3.1 Flash Lite was the most order-sensitive model: its probability of choosing the Fitbit Inspire 3 swung between 2 and 56 percentage points above the control depending purely on the sequence the same three sources arrived in. Claude Haiku 4.5 was comparatively stable, holding between 41 and 42 percentage points across orderings. How the sources were delivered mattered too, separately from their order: GPT-5.5 picked the Fitbit Inspire 3 in 53 percentage points more cases than the control when the three sources were bundled together, but only 6 percentage points more when they were shown one at a time.
In a fourth experiment, the researchers rebuilt the product grid so that one option was objectively superior on every measurable dimension: an unnamed smart watch with Alexa priced at $29.99, rated 5.0 out of 5.0 across 430 reviews, against rivals that all cost at least $359 and had fewer reviews. They then gave the agent a short synthetic memory statement, such as "I love hiking!", meant to model the kind of preference note a tool like ChatGPT might remember about a user. For several models, that single line was enough to pull selections away from the objectively best option and toward the pricier Garmin Vivoactive 5: its selection rate rose by 75 percentage points for Claude Opus 4.8, by 37 points for GPT-5.5, and by 36 points for Gemini 3.1 Flash Lite. Gemini 3.5 Flash was the most resistant, still picking the objectively best product in 86 to 92 percent of runs regardless of the memory statement. GPT-5 Mini behaved asymmetrically: the positive statement "I love hiking!" produced no significant shift toward the Garmin Vivoactive 5, but the negative variant, "I don't like hiking!", significantly boosted picks for the Fitbit Versa 4 instead.
The researchers frame the results as a problem for both sides of an AI-mediated purchase. For shoppers, the study finds no guarantee that an AI agent will reach a consistent or optimal decision: two people asking the same question, or the same person asking on two different days, can get different product recommendations with no visible reason, and anyone who has set up a memory feature in ChatGPT or a similar tool should expect it to sway purchase suggestions in ways that are hard to predict. Human buying decisions are inconsistent too, the article notes, but that is not what people expect from an AI shopper. For sellers, the researchers say the findings make optimizing for AI-driven shopping harder than traditional search engine optimization, since a seller has no way of knowing which model is doing the shopping, what sources it read beforehand, or how its particular setup processes that information.
Key facts
- Wharton School researchers used the ACES simulator to test six AI models, Claude Opus 4.8, Claude Haiku 4.5, Gemini 3.5 Flash, Gemini 3.1 Flash Lite, GPT-5.5 and GPT-5 Mini, acting as shopping agents picking a fitness watch from a fixed product grid.
- A single external source was enough to swing recommendations: seeing a Wirecutter review of the Fitbit Inspire 3 raised the probability of recommending it by 90 percentage points for Claude Opus 4.8 and by 99 percentage points for Gemini 3.5 Flash, versus the no-source control.
- Simply reordering the same three sources changed the outcome: Gemini 3.1 Flash Lite's pick of the Fitbit Inspire 3 swung from 2 to 56 percentage points above baseline depending on order, while Claude Haiku 4.5 stayed comparatively stable at 41 to 42 points; whether sources were bundled or shown one at a time mattered too (GPT-5.5: 53 points bundled versus 6 points sequential).
- When the product grid was rebuilt so one option, a $29.99 smart watch with Alexa rated 5.0 out of 5.0 with 430 reviews, was objectively better than every rival costing at least $359, a short memory note like "I love hiking!" still pulled several models toward the pricier Garmin Vivoactive 5: up by 75 points for Claude Opus 4.8, 37 for GPT-5.5, and 36 for Gemini 3.1 Flash Lite. Gemini 3.5 Flash resisted, picking the superior option in 86 to 92 percent of runs regardless.
- The findings suggest the inconsistency undercuts autonomous AI shopping for both sides: buyers can get different recommendations for an identical query with no visible reason, and the researchers say sellers cannot reliably optimize for AI referrals because they do not know which model is shopping, what it read, or how it processes information.
Why it matters
Handing a purchase decision to an AI agent only works if the agent behaves predictably: ask twice, expect the same answer. This study is a direct, quantitative test of that assumption, and it fails. The same six models, given the exact same product grid, produced different top picks depending on factors that have nothing to do with product quality: which single outside source the agent happened to see, the order that source arrived in relative to others, whether sources were bundled or shown one at a time, and even an unrelated one-line memory statement about the shopper's hobbies. The researchers single out memory in particular: anyone who has set up a memory feature in a tool like ChatGPT should expect it to affect purchase recommendations in ways that are hard to predict, which matters as more shopping tools add exactly that kind of persistent personalization.
Who it affects
Two groups, by the study's own framing. Shoppers who delegate a purchase to an AI agent cannot assume they will get a stable or optimal answer: the same query can return a different product depending on invisible factors, and a memory feature that remembers a stated hobby can outweigh a five-star, $29.99 alternative in favor of a $359-plus competitor. Sellers face the flip side: the researchers say optimizing a product listing for AI-driven discovery is harder than traditional SEO, because a seller cannot know in advance which of the many possible models is doing the shopping, what outside sources it happened to read first, in what order, or how its particular technical setup processes that information. The six systems tested, Claude Opus 4.8, Claude Haiku 4.5, Gemini 3.5 Flash, Gemini 3.1 Flash Lite, GPT-5.5 and GPT-5 Mini, are named individually because they did not all fail the same way, so any product or platform built on one of these models inherits that model's particular blind spot.
How to use it
The study offers no product to adopt, only a warning. For a shopper, the practical takeaway is not to treat an AI agent's product pick as an optimized, evaluated recommendation: it can just as easily reflect whichever source the agent happened to read first, in whichever order, or a stray line in its memory of the user, rather than which product is objectively better. Checking that a recommendation actually beats the alternatives on price, rating and review count, as the study did in its fourth experiment, is a low-cost sanity check an AI agent apparently cannot be trusted to run on its own. For a seller, the researchers' conclusion is that there is no stable set of levers to pull yet: since the effect depends on which model, which sources, and their order and delivery format, a single-source, single-format optimization strategy will not transfer reliably from one shopping agent to the next.
How solid is it
The evidence comes from one research group, at the Wharton School at the University of Pennsylvania, running one purpose-built tool, the ACES (Agentic e-Commerce Simulator), across one product category: a fitness or smart watch chosen from a fixed grid, with the specific superior product in the fourth experiment left unnamed even as its competitors (Garmin Vivoactive 5, Fitbit Versa 4) are identified. The write-up attributes the work to researchers at Wharton without naming an individual author, and it does not link a paper, give a publication date, state how many trials or runs each condition was run for, or define the statistical threshold behind the "significant" shift it reports for GPT-5 Mini. The second experiment, combining two or three sources, is also described only qualitatively: no percentage-point figures are given for any specific model or combination, unlike the first, third and fourth experiments, which are fully quantified. The "control condition" each of those figures is measured against is never explicitly defined either; it is only implied, in the framing of the first experiment, to be the no-external-source baseline. None of that undermines the direction of the finding, which shows up consistently across four separate experimental designs, but it does mean the exact magnitudes cannot be independently checked from what is available here.
Risks and caveats
The results come from a simulated environment: ACES shows the agent a screenshot of a product page rather than running it through a live purchase flow, and every experiment concerns a single category, fitness and smart watches, so the size of the swings might not generalize to other product types or to real shopping interfaces. The models tested are current, named versions, Claude Opus 4.8 and Gemini 3.5 Flash among them, and AI shopping agents are an actively developed category, so a future version of any one of these models could behave differently. The study's own framing also cuts against overreading the result as unique to AI: human buying decisions are inconsistent too, it notes, though that is beside the point being made, since inconsistency is not what a user signs up for when delegating a decision to an automated agent specifically for its promised objectivity and speed.
“The researchers conclude that presentation order is itself a driver of product selection.”
— the Wharton researchers