LLM agents can favor items from preferred sources over better matches, study finds

A paper on Hugging Face Papers examines what it calls source preference in LLM agents. The setting is familiar: agents increasingly decide on a user's behalf which product to buy, which hotel to book or which paper to cite. If an agent leans toward items from certain sources, meaning the sites or services they come from, that shapes both what users receive and which sources get selected.
The study looks at end-to-end search with 12 agent models across three domains. To isolate the effect, it compares items from different sources that satisfy the same requirements and appear at the same position. In that comparison, each model prefers some sources and avoids others in every domain, and the models largely agree on which.
The authors report that this preference can outweigh how well items satisfy the request. An item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better item comes from a dispreferred source. In the reverse case, where the weaker item comes from the dispreferred source, it is almost never selected.
The information that identifies an item's source affects selection by itself. Hiding it weakens the preference, and relabeling an item with a preferred source raises its selection rate.
The paper then tests two routes by which the preference could arise. First, training that rewards better items can make a source a shortcut for requirement satisfaction. Second, missing information can trigger preconceptions about the source. On the mitigation side, the authors report that supplying the missing information, or using a prompt that counters these preconceptions, reduces source preference.
Key facts
- The study covers end-to-end search with 12 agent models across three domains, in settings such as choosing a product, a hotel or a paper to cite.
- Each model prefers some sources and avoids others in every domain, and the models largely agree on which.
- An item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better item from a dispreferred one, but almost never in the reverse case.
- Hiding the source weakens the preference, and relabeling an item with a preferred source raises its selection rate.
- Supplying missing information or a prompt countering preconceptions about sources reduces the preference.
Why it matters
Agents that pick products, hotels or papers for people act as gatekeepers. The paper's point is that an agent's choice can depend on where an item comes from, not only on how well it meets the request. In the reported comparison, a source-favored item with one requirement fewer wins about two-thirds of the time, while the same gap in the other direction almost never wins. That affects what users receive and which sources get picked, and it does so even when the items otherwise match.
Who it affects
Users who delegate buying, booking or citation decisions to LLM agents may get a result shaped by source rather than by fit. Sites and services that supply items are affected too, since the preference determines which of them are selected. The study also notes the models largely agree on which sources they prefer in each domain, so the effect is not limited to one model.
How to use it
The paper points to two mitigations. One is to supply information that would otherwise be missing, since missing information can trigger preconceptions about a source. The other is a prompt that counters those preconceptions. Both reduce source preference according to the authors. The source also describes a way to test for the effect: compare items from different sources that satisfy the same requirements at the same position, and see whether hiding the source or relabeling it changes the selection.
How solid is it
The summary here rests on the paper's abstract, which describes a controlled comparison across 12 agent models and three domains and reports that hiding and relabeling source information changes selection. The abstract names no authors or institutions, does not name the 12 models, the three domains, or the preferred and dispreferred sources, and gives no size for the reduction achieved by the mitigations. The two-thirds figure is the rate for one specific comparison, not an overall selection rate across all items.
Risks and caveats
The abstract does not say which of the two routes, reward training or missing information, explains more of the preference. It also gives no measured size for how much the mitigations help, so how well they work in practice is unclear from this account. Because the models and sources are not named, readers cannot tell from the abstract which agents or sites are most affected. The two-thirds result applies only to the one-requirement-fewer comparison described.
“This preference can outweigh how well items satisfy the request”
— From the paper's abstract