RealCompanion benchmark finds AI companions rarely need past messages

RealCompanion benchmark finds AI companions rarely need past messages

A companion that talks with a person for months should understand them: remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing that needs a real person's record, and such records are private. The authors say existing benchmarks therefore generate the person and the questions and settle in advance what matters.

This paper takes the other route. It releases a benchmark (called RealCompanion in the title, and \bench in the abstract text) made of ten real relationships between people and an AI companion: 27,218 messages over up to 120 days. The release is the conversation itself plus four files derived from it: a profile, a persona, a chat ground truth and a question set. Each file cites the messages it rests on. Every chat label also carries the reasoning trace that produced it, checked stage by stage against the conversation.

Three findings follow.

First, the past is rarely needed and far away, and pooled measures mislead. A recency window finds the required message for 95.9% of probes pooled together, but for only 2.2% of the probes that actually need memory. At the natural rate, 96% of the gain from supplying recorded evidence comes from messages that need no memory at all.

Second, no detector the authors tried can tell when memory is needed on real messages. Questions written by authors over the same histories leak the cue. Labeling the same messages as memories raises their use by ten to fourteen points.

Third, three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.

Key facts

  • The benchmark covers ten real relationships with an AI companion: 27,218 messages over up to 120 days, released with a profile, a persona, a chat ground truth and a question set, each citing its source messages.
  • A recency window finds the required message for 95.9% of probes pooled, but only 2.2% of those that need memory.
  • At the natural rate, 96% of the gain from supplying recorded evidence comes from messages that need no memory.
  • No detector the authors tried can tell when memory is needed on real messages; labeling messages as memories raises their use by ten to fourteen points.
  • Three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.

Why it matters

Companion products are judged on whether they come to understand a person over months, but the private records needed to test that are rarely available. The authors argue that existing benchmarks work around this by generating the person and the questions and deciding in advance what matters. A set built from real conversations, with every label tied to the messages and a reasoning trace behind it, lets that assumption be checked. The headline result is that the past is rarely needed and far away, so scores pooled over all messages can look good while hiding how a system does on the few messages that truly need memory.

Who it affects

Teams building AI companions and long-term memory features, and researchers who design memory benchmarks. The findings bear on anyone who justifies a memory component with pooled accuracy gains, since the authors report that 96% of the gain from supplying recorded evidence at the natural rate comes from messages that need none. Anyone choosing between agent systems for building a persona is also affected: the paper reports the same F1 across three systems at a 31-fold cost difference.

How to use it

The authors say they release the conversations and the four derived files. They frame the results as a way to test memory on real messages rather than authored questions, and to compare agent systems on persona reconstruction with cost in view. The source text does not say whether the dataset is publicly downloadable, nor under what license or privacy protections beyond "We release", so check the paper page before planning to use it.

How solid is it

This account rests on the paper's abstract alone. The design is concrete: real conversations, derived files that each cite their source messages, and chat labels whose reasoning traces were checked stage by stage against the conversation. The scale is small, though: ten relationships, up to 120 days. The abstract does not explain the 95.9% and 2.2% figures further, such as the number of probes or the window size, and it does not give the F1 value or the dollar costs behind the 31-fold difference.

Risks and caveats

The abstract does not name the AI companion product, the detectors, the models or the three agent systems, and it does not say which system is cheapest or most expensive, so the findings cannot be tied to specific products from the abstract alone. Ten relationships is a small sample, and it is not stated how participants were recruited. The claim that no detector can tell when memory is needed covers only the detectors the authors tried. Read the pooled and memory-only figures separately: 95.9% and 2.2% describe different sets of probes.

“three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost”

— RealCompanion paper abstract