VoxPolyMem gives agents long-term memory for multi-party spoken talk

The paper starts from a gap. Long-term memory lets agents accumulate information and reason across sessions, but existing research, the authors say, mostly covers dyadic conversations in text or image-text form. Long-term memory for multi-party spoken conversations is underexplored. That setting asks for three things at once: preserving what was said, identifying participants across sessions, and retaining who speaks to whom.
The authors' answer is VoxPolyMem, an interaction-aware multimodal memory framework. It combines incremental speaker identification with a memory hierarchy of three layers: interaction memory, fact memory, and participant profiles.
Retrieval is framed as sequential decision-making. An agent rewrites queries and chooses which retrieval tools and memory layers to use, based on the evidence it has gathered so far, in order to close information gaps. To train this behaviour the authors introduce Evidence-Gain GRPO (EG-GRPO), which uses round-wise credit assignment to encourage the agent to acquire complementary evidence.
The authors also build VoxPolyBench, a benchmark for multi-party spoken conversations. It evaluates memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution.
On VoxPolyBench, VoxPolyMem reaches an overall score of 85.0, surpassing the strongest evaluated baseline by 23.6 points. On two other benchmarks, Mem-Gallery and H2HMem-Multi, it scores 89.6 and 74.4 respectively, exceeding the strongest evaluated public memory baselines by over 8 points each. The authors conclude that the results highlight the framework's potential for persistent, personalized assistance in multi-party multimodal interactions. Code and datasets are available at https://voxpolymem.github.io/VoxPolyBench/demo/.
Key facts
- VoxPolyMem is an interaction-aware multimodal memory framework for multi-party spoken conversations, pairing incremental speaker identification with three memory layers: interaction memory, fact memory and participant profiles.
- Retrieval is treated as sequential decision-making: an agent rewrites queries and picks retrieval tools and memory layers from accumulated evidence, trained with Evidence-Gain GRPO (EG-GRPO) and its round-wise credit assignment.
- The new VoxPolyBench covers memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution.
- VoxPolyMem scores 85.0 overall on VoxPolyBench, 23.6 points above the strongest evaluated baseline.
- On Mem-Gallery and H2HMem-Multi it scores 89.6 and 74.4, over 8 points above the strongest evaluated public memory baselines on each.
Why it matters
Agent memory work has mostly assumed a conversation between two parties, in text or image-text form. Real group talk adds a harder problem: the system has to remember not only what was said but who said it and to whom, and recognise the same people in later sessions. This paper targets that setting directly and supplies both a method and a benchmark for it. The reported margins, 23.6 points on the authors' own benchmark and over 8 points on each of two existing ones, suggest the approach is worth a look, though the abstract does not name the baselines.
Who it affects
Mainly researchers and engineers building long-term memory for conversational agents, especially those working on spoken, multi-speaker or multimodal settings. The authors point to persistent, personalized assistance in multi-party multimodal interactions as the intended direction.
How to use it
The authors say code and datasets are available at https://voxpolymem.github.io/VoxPolyBench/demo/. Practitioners can use VoxPolyBench to test memory systems on memory evolution, personalized answering, retrieval and reasoning, and interaction attribution. The design ideas that carry over are the three-layer memory (interaction, fact, participant profile) and retrieval as an agent that rewrites queries and chooses tools and layers each round.
How solid is it
The claims come from the paper's abstract, and the numbers are the authors' own. VoxPolyBench was built by the same authors who propose VoxPolyMem, so the headline 85.0 is on their own benchmark. The two outside results, 89.6 on Mem-Gallery and 74.4 on H2HMem-Multi, are compared against the strongest evaluated public memory baselines. The abstract gives only margins (23.6 points; over 8 points each), not baseline scores, and it does not name the baselines or say how the overall score is computed.
Risks and caveats
The abstract reports no real-world deployment or user study; "potential" for persistent, personalized assistance is the authors' wording. It gives no latency, cost or memory-size figures, and no model sizes, underlying LLM or speech models, or training data and compute. It does not say whether Mem-Gallery and H2HMem-Multi are multi-party or spoken benchmarks, so how far the gains on those two transfer to the paper's target setting is unclear. A system that identifies and profiles participants across sessions is also the kind of thing whose practical use would need scrutiny, but the abstract itself raises no such concerns.
“These results highlight its potential for persistent, personalized assistance in multi-party multimodal interactions.”
— Paper abstract