Microsoft Research Podcast: Jennifer Neville on testing AI beyond benchmarks

This Microsoft Research Podcast episode is a conversation between host Chad Atalla and Jennifer Neville, a partner research manager at Microsoft Research who is also the Samuel D. Conte chair professor of computer science and statistics at Purdue. The episode description says Neville explores the role evaluation plays in pushing the performance boundaries of today's AI systems to meet user needs, and the "surprising failures" that emerge when models are tested beyond traditional benchmarks. It also says she shares practical guidance for working with current AI systems, explains why looking closely at data matters when results defy expectations, and reflects on what decades of AI progress have taught her about predicting what comes next. The text available covers the first part of the transcript, which is mostly about her career and her team's charter.
Atalla introduces her research as machine learning and AI for interactive domains and structured data, looking at how the data AI systems are trained on affects their behavior and how that matches what users want. He says she has published more than 130 papers with over 10,000 citations, and lists a National Science Foundation CAREER Award, a place on IEEE's "10 to Watch" in AI list, and best paper awards from the International Conference on Data Mining and the International Conference on Learning Representations.
Neville's route into AI was indirect. Growing up, she wanted to do anything but computer science, because her father worked in it. At college she majored first in math and then in physics, but could not settle in either. She dropped out for a while, returned to study cognitive science, and found it too "squishy" with not enough math. After working for a time, she went back again and chose computer science because she wanted to work with data. That is where she found AI, "just by happenstance." She jokes that if someone had told her AI combines cognitive science with math and computational thinking, she would have saved herself a lot of time.
Research came by a similar accident. Her computer science honors program required a research project, and she chose to look at data mining on interconnected web data. She was pointed to a faculty member who had just started working in a young field called statistical relational learning. She published a paper at a workshop and was hooked, even though she had not planned on grad school. What hooked her, she says, was the process: long stretches of spinning her wheels and feeling dejected, with a supportive adviser saying to keep going, then the thrill of understanding something no one else yet understood. She calls it "like a drug almost" and says it has kept her in research.
On academia versus industry, Atalla notes she joined Microsoft in 2021 from academia and has been teaching and advising for 20 years. Neville describes two sabbaticals. After getting tenure she went to the Simons Institute in Berkeley, a very theoretical stint that made it easy to return to Purdue. On the second she came to Microsoft Research and saw how algorithms behave in real systems with real users and real data, which she found hard to turn back from. She says the questions feel fundamentally similar in both places, but industry lets her apply work at scale on real systems, with products and company interests as the target. Academia tends toward abstractions that cover many domains at once, which is how funding from bodies like DARPA and NSF is won. She sees synergies and room to move between the two, but says that right now, for work on AI systems, "industry is really the place to be."
Atalla, who says he has been at Microsoft a little over six years and has worked on AI evaluation for much of that time, then asks about her team's charter. Neville says the group is called the AI Interaction and Learning team. It tries to push the frontier of AI system behavior in realistic work environments: where the performance boundary of current systems lies, how users experience it in real workflows, and how to improve models on complex tasks. Evaluation is the focus, she says, because the standard way ML and AI people evaluate is through benchmarks that are fairly simple compared with how people actually use these systems. So the team first asks what people want from the systems and designs practical evaluations to match, including multiturn behavior, collaborative environments and long-horizon tasks. The gaps those evaluations reveal show where the systems need theoretical or algorithmic improvement. The team starts with evaluation, she says, but the goal is better algorithms, models and estimation methods to raise model performance.
Key facts
- Jennifer Neville, a partner research manager at Microsoft Research and a professor at Purdue, is interviewed by Chad Atalla on the Microsoft Research Podcast.
- She leads the AI Interaction and Learning team, which studies AI behavior in realistic work environments such as multiturn, collaborative and long-horizon tasks.
- Her stated reason for focusing on evaluation: standard benchmarks are fairly simple compared with how people actually use AI systems.
- The team starts with evaluation but aims to develop better algorithms, models and estimation methods.
- She joined Microsoft in 2021 from academia and says that right now, for work on AI systems, industry is the place to be.
Why it matters
The episode puts a working researcher's case that simple benchmarks do not show how AI systems behave when people use them for real work. Neville argues that evaluations should reflect multiturn exchanges, collaboration and long-horizon tasks, and that the gaps they expose tell researchers where algorithms and models need to improve. The episode description also points to "surprising failures" that appear when models are tested beyond traditional benchmarks.
Who it affects
Most directly, researchers and engineers who evaluate or build AI systems, and people weighing a career in academia or industry research. Neville gives her own view on that choice. The description also says the conversation offers practical guidance for people working with current AI systems.
How to use it
It is a listening and reading resource: the episode comes with a transcript. Listeners designing evaluations can take her framing as a starting point: ask what you want out of the system, then build tests that put it in those conditions. The episode is part of the Microsoft Research Podcast, which the page invites readers to subscribe to.
How solid is it
This is an interview, so the claims are Neville's own views and experience, not a published study. Her credentials are given by the host: more than 130 papers, over 10,000 citations, and several awards. The text available covers the opening of the conversation. The "surprising failures" and the practical guidance are described only in the episode description, not in the part of the transcript available here, so no specific failures can be reported.
Risks and caveats
No benchmark names, model names or quantitative evaluation results appear in the text available, so nothing here can be checked against data. Her view that industry is the place to be for AI systems work is an opinion stated for the current moment. The host is a Microsoft colleague and the show is produced by Microsoft Research, so this is an in-house conversation.
“What we really focus on is trying to push the frontier of behavior of these AI systems in realistic work environments.”
— Jennifer Neville, Microsoft Research Podcast