Inherent's Faraday agent outperforms Claude Opus 4.8 and GPT-5.5 at replicating research

Inherent, a London AI lab founded by Google DeepMind alumni, says its newly released AI agent Faraday beat much larger models from Anthropic and OpenAI at a specific task: independently reproducing the findings of published scientific papers without being told the answer in advance. The startup emerged from stealth just weeks earlier with a $50 million seed round. Cofounder and chief scientist Edward Hughes said paper replication is a standard training exercise for human scientists too, noting that many PhD students start their careers by doing exactly this. Beating other AI systems at the task was not the point, Hughes told TechCrunch; how Inherent got there was. Faraday was measured against Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5, both described as much larger, frontier-scale systems, while Faraday itself runs on a comparatively tiny model called Qwen 3.6 with 27 billion parameters, a fraction of the size of its rivals. Inherent's bar for success went beyond raw accuracy: it wanted Faraday to demonstrate what Hughes called 'research taste,' an instinct for which experiments are worth running and how to design them well. To teach that, Inherent trained Faraday primarily with reinforcement learning, a method that rewards the system for good outcomes rather than spelling out rules to follow, betting the approach will generalize better toward the company's longer-term goal of agents that can contribute across many scientific fields. Rather than build its own coding tool, Inherent had Faraday use OpenAI's GPT-5.5 Codex instead, the way human scientists lean on existing software rather than building everything themselves. Hughes said Inherent is also trying to avoid building agents that simply tell users what they want to hear, modeling the goal instead on a teammate who comes back and says they got curious, ran some experiments, and wants to know what you think of the results. The company's dozen employees all work in person from an office in King's Cross, the London neighborhood that Google DeepMind's presence helped turn into a major AI hub; Hughes said he believes London is the place to be. Hughes has also added his voice to calls to end 'garden leave,' the UK practice of barring departing employees from joining or starting a rival company for months after they resign, saying he was personally affected by it before starting Inherent alongside three other cofounders: Louis Kirsch, Kaloyan Aleksiev and Tantum Collins. Inherent plans to grow its headcount to about 20 to 25 people by the end of the year, a hiring push the article suggests could appeal to Google DeepMind staff unsettled by a new role recently taken on by Demis Hassabis, though that role itself is not described.
Key facts
- Inherent's AI agent Faraday, running on a 27-billion-parameter Qwen 3.6 model, outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at independently replicating the findings of published scientific papers.
- Inherent, founded by Google DeepMind alumni, raised a $50 million seed round just weeks before releasing Faraday.
- Faraday was trained mainly with reinforcement learning to develop 'research taste,' an instinct for which experiments to run and how to design them, rather than just accuracy.
- Instead of building its own coding tool, Inherent has Faraday use OpenAI's GPT-5.5 Codex.
- Inherent's dozen employees work in person from King's Cross, London, with plans to grow to about 20 to 25 staff by year end; cofounder Edward Hughes has criticized the UK's 'garden leave' noncompete practice.
Why it matters
A 27-billion-parameter model beating frontier-scale systems like Claude Opus 4.8 and GPT-5.5 at a scientific task, if it holds up, would be a data point that training method and task-specific reward design can substitute for raw scale on at least some research-adjacent work. Inherent frames paper replication as a stepping stone toward a much larger goal, an AI agent that can independently discover new scientific knowledge rather than just verify existing results, and treats reinforcement learning aimed at 'research taste' as the mechanism it is betting on to get there.
Who it affects
The comparison directly involves Anthropic and OpenAI, whose models were used as the benchmark Faraday beat. It matters to AI labs and investors watching whether smaller, cheaper models trained with targeted reinforcement learning can compete with frontier-scale systems on specific tasks. It also touches the London and UK AI talent market: Inherent staffs entirely in person out of King's Cross, is hiring toward 20 to 25 employees by year end, and its cofounder has publicly criticized the 'garden leave' noncompete practice that constrains researchers moving between UK AI companies.
How to use it
The source does not describe Faraday as a product available for purchase or public access, and gives no pricing, tiers or release channel; nothing here should be read as an offering people can sign up for. What is described is the underlying setup: Faraday runs on the Qwen 3.6 model and uses OpenAI's GPT-5.5 Codex for coding tasks rather than a tool Inherent built itself.
How solid is it
The claim comes from Inherent itself, relayed by TechCrunch through cofounder Edward Hughes, with no independent benchmark, published methodology or third-party verification cited in the article. No accuracy score, success rate or other quantified result is given for Faraday's performance, only the qualitative claim that it outperformed Claude Opus 4.8 and GPT-5.5; the specific papers, scientific fields or benchmark suite used are not named either.
Risks and caveats
Because the result is self-reported and unaccompanied by a number, an independent replication, or details of the test set, it is not possible to judge from the article how large or consistent the gap over Claude Opus 4.8 and GPT-5.5 actually is, or whether it would hold on a broader or independently chosen set of papers. The article also does not specify when Faraday was released beyond 'just weeks' after Inherent left stealth, and does not describe the new role of Demis Hassabis that it says has unsettled some DeepMind staff.
“Many PhD students actually start by doing this.”
— Edward Hughes, Inherent cofounder and chief scientist