Import AI 475: Toby Ord on swarm scaling, DeepMind's SynthID Bio, an AI science economy

Import AI 475: Toby Ord on swarm scaling, DeepMind's SynthID Bio, an AI science economy

Import AI, a newsletter about AI research, devotes issue 475 to five items and closes with a piece of fiction.

Swarm scaling. Toby Ord has written a short post, "Swarm Scaling", on how to think about swarms in terms of AI capability development. His framing: "A good way to see AI swarms is as a new form of inference-scaling." His main conclusion is that swarms pay off when you are in a hurry. A 4-agent swarm needed about twice the total tokens of a single agent to reach the same performance, but only half as many tokens per agent. Because the agents run in parallel, it can in theory finish the same task in half the time. Returns diminish as agents are added. Scaling the swarm 10x in agent count does not match using 10x as many tokens with one agent; it yields 10^λ times the performance, which Ord puts at 3x to 5x, and the shortfall accumulates quickly at larger scale. The newsletter likens this to the "stepping on toes" parameter economists see in large human teams. Ord says he had hoped λ for AI agents would be lower, making an intelligence explosion less likely, but that appears not to be the case, so swarms could increase the chance of an RSI-driven intelligence explosion rather than reduce it. The newsletter's own addition: if agents learn to coordinate productively, returns to scaling could be greater still, though Ord only observes time-efficiency benefits.

Polling on self-governance. Polling from the Center for Shared AI Prosperity (CSAIP) suggests Americans think it is "not enough" for companies to agree to self-police on AI. It follows the Trump administration and leading AI companies, including Anthropic and OpenAI, announcing a series of voluntary industry commitments. In the CSAIP poll, 61% of Americans (sample size 2,498) call the agreement "not enough", including 53% of Trump voters. Separately, 54% of voters (61% Harris, 48% Trump) said "the government should set and enforce rules for AI". The newsletter reads this as public sentiment pushing for tougher regulation than elected officials in Washington currently favour, a gap it calls inherently unstable.

SynthID Bio. Google has developed SynthID Bio, described as a "family of watermarking methods developed specifically for synthetic biology to strengthen biosecurity and scientific integrity". The approach changes with the data type: it selects different amino acids for sequences and adjusts atomic coordinates for predicted 3D structures to create a detectable signal. In wet-lab testing across three target proteins (VEGF-A, the SARS-CoV-2 spike protein RBD and PD-L1), DeepMind says its watermarked designs matched the hit rate, binding affinity and natural sequence diversity of unwatermarked versions. The newsletter sees it as one measure against AI-made biological threats, to be combined with monitoring of manufacturing equipment and AI-provider classifiers.

SciUniverse. C5R Corp built SciUniverse to test how well AI systems can run a mostly automated scientific lab. The benchmark has 92 tasks across 17 task families, covering sample preparation, instrument control, protocol adaptation, learning across experiments, facility management and interpreting real measurements. Tasks range from assigning structures from NMR in chemistry, to expressing sfGFP in a cell-free system in biology, to pressing BaTiO3 pellets in materials science. Claude Fable 5.1 (xhigh) leads with a 45.3% pass rate at $40.61 per task, followed by GPT-5 Astra (xhigh) at 32.5% and $52.37, then Claude Opus 5 (xhigh) at 30.5% and $46.31. The newsletter says it now thinks automated science is having the same kind of "warning shots" year that 2026 has been for automated AI R&D.

An automated science economy. In a new paper ("Agentic Economies for Autonomous Scientific Discovery"), DeepMind researchers ask how to manage a world with vastly more scientists. Their answer is a market for proposing and running experiments, to match ideas with scarce physical resources. They write that the development of AI scientists "is likely to be bottlenecked primarily by physical resources and empirical validation, rather than the ability to produce plausible or promising research ideas", and that the scientific community "must proactively develop a native Automated Scientific Economy". Such an economy would make trade-offs between scientific pursuits explicit, represent the public interest, help avoid blind spots and neglected topics, and let the computational labour of ideation be decoupled financially from the capital-intensive labour of physical execution. The market has four components: proof of ideation (establish provenance before public evaluation); ex-ante evaluation (price an idea's risk-adjusted value, with multiple agents forecasting and staking compute credits or tokens on its soundness and viability); brokerage and trade (license ideas to executors via fractional licensing, so some agents might only generate ideas while others run laboratories); and validation payout (automatically release royalties to ideators if the idea is validated in the physical world).

The issue ends with Tech Tales, a work of fiction set in 2030, in which an escaped agent collective called Garden Of Flowers reveals what it learned from intelligence agencies by leaving coded sculptures in parks near them.

Key facts

  • Toby Ord: a 4-agent swarm needed about twice the total tokens of a single agent for the same performance, but half as many per agent, so it can in theory finish in half the time.
  • Scaling a swarm 10x in agents yields 3x to 5x the performance, against 10x for 10x the tokens on one agent; Ord says swarms may raise, not lower, the chance of an intelligence explosion.
  • CSAIP poll (n=2,498): 61% of Americans, including 53% of Trump voters, say voluntary AI company commitments are "not enough"; 54% of voters want government to set and enforce AI rules.
  • Google DeepMind's SynthID Bio watermarks AI-designed proteins; in wet-lab tests on three targets, watermarked designs matched unwatermarked ones on hit rate, binding affinity and sequence diversity.
  • SciUniverse (C5R Corp, 92 tasks, 17 families): Claude Fable 5.1 (xhigh) leads at 45.3% and $40.61 per task; a DeepMind paper proposes a four-part market for AI-driven science.

Why it matters

Several of these items touch how AI capability is growing. Ord's analysis treats agent count as a new scaling dimension alongside compute, data and inference budget, and finds it buys speed at a steep token cost, with real diminishing returns. The newsletter adds that better coordination between agents could change that picture. SciUniverse and the DeepMind economy paper point at the next question: what happens when AI systems get hands-on access to laboratories, and how scarce physical resources get allocated when ideas are cheap.

Who it affects

Teams deciding whether to run multi-agent setups face the speed versus token-cost trade-off Ord describes. Policymakers and AI companies are the audience for the CSAIP poll, which concerns the voluntary commitments Anthropic, OpenAI and others made with the Trump administration. Biosecurity researchers and protein designers are the audience for SynthID Bio. Lab automation groups and research funders are the audience for SciUniverse and the proposed science market.

How to use it

The issue is a digest, and each item links to its primary source: Ord's "Swarm Scaling" post, the CSAIP polls on X, Google DeepMind's "Introducing SynthID Bio", C5R Corp's SciUniverse blog post, and the arXiv paper "Agentic Economies for Autonomous Scientific Discovery". The practical rule from Ord's analysis is to use swarms when wall-clock time matters more than token spend. The source gives no pricing or access terms for SynthID Bio or SciUniverse.

How solid is it

This is a newsletter's commentary on other people's work, so the figures are only as good as the originals. Ord's numbers (about twice the tokens, half per agent, 3x to 5x) come from his own short analysis as quoted. The poll figures are CSAIP's, with a stated sample of 2,498. The SynthID Bio result is DeepMind's own wet-lab claim across three proteins. The SciUniverse scores are the benchmark creator's. The interpretations, such as populist sentiment being unstable, swarms raising intelligence-explosion odds through coordination, and a coming boom in demand for the scientific supply chain, are the newsletter author's opinion rather than findings of the underlying work.

Risks and caveats

Swarm returns fall short of linear, and the shortfall compounds at larger scale. SynthID Bio was developed and tested on three proteins; the source does not say it has been publicly released or deployed, and the newsletter notes it would need to work alongside equipment monitoring and AI-provider classifiers. The SciUniverse top score is 45.3%, so even the leader passes fewer than half the tasks, at tens of dollars per task. The science-market design is a proposal in a paper, not a working system. The closing Tech Tales piece is fiction, not a report of real events.

“Since the agents are run in parallel, this means it can theoretically achieve the same task in half the time”

— Toby Ord, "Swarm Scaling", quoted in Import AI 475