Prime Intellect rewrites Prime Agent in Rust with a swarm of over 2,000 agents

Prime Intellect has shipped a new version of Prime Agent, its coding agent harness, rewritten from the ground up in Rust. The company says Prime Agent has been downloaded more than 300,000 times since its August launch and has processed over 8 trillion tokens. The rewrite was done largely by Prime Agent itself: over two weeks it orchestrated a swarm of over 2,000 agents that rewrote the TypeScript codebase end to end, operating across 10,000+ Prime Sandboxes and consuming over 200 billion tokens from Prime Inference's GLM-5.3 endpoint. All orchestrators and their agents ran on two 8-core on-demand CPU nodes, each supporting over 100 concurrent subagents and their CPython kernels, while heavy work such as compilation, type-checking and diffing was pushed out to Prime Sandboxes.
The company's case for Rust: TypeScript got the product out quickly, but its types are optional and vanish at runtime, errors travel as unchecked exceptions, CPU-heavy work like rendering and parsing large sessions competes with keyboard input on a single event loop, and every process pays for a JavaScript runtime and garbage collector. Prime Intellect says native code without a garbage collector accounts for most of the memory and startup gains, that Rust's Send and Sync traits let the compiler check what can move between threads, and that exhaustive enums, ownership and Clippy's pedantic lints rule out whole classes of bugs, which matters when agents write most of the code. The rewrite also brought Windows support, session crash isolation and a more consistent daemon protocol.
The goal was for agents to do the rewrite with as little human intervention as possible, so the human work was building verification. Four kinds of parity were checked. TUI parity: a differential test suite runs the TypeScript and Rust binaries side by side against the same scripted model and diffs the terminal frames. Harness parity: session transcripts and the requests each binary sends to the model provider are compared. Protocol parity: all daemon protocol message types are checked against the TypeScript implementation. Feature parity: agents audited the TypeScript product component by component, classifying each as matching, partial or missing. With an objective check for each, the agents could measure their own progress and catch regressions, which let the team cut back on human review.
A single root agent split the work into a topological ordering of tasks and wrote no product code, leaving it free to monitor tasks, merge finished work and work with the humans on priorities. Each task passed through four agents: a Planner that writes the specification and the parity check, an Implementer that writes Rust in its own worktree, a Reviewer that inspects the pull request adversarially using a different model and a separate context, and a Verifier that runs the parity checks in a fresh Prime Sandbox. A failure sends the feature back to the Implementer; the PR merges when both review and verification pass. Prime Intellect says it is building this pattern into Prime Agent as a factory of finite state machines so such workflows can be defined once and reused.
The new architecture is split into nine crates with a one-way dependency graph that Cargo enforces. The largest source file is about 2,500 lines, down from about 15,000 in TypeScript, and no file exceeds 5,000 lines, against four in TypeScript. Each session runs in its own worker process under a small supervisor, so one failure leaves other sessions running and sessions persist on disk for reattachment. Transport, process control and file locking sit behind platform-specific interfaces, the client, daemon and workers share one set of protocol message types, and model and MCP lists moved to a catalog fetched at runtime, so new models and plugins can ship without a Prime Agent release.
Parity was not the end. The company says the parity checks only verified the behavior they exercised. Once the main loop reached parity, the rewrite agents themselves moved onto the Rust build, then the internal team did, and that dogfooding exposed bugs and missing behavior outside the differential tests. Agents reviewed logs and traces from beta users to diagnose issues. Getting to release quality took weeks of follow-up work that Prime Agent still wrote and tested, with humans finding problems, directing changes and reviewing all results.
A second orchestrator then ran a performance hillclimbing loop for three days. A benchmark harness runs Prime Agent in Rust and TypeScript, plus other harnesses, on a fresh 4-core, 8 GB Prime Sandbox per benchmark, with noise checks that withhold unstable results; agents run in a real terminal against a scripted model, so timings exclude inference. Each experiment profiled the benchmarks, built the current code and the candidate change on the same sandbox, ran them in alternating order, and had two reviewer agents, each on a different frontier model, check parity before merging. The loop was deliberately given no numeric targets. It logged over 144 experiment and audit records and merged over 69 valid changes. Most of the gains came from moving work off startup and render paths, replacing polling loops with event-driven waits, and releasing memory as soon as large sessions finished loading.
The headline result: time to input roughly 14x faster than TypeScript and over 80% less memory after startup, which the company says puts Prime Agent amongst the fastest coding agent harnesses available. It adds that external harness results come from its own custom runtime suite and that, without a common benchmark standard, comparisons should be read with caution. Next, Prime Intellect says it will focus on capabilities and evals, make the multi-agent workflows behind the rewrite available to users, and tie Prime Agent more closely into the Prime Intellect ecosystem.
Key facts
- Prime Agent was rewritten from TypeScript to Rust over two weeks by a swarm of over 2,000 agents orchestrated by Prime Agent itself, across 10,000+ Prime Sandboxes and over 200 billion tokens from Prime Inference's GLM-5.3 endpoint.
- Each task went through a Planner, Implementer, Reviewer (a different model, separate context) and Verifier, checked against four kinds of parity with the TypeScript version: TUI, harness, protocol and feature.
- A three-day agent-driven hillclimb logged over 144 experiment and audit records and merged over 69 changes; the company reports time to input roughly 14x faster than TypeScript and over 80% less memory after startup.
- The codebase is now nine crates; the largest file is about 2,500 lines versus about 15,000 in TypeScript, and the rewrite added Windows support and session crash isolation.
- The benchmarks come from the company's own custom suite, and it says comparisons with other harnesses should be interpreted with caution.
Why it matters
This is a concrete, large example of a coding agent porting its own product to another language with a swarm of subagents, and the company frames it as a stress test of its sandbox and inference infrastructure for large-scale agent swarms, autonomous research and reinforcement learning. The method is as notable as the result: objective parity checks of four kinds let agents measure their own progress and catch regressions, so human review could be reduced, and generation was kept separate from verification because an agent that writes code is biased when judging it. The claimed outcome is roughly 14x faster time to input and over 80% less memory after startup.
Who it affects
Users of Prime Agent, which the company says has been downloaded more than 300,000 times, get the Rust build, including Windows support, session crash isolation and a more consistent daemon protocol. Teams that run coding agents or plan large agent-driven migrations can read the workflow as a worked example: a root orchestrator, a four-stage task pipeline and parity suites built before the work began. The post also points to tighter links between Prime Agent and the Prime Intellect ecosystem, including cloud agent swarms.
How to use it
Prime Intellect says the faster, cleaner Prime Agent is shipped today. The post gives no install steps or pricing, so there is nothing to add on cost. For builders, the transferable pattern is explicit: write a differential test suite against the old implementation, split work into a dependency-ordered task list, keep the implementer and an adversarial reviewer on different models with separate contexts, verify in a fresh sandbox, and let a benchmark harness drive optimisation without numeric targets. The company says it is building this pattern into Prime Agent as a factory of finite state machines and plans to make the multi-agent workflows behind the rewrite available to users. Model and MCP lists now come from a catalog fetched at runtime, so new models and plugins can arrive without a release.
How solid is it
This is a first-party engineering post, so the account of the process and the numbers are the company's own. The performance figures come from its custom runtime suite, and the company itself warns that without a common benchmark standard, comparisons should be interpreted with caution. The text states the 14x figure as time to input against TypeScript without saying whether it is cold or warm start, and gives no absolute milliseconds. The charts' underlying numbers are not in the text, and the other harnesses compared are not named. No independent third-party verification of the benchmarks is mentioned. The process counts, such as over 2,000 agents, over 144 records and over 69 merged changes, are stated as lower bounds.
Risks and caveats
The rewrite was not fully autonomous: humans set up verification, worked with the root agent on priorities and decisions, and reviewed results, and release quality needed weeks of follow-up with humans in the loop. The parity checks only verified the behavior they exercised, and remaining bugs and differences were largely found through internal use and beta-user traces. The strong claim 'amongst the fastest' sits next to the softer 'faster and uses fewer resources than most coding agent harnesses', so the comparison is not against everything. No cost, human-hours or reviewer count is given, though the run used over 200 billion tokens and 10,000+ sandboxes, which suggests this approach needs substantial infrastructure. The Reviewer models are described only as different frontier models, not named.
“We deliberately gave the loop no numeric targets, since a fixed threshold tends to become a stopping point.”
— Prime Intellect, 'Rewriting Prime Agent in Rust'