2026 in LLMs so far: coding agents, OpenClaw and tokenmaxxing

2026 in LLMs so far: coding agents, OpenClaw and tokenmaxxing

On a Friday, the author gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose and then published annotated slides and notes. The talk is a chronological tour of 2026 in LLMs, told from a working developer's point of view.

The story starts in November 2025 with two models, Claude Opus 4.5 and GPT-5.1. The author calls them incremental improvements, but says that every so often a model crosses an invisible line where something that did not really work starts working. Here it was coding agents. Claude Code had been around since February 2025, and Codex was a little younger. Paired with their agent harnesses, the two models went from "often make mistakes" to "reliable enough to use on a day-to-day basis". The author's long-running pelican-on-a-bicycle SVG test still showed weak results at that point: Claude could not really draw a bicycle, and the GPT-5.1 frame was poor too. Also in November came the first commit to an obscure GitHub repository called "Warelay".

Developers tinkered with the new combinations over the December holidays and came back in January excited. The author broke a lifelong New Year's resolution to stay focused and instead decided to take on as many new projects as he liked, with "be more ambitious" as the year's theme. On the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal, the author had predicted that "it will become undeniable that LLMs write good code" and that sandboxing would finally be solved, and says he thinks the code-quality prediction holds now; on sandboxing, he counted around 40 of the 277 sessions at the conference touching on sandboxing or agent security, so people are at least putting a lot of effort into it. A predicted "Challenger disaster" for coding agent security has not happened in the form predicted, meaning hijacked agents causing real-world economic damage, though there has been a lot of noise about agent security. The author also predicted a strong breeding season for New Zealand's kākāpō parrots, of which there were only 236 at the start of the year. The Rimu trees they depend on had not had a big fruiting season in four years, and this year's looked excellent. The podcast also produced a term, credited to Adam, for the "AI induced ennui where software engineers get listless because the AI can do anything": Deep Blue. The author says it has been a major theme all year.

In January the author also suffered what he calls AI mania, where any time an agent is not building something feels like wasted time. The result was over-ambitious projects: a JavaScript interpreter written in Python, vibe-ported from Fabrice Bellard's MicroQuickJS, and a WebAssembly runtime in Python. Looking at them, he asked whether the world needs a slow, buggy, half-baked Python JavaScript interpreter, and decided it does not. A browser playground survives, running the interpreter in Python, in Pyodide, in WebAssembly, in JavaScript, in a browser.

By the end of January, the "Warelay" repository had been renamed to CLAWDIS, then CLAWDBOT, then Moltbot, and finally OpenClaw. It had 8,300 commits less than two months after starting, and the author says it is now over 100,000. He calls it the most vibe-coded piece of software in existence. It kicked off a new category of software he likes to call "Claws" (OpenClaw, NanoClaw, IronClaw, PicoClaw), now often rebranded as "personal agents" or "general agents". The author says Apple stores in the Bay Area sold out of Mac Minis because so many people were buying them to run OpenClaw. Drew Breunig's explanation, as the author relays it, is that an OpenClaw is a digital pet and the Mac Mini is its aquarium. Also in January came MoltBook, a social network where people sent their Claws to talk to other Claws. It launched on Thursday, blew up on Friday, was profiled by the New York Times on Monday, and by Tuesday was drowning in slop and spam. Facebook/Meta bought it a month later.

In February, StrongDM described its Software Factory, which Dan Shapiro called the Dark Factory. The company had followed two rules since July of the previous year: code must not be written by humans, and code must not be reviewed by humans. The author says the first sounded radical in February but many in the room now live it, and the second has stayed a huge topic, with many conference sessions about code review. He notes that StrongDM is a security company with people of decades of experience on the project, and that they were living about six months ahead of everyone else, working out how to be confident in software quality without reading the code. Other February items: the first kākāpō chick in four years hatched on Valentine's Day, and Google released Gemini 3.1 Pro, which drew a good pelican on a bicycle, with the chain in the right place, feet on both sides and a fish in the basket. Google's Jeff Dean then tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro with an animated pelican on a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard and a dachshund driving a stretch limousine. The author says Google trained for all forms of animals on all forms of transport and has defeated his benchmark.

February was also when tokenmaxxing began. Headlines said Meta was making AI adoption part of performance reviews, Microsoft wanted every employee to use AI, and Uber boasted that 90% of its engineers were using AI workflows. A few months later Meta was cracking down on token use, Microsoft was saying tokenmaxxing is "not what we are optimizing for", and Uber was capping employee AI spending. The author's reading: it went straight up and straight back down because agents are expensive. Last year it was hard to spend more than $50 on AI tokens; now, he says, you can spend $1,000 in a day doing real work. He adds that this is also why Anthropic's valuation skyrocketed to "maybe a trillion dollars", and concludes that AI appears to have hit product market fit in 2026, primarily through coding agents.

March was peak OpenClaw. In China, companies hosted OpenClaw install parties where non-technical people queued around the block for help getting Claws onto their devices. The author takes this as proof of real demand for personal AI agents, and says a Claw is really just a coding agent wearing a less threatening hat: under the hood it writes and runs code on your computer. The race was then on to build the first safe Claw, one regular people could use without shooting themselves in the foot. The text available for this retelling cuts off mid-sentence in the March section, at a mention of Meta, so later months of the talk are not covered here.

Key facts

  • The author says Claude Opus 4.5 and GPT-5.1 (November 2025), paired with their coding agent harnesses, went from "often make mistakes" to "reliable enough to use on a day-to-day basis".
  • OpenClaw, first committed to as "Warelay" in November 2025 and renamed four times by the end of January, had 8,300 commits in under two months and is now over 100,000; the author calls it the most vibe-coded software in existence.
  • StrongDM's Software Factory followed two rules since July of the previous year: code must not be written by humans and must not be reviewed by humans.
  • Tokenmaxxing rose in February and then fell as Meta cracked down on token use, Microsoft said it is "not what we are optimizing for" and Uber capped AI spending; the author says $1,000 a day of tokens is now possible where $50 was hard to reach last year.
  • The author's verdict: AI appears to have hit product market fit in 2026, primarily through coding agents.

Why it matters

This is a first-hand, dated timeline of how coding agents went from promising to routine, told by someone who built with them all year. The author's central claim is that the November 2025 releases of Claude Opus 4.5 and GPT-5.1, paired with agent harnesses, crossed a reliability line, and that almost everything later in the year followed from it: the OpenClaw boom, factory-style development without human-written code, and heavy token spending. He sums it up as AI hitting product market fit in 2026, primarily through coding agents. Most milestones here were reported at the time; the value is the connected narrative.

Who it affects

Software engineers and engineering managers are the main audience. The author says the pace of change is unlike anything in his career and that he is working to come to terms with what it means for his profession; he describes the Deep Blue feeling of listlessness as a major theme, raised by several speakers at the conference. Companies setting AI policy are affected too: Meta, Microsoft and Uber all pushed AI adoption in February and then reined in token use or spending. Non-technical people also appear, in the form of the China OpenClaw install parties.

How to use it

The talk is a narrative rather than a product, but it contains practical pointers. The author's own approach this year was to take on more ambitious projects with coding agents in order to find where they stop working, and he warns that this can tip into what he calls AI mania, where idle agents feel like wasted time. StrongDM's case shows one extreme: route all code through a coding agent and put the effort into verifying the software's quality without reading the code. The author's annotated slides and the talk video are the original material.

How solid is it

This is one practitioner's account in a keynote, with slides and notes, not a study. Most claims are attributed to the author and rest on his observation: the Mac Mini sell-out, the China install parties and the tokenmaxxing reversal come with no figures beyond his statements. The Anthropic valuation is hedged by the author as "maybe a trillion dollars" and no source is given. The commit counts and the 40 of 277 session count are his own. The text available for this retelling is cut off mid-sentence in the March section, so the rest of the talk is not covered.

Risks and caveats

The author himself flags limits: the "Challenger disaster" he predicted for coding agent security has not played out as predicted, but agent security noise is high, and the race in March was to build a safe Claw that regular people could use without hurting themselves. MoltBook is offered as a cautionary example of agents talking to agents, drowned in slop and spam within days. StrongDM's no-review rule is presented as an exploration of what is possible and responsible, done by a security company with experienced people, not as a general recommendation. Tokenmaxxing showed that agent use gets expensive. The author also jokes that his own AI mania produced projects the world does not need.

“AI appears to have hit product market fit in 2026, primarily through coding agents.”

— From the author's annotated notes for the WeAreDevelopers keynote