Essay: four years of AI coding tools, still no new Airbnbs

Jordan Andersen, writing on his consultancy's blog, responds to an earlier post by Disesdi Shoshana Cox that did the math on how many breakout tech companies AI should have produced by now. Borrowing that arithmetic, Andersen notes that after four years of open source LLMs and a claimed 10x improvement in development speed, the math implies the world should have three Airbnb-equivalents, two Stripe-equivalents and three Dropbox-equivalents. Instead, he argues, the biggest new companies since generative AI arrived are the AI companies themselves: OpenAI, Anthropic and High-Flyer (maker of DeepSeek), all valued at record sums, while the most visible effect on everyday life has been worse emails, more social media garbage and degraded search results.

To test the productivity claims himself, Andersen spent $10 on DeepSeek credits (bought by his co-founder Nik) on a real project. He describes the experience as infuriating: the chatbot suggested bad architectural choices, and while the resulting code ran, he compares it to a clown car held together by duct tape. He preempts the obvious rebuttal, that he simply prompts badly, by noting he has shipped a production product before, which he says puts him ahead of most people who attempt this.

He describes walking out of a conference panel titled something like No code, No problem, where four startup founders discussed having vibe-coded their way to an MVP and, he assumes, are already charging customers. The moment that drove him out was one founder stating he did not have, and would never need, a CTO because he could just use Claude to solve his technical problems. Andersen compares this to skipping a structural engineer for a building or a surgeon for an operation.

His broader argument is that writing lines of code was never the part of software development that consumed the most time, so removing that step does not remove the real bottleneck. He extends the argument beyond coding to information-seeking generally, citing a statistic from Semrush writer Carlos Silva that Reddit outranks financial experts 176% of the time when ChatGPT answers finance questions, despite guidelines meant to prioritize authoritative sources for topics that affect a person's money or life. He argues chatbots reproduce whatever a mix of mostly non-expert opinions says, without being able to verify any of it, and that the danger multiplies once non-experts use these tools to build apps that collect payment details and personal data. He closes by warning that GenAI can help develop an idea but cannot substitute for expertise, and that AI will only make a bad developer produce bad work faster.

Key facts

  • Borrowing arithmetic from an earlier post by Disesdi Shoshana Cox, the author argues that four years of open source LLMs and a claimed 10x development speedup should have produced three Airbnb-equivalents, two Stripe-equivalents and three Dropbox-equivalents, and asks where they are.
  • He argues the biggest new companies since generative AI arrived are the AI companies themselves, naming OpenAI, Anthropic and High-Flyer (developer of DeepSeek) as being valued at record sums.
  • The author spent $10 on DeepSeek credits (bought by his co-founder Nik) on a real project and describes the resulting code as running but held together like a clown car with duct tape.
  • At a conference panel titled something like No code, No problem, one of four vibe-coding startup founders said he did not have, and would never need, a CTO because he could just use Claude to solve his technical problems.
  • Citing Semrush writer Carlos Silva, the essay states Reddit outranks financial experts 176% of the time when ChatGPT answers finance questions, despite guidelines meant to prioritize authoritative sources.

Why it matters

The essay pushes back on a specific, quantifiable claim: that a 10x AI-driven productivity gain should have already produced a fresh crop of Airbnb- or Stripe-scale companies. Instead, the author argues, the only new giants are the AI vendors themselves, and the visible downstream effects are degraded email, social media and search quality rather than a wave of new products.

Who it affects

The piece targets startup founders and executives who treat AI coding tools as a substitute for technical hires, and by extension their customers and employees, whose money, data and jobs the author says are put at risk when nobody with software expertise is checking the work. It also implicates any user who defers to a chatbot's answer over an expert's, in fields from software to home electrical wiring to running programs.

How to use it

The author's own practice is narrower than the industry hype: he uses AI agents for boilerplate code and repetitive SQL, not as a replacement for engineering judgment. His argument is that AI tools are useful for mechanical, low-stakes tasks but should not be trusted to make architectural decisions or stand in for a CTO on a production system handling payments or personal data.

How solid is it

The piece is a personal essay, not a study: its central data points are a single $10, informal DeepSeek trial with no benchmark or comparison group, one anecdote about a conference panel, and one third-party statistic (the Reddit-vs-experts figure attributed to Semrush writer Carlos Silva) cited without a link to the original study. The core claim, that headline-scale AI-driven startups have not materialized, is presented as an observation rather than something formally measured.

Risks and caveats

The essay is an opinion piece and reads as such throughout, including profanity and rhetorical exaggeration; its numbers on Airbnb, Stripe and Dropbox's time to IPO or profitability are given as footnoted context for the borrowed arithmetic, not as claims the author independently verified. The DeepSeek experiment is a single anecdote from one user on one project, and the conference panel and its founders are not named, so neither can be checked against other sources.

“Reddit outranks financial experts 176% of the time when ChatGPT answers finance questions, despite YMYL guidelines prioritizing authoritative sources.”

— Carlos Silva, writer for Semrush, as cited in the post