Lemire proposes ephemeral testing: let AI agents build throwaway apps on your code

Daniel Lemire, in a blog post dated 5 October 2026, proposes a testing method he calls ephemeral testing. He says it was unthinkable before AI agents, and he defines "ephemeral" as a fancy word for "throw away" or "temporary".
The procedure is simple. You write your code or build your software component; the AI agent can do that too, he says, and it does not matter which. Then you ask an AI agent to build on it: an application, another layer, maybe several. You have the agent test what it built. You do not assess the original work directly. You assess how good the software built on top of it is.
Lemire calls this a form of integration testing. The difference is that the software on top is entirely ephemeral: you throw it away when you are done.
The quality of the foundation shows in how the exercise goes. A library with a clean API, stable invariants and useful errors lets the agent produce something that works quickly. A library with hidden state, surprising defaults or incomplete docs produces a pile of patches and failures. His point is that those failures are evidence about your code, not about the agent.
The test can be repeated with different agents and different tasks on the same foundation. In effect, he says, instead of building the core while trying to anticipate what might be needed at the other layers, you simulate those layers by actually building them.
He addresses one objection: with AI you could argue that you can rebuild everything whenever you need to. That is not practical, he writes, and you need some form of stability.
Lemire says he has been applying the trick to various projects. When he considers a new feature, he asks his AI to quickly prototype what he might later build on top of what he is doing. His verdict: ephemeral testing works for him thus far.
Key facts
- Ephemeral testing: have an AI agent build an application or further layers on top of your component, have it test what it built, then throw that software away.
- You assess the quality of the software built on top, not the original work directly; Lemire calls it a form of integration testing with throwaway software.
- A clean API, stable invariants and useful errors let the agent succeed quickly; hidden state, surprising defaults or incomplete docs produce patches and failures.
- The test can be repeated with different agents and different tasks on the same foundation.
- Lemire says it works for him thus far on various projects, including quick prototypes of features he might build later.
Why it matters
Lemire's idea changes who the first user of a component is. Instead of guessing what higher layers will need, you let an agent build those layers and see where it struggles. The cost of that experiment is low because the result is discarded. He presents it as a method that was unthinkable before AI agents could build applications on demand.
Who it affects
Anyone who writes libraries or software components and has access to an AI agent. The post frames it around code you write, or code the agent writes for you, and the layers that would sit on top of it. It also touches developers weighing a new feature, since Lemire uses the same trick to prototype what he might build later.
How to use it
Write or build your component. Ask an AI agent to build an application or another layer on top of it, perhaps several. Have the agent test what it built, and judge the result by how good that software is, not by inspecting the original work directly. Throw the layer away when you are done. Repeat with different agents and different tasks on the same foundation. Lemire also asks his AI to prototype quickly what he might later build when he considers a new feature.
How solid is it
This is a short opinion post resting on the author's own experience: Lemire says it works for him thus far. The post gives no benchmarks, measurements or numbers, and it does not name specific projects, libraries, agents or models. Treat it as a practitioner's proposal, not a validated method.
Risks and caveats
Lemire notes that you cannot simply rebuild everything whenever needed: that is not practical, and some form of stability is required. Since the post offers no evidence beyond his own use, how well the approach generalises is untested in the source. His own framing also ties the signal to the agent: failures count as evidence about your code, which is his claim, not something the post demonstrates.
“The failures are evidence about your code, not about the agent.”
— Daniel Lemire, "Ephemeral testing"