FrugalEvo pairs a strong and a cheap LLM to cut the cost of program evolution

LLM-guided evolutionary methods such as AlphaEvolve have become a powerful way to attack hard computational optimization problems, circle packing among them. The authors of a new paper argue that prior work usually measures performance gain over a fixed number of iterations, when practical optimization should maximize gain per unit cost.
Their answer is FrugalEvo, a cost-aware evolutionary framework built on a division of labor between two models. A stronger, higher-cost LLM explores solution strategies. A cheaper LLM implements those strategies and iteratively refines the resulting code.
The framework also includes a cache-efficient evolution process. The harness and prompts are designed to maximize the sharing of prefixes across different evolution steps, which improves cache reuse.
To measure quality while spending is capped, the authors introduce Budget-Aware Area Under the Curve (BA-AUC). It is defined as the area under the best-so-far evaluation score curve over cumulative LLM cost, up to the budget.
Results are reported on 10 mathematical and systems optimization tasks. In final solution quality, FrugalEvo matches or surpasses state-of-the-art baselines including OpenEvolve, ShinkaEvolve, AdaEvolve and EvoX, and it achieves higher BA-AUC on 9 of the 10 tasks. On 10 algorithmic optimization tasks from ALE-Bench-Lite, it achieves higher average performance than the same baselines.
The headline result is circle packing. FrugalEvo reaches new state-of-the-art performance with GPT-5.6 Terra and Luna for only 1.68 USD, and with GLM-5.3 and its Flash variant for only 0.55 USD. That matches or surpasses all baselines, including the multi-agent methods CORAL and SwarmResearch, which cost approximately 50 USD on average.
Key facts
- FrugalEvo splits the work: a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them and iteratively refines the code.
- Prompts and harness are built to share prefixes across evolution steps, improving cache reuse.
- The paper introduces BA-AUC, the area under the best-so-far score curve over cumulative LLM cost, up to the budget.
- Across 10 math and systems optimization tasks FrugalEvo matches or surpasses OpenEvolve, ShinkaEvolve, AdaEvolve and EvoX in final quality, with higher BA-AUC on 9 tasks.
- On circle packing it sets a new state of the art for 1.68 USD (GPT-5.6 Terra and Luna) or 0.55 USD (GLM-5.3 and its Flash variant), versus about 50 USD on average for CORAL and SwarmResearch.
Why it matters
LLM-guided evolutionary search has delivered strong results on problems like circle packing, but the paper's point is that cost has been an afterthought: earlier work typically optimizes gain over a fixed number of iterations. FrugalEvo reframes the goal as gain per unit cost and shows, in the authors' results, that a state-of-the-art circle packing result is reachable for 1.68 USD or 0.55 USD, while multi-agent methods such as CORAL and SwarmResearch cost approximately 50 USD on average.
Who it affects
Researchers and engineers who run LLM-driven evolutionary or program-search systems on optimization problems and pay for every model call. Teams comparing methods such as OpenEvolve, ShinkaEvolve, AdaEvolve and EvoX are the direct audience, along with anyone benchmarking these methods under a fixed spending cap.
How to use it
The abstract describes the recipe rather than a release. Use a stronger, higher-cost model to propose and explore solution strategies, hand implementation and iterative code refinement to a cheaper model, and structure prompts so prefixes are shared across evolution steps to improve cache reuse. For evaluation, BA-AUC gives a way to compare methods under a fixed cost budget: the area under the best-so-far score curve over cumulative LLM cost. The abstract does not state that FrugalEvo is released as open source or give a code link.
How solid is it
The claims come from the paper's abstract, so they are the authors' own reporting. The evidence base is 10 mathematical and systems optimization tasks plus 10 algorithmic tasks from ALE-Bench-Lite, with comparisons against named baselines. FrugalEvo matches or surpasses baselines in final quality and has higher BA-AUC on 9 of the 10 tasks. The abstract does not say which task did not get higher BA-AUC, and it does not say whether the baselines were run with the same LLMs. The 50 USD figure is an approximate average for the multi-agent methods CORAL and SwarmResearch, not for all baselines.
Risks and caveats
The abstract gives costs but no numeric scores such as circle packing sum of radii or BA-AUC values, so the size of the quality margins cannot be judged from it. It does not say how much of the saving comes from the cheaper model versus cache reuse, and it does not state the cost budget used in the experiments. No authors or institutions are named in the abstract. The headline cost comparison is on one task, circle packing, with a result described as matching or surpassing the baselines.
“We argue that practical optimization should maximize gain per unit cost.”
— FrugalEvo paper abstract