AgenticBBO-Bench tests LLM agents on black-box optimization in five domains

AgenticBBO-Bench tests LLM agents on black-box optimization in five domains

Black-box optimization (BBO) covers scientific and engineering problems where each evaluation of the objective is expensive and the number of evaluations is limited. The paper argues that recent LLM agents offer a new way to approach it: they combine task semantics, computation, optimization tools and feedback-driven decision making, and the integration with mathematically rigorous tools shows great potential.

The problem the authors point to is comparability. Existing agentic BBO studies use different task domains and system configurations, so their results are difficult to compare and the effect of individual design choices is hard to isolate. To fix that, the paper introduces AgenticBBO-Bench, a cross-domain benchmark for agentic BBO. It spans five areas: synthetic functions, hyperparameter optimization, database tuning, chip design and molecular design, all under a unified finite-budget evaluation protocol.

In the experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains. It outperforms the best numerical optimizers in four of the five.

The authors then study three factors that shape agent performance: the optimization tools available, the task information and prior knowledge given to the agent, and the role of the LLM during search. Three findings come out of that. Additional numerical tools do not consistently improve performance. Task semantics are broadly useful, while more specific priors are less reliable. And numerical optimizers can effectively absorb gains from search trajectories established by the agent.

Finally, the paper adds a five-task frontier challenge inside AgenticBBO-Bench. It evaluates seven LLMs under the Codex agent harness. Among the evaluated models, GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost. The code is available at https://github.com/lamda-bbo/agentic-bbo.

Key facts

  • AgenticBBO-Bench is a cross-domain benchmark for LLM-agent black-box optimization covering synthetic functions, hyperparameter optimization, database tuning, chip design and molecular design under one finite-budget protocol.
  • Agentic BBO scores higher than direct LLM-based methods (family-averaged) in all five domains and outperforms the best numerical optimizers in four.
  • Extra numerical tools do not consistently help; task semantics are broadly useful, while more specific priors are less reliable.
  • Numerical optimizers can absorb gains from search trajectories established by the agent.
  • In a five-task frontier challenge with seven LLMs under the Codex agent harness, GPT-6 Astra and DeepSeek-V4.1-Flash sit on the performance and cost Pareto frontier.

Why it matters

Agentic approaches to black-box optimization have been studied on different tasks with different system setups, which makes it hard to say what actually helps. A shared benchmark with one finite-budget protocol gives those studies a common yardstick. The headline result is that agents beat direct LLM-based methods in all five domains and the best numerical optimizers in four, which suggests that wrapping an LLM in an agent with tools and feedback matters.

Who it affects

Researchers building or comparing LLM agents for expensive optimization problems, especially in the five areas the benchmark covers: synthetic functions, hyperparameter optimization, database tuning, chip design and molecular design. It also concerns anyone choosing between a classical numerical optimizer and an agent for such work, since the paper reports that agents outperform the best numerical optimizers in four of the five domains.

How to use it

The benchmark code is released at https://github.com/lamda-bbo/agentic-bbo. The paper's findings on design choices are practical starting points: do not assume more numerical tools will raise performance, give the agent task semantics, treat specific priors with caution, and consider handing the agent's search trajectory to a numerical optimizer. For model choice under the Codex agent harness, the paper names GPT-6 Astra and DeepSeek-V4.1-Flash as lying on the performance and cost Pareto frontier among the models it evaluated.

How solid is it

This is a preprint-style paper whose claims here rest on its abstract. The comparisons are reported qualitatively (higher, outperforms); no actual scores or margins of improvement are given. The abstract does not say which one of the five domains is the exception where agents do not beat the best numerical optimizers. The seven evaluated LLMs are not listed beyond GPT-6 Astra and DeepSeek-V4.1-Flash, and no cost or performance figures are given for the models on the Pareto frontier. The abstract does not state whether the work is peer-reviewed.

Risks and caveats

The findings are tied to the benchmark's five domains, its finite-budget protocol and the Codex agent harness used in the frontier challenge, so they may not carry over to other settings. The Pareto-frontier statement covers only the models evaluated, not models in general. The finding that additional numerical tools do not consistently improve performance means tool-heavy designs need testing rather than assuming a gain, and the weaker reliability of specific priors means injected domain knowledge can be a liability.

“additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable”

— AgenticBBO-Bench paper abstract