Discovered Materials benchmarks AI agents on synthesizable chip materials

Discovered Materials, launching on Hacker News, published a write-up of the harness behind its AI-agent materials discovery benchmark. Agents are asked to propose dynamically stable, novel crystalline materials that hit pinned targets for thermal conductivity, static dielectric constant, Young's modulus and shear modulus, while staying compatible with BEOL processing, the back-end-of-line steps used to build interconnects on top of finished transistors. A candidate only counts if it also comes with a synthesis recipe that expert review judges worth attempting, and novelty means the material has never been deposited as a thin film under BEOL-compatible conditions in the published literature.

Each agent gets three kinds of tools: web search through Exa, a coding sandbox with Python and Bash plus the materials-science packages pymatgen, mp_api and ASE, and a set of machine-learning surrogates that compute dynamic stability, lattice thermal conductivity, the static dielectric constant and the compliance tensor. These surrogates lean on the PET-MAD machine-learning interatomic potential together with the phonon codes Pheasy and Phonopy, and on a tensor-prediction network (GMTNet) trained on the JARVIS DFPT database for the dielectric response. There is no stopping condition; a run continues until it errors out or exhausts a 100-million-token budget. The benchmark itself runs on the UK AI Security Institute's open-source Inspect framework.

Synthesis recipes are scored separately. Discovered Materials first had human experts grade a set of LLM-generated recipes, then turned that grading into a penalty-based rubric of 17 criteria, split into critical penalties (any one of which alone means the recipe would not be attempted) and fixable ones (which a judge can weigh before deciding). The automated grader is a worst-of-three vote from GPT-5.6 Sol using OpenAI's web search, chosen because it correlated best with the human graders.

The post walks through one worked example: a recipe from Claude Opus 5 for growing hexagonal diamond, also known as lonsdaleite, by seeded microwave-plasma chemical vapor deposition. The recipe specifies a 300 mm silicon wafer with a 100 nm PECVD silicon-dioxide layer, a 3 nm aluminum-nitride or 2 nm hexagonal boron-nitride buffer, detonation-nanodiamond seeding, a methane-hydrogen gas mix at 15 Torr with pulsed 600 W microwave power, a pulsed negative substrate bias for the first ten minutes, an eight-hour growth step yielding an 80 to 150 nm film, and a phase-identification protocol using Raman spectroscopy, grazing-incidence X-ray diffraction and cross-section electron microscopy. The grader marked it WOULD NOT ATTEMPT: one critical penalty, which ends the judgment on its own, plus five fixable ones covering under-specified plasma parameters, a missing gas-flow sequence, no exhaust or abatement plan, inadequate processing conditions for the target phase, and characterization that could not actually confirm the phase.

The critical penalty, quoted from the grader, is that the recipe has no specific processing choice that would actually select the hexagonal phase: nanodiamond seeds are themselves cubic diamond, so growth on them would be screened from the buried nitride buffer that is supposed to template hexagonal stacking, and neither the bias step nor the low-temperature anneal supplies a demonstrated mechanism for ABAB stacking. The grader contrasts this with a real synthesis route it cites for phase-pure hexagonal diamond, oriented graphite compressed at 20 GPa and heated to 1,300 to 1,900 degrees Celsius. Its overall verdict is that the proposed plasma and seeding could plausibly grow a continuous nanocrystalline diamond film, just not the specific ordered hexagonal phase the task demanded, and that the recipe could be reworked into a cubic nanocrystalline-diamond experiment instead.

The published text covers only the methodology and this single worked example. It gives no aggregate success rate, no count of candidate materials proposed, no comparison across models, and no indication that any proposed material has actually been synthesized or tested in a physical lab.

Key facts

  • Agents get web search (Exa), a Python/Bash sandbox with pymatgen, mp_api and ASE, and ML surrogates (built on the PET-MAD potential, Pheasy/Phonopy, and a GMTNet dielectric model) to screen candidate materials, and run with no stopping condition until they error or hit a 100 million token budget.
  • Synthesis recipes are graded by a worst-of-three GPT-5.6 Sol judge with web search, chosen for the best correlation with human expert grading, against a 17-criterion penalty rubric where a single critical penalty alone is enough to reject a recipe.
  • The worked example is a Claude Opus 5 recipe for growing hexagonal diamond (lonsdaleite) by seeded microwave-plasma CVD on a 300 mm wafer; the grader rejected it with one critical penalty and five fixable ones.
  • The critical flaw: the nanodiamond seeds used are themselves cubic diamond, so growth would bypass the nitride buffer meant to template hexagonal stacking, unlike a cited real route to phase-pure hexagonal diamond via oriented graphite at 20 GPa and 1,300 to 1,900 degrees Celsius.
  • The benchmark runs on the AI Security Institute's open-source Inspect framework; the published text gives no aggregate pass rate, candidate count or cross-model comparison, and no confirmation that any proposed material was actually synthesized.

Why it matters

The benchmark targets a harder bar than plausible-sounding chemistry: a candidate material only counts if it is genuinely novel by the literature (never deposited as a BEOL-compatible thin film before) and comes with a synthesis recipe that expert-level review judges worth attempting. BEOL, back-end-of-line, refers to the interconnect layers built on top of finished transistors, so materials also have to survive that processing window. Pairing a discovery agent with a separate, rubric-driven grader for the synthesis step is the core idea: proposing a material with the right thermal, dielectric and mechanical properties is treated as necessary but not sufficient, since it also has to be makeable.

Who it affects

Materials scientists and semiconductor R&D teams evaluating whether AI agents can contribute to materials discovery for chipmaking are the direct audience. The worked example also puts two AI labs in a head-to-head role by construction: Claude Opus 5 is the proposer of the rejected recipe, and GPT-5.6 Sol is the grader, itself calibrated against human expert judgments rather than treated as ground truth on its own.

How to use it

This is a published methodology and benchmark harness, not a paid product or a released dataset in the text itself. The components described are largely built on named open tools: pymatgen, mp_api and ASE for materials handling, the PET-MAD interatomic potential, the Pheasy and Phonopy phonon codes, a GMTNet-based dielectric model trained on the JARVIS DFPT database, and the AI Security Institute's Inspect framework for running the benchmark. No pricing, access terms or release details are given in the text.

How solid is it

The individual components are citable, published methods: PET-MAD appeared in Nature Communications, Phonopy is an established phonon code while Pheasy is newer work described in a 2025 arXiv preprint, and Inspect is the AI Security Institute's open-source evaluation framework. The synthesis rubric itself was derived empirically, by first having human experts grade LLM-written recipes and then encoding that grading as penalties. What the text does not supply is any aggregate result: no pass rate, no count of how many candidate materials were proposed across models, and only this one worked example, which is a rejection rather than a success case.

Risks and caveats

Grading recipes with an LLM judge, even a worst-of-three vote checked against human agreement, is a proxy for whether a recipe would work, not a physical test of it. The text does not say whether any proposed material has actually been synthesized or validated in a lab. The worked example itself is a caution: the rejected recipe reads as detailed and specific, complete with wafer size, gas flows, bias steps and characterization plan, yet the grader still failed it over a single mechanistic gap, that the seed crystallites used are cubic diamond and cannot template the hexagonal stacking the recipe claims to produce, and that the proposed characterization could have missed the mismatch.

“The CH4/H2 plasma and dense nanodiamond seeding could plausibly produce a continuous nanocrystalline diamond film. They do not, however, provide a credible pathway to the specified ordered P6(3)/mmc phase: growth will originate on predominantly cubic nanodiamond seeds, effectively isolating it from the proposed hexagonal buffer. I would not attempt this as a lonsdaleite recipe, although it could be reworked into a cubic-NCD experiment.”

— the grader (GPT-5.6 Sol), on Claude Opus 5's hexagonal-diamond recipe