GPT-5.6 Sol runs routine qubit calibration measurements at MIT

Beatriz Yankelevich, a graduate student in MIT's Engineering Quantum Systems Group (EQuS), connected OpenAI's GPT-5.6 Sol, harnessed to Codex, to the software that controls her quantum computing experiments. She wanted to know whether an AI agent could take over the routine calibration measurements that quantum experiments require, work that can take months and hundreds to thousands of preliminary measurements before it is done.
EQuS studies superconducting qubits, which are cooled to near absolute zero inside dilution refrigerators and controlled with microwave pulses. Once a chip is fabricated, packaged and cooled, researchers interact with it entirely through software, which is what made it a workable target for an AI agent. Yankelevich gave GPT-5.6 Sol, through Codex, a set of measurement-specific skills describing how to run and evaluate each experiment, along with the chip's design targets. The agent then chose measurement parameters, operated the hardware, analyzed the resulting data and decided whether to refine a measurement or save the result and move to the next step.
The test chip was an uncalibrated six-qubit device, one of a standard type EQuS fabricates to benchmark its manufacturing process; characterizing one of these chips normally takes a researcher several days. When the measurement signals were clear, GPT-5.6 Sol completed a standard sequence with little intervention: it identified the qubits' transition frequencies, calibrated the pulses used to control and read them, and determined how long each qubit retained quantum information. When signals were weak or noisy, the agent struggled more, taking longer to find workable parameters and sometimes needing guidance from Yankelevich.
Yankelevich said she can now leave agents running measurements for many hours overnight or while she works in the cleanroom, checking progress from her phone and stepping in only when something needs fixing or she wants to try a different direction. EQuS now uses agents regularly to handle routine chip characterization. For less standard experiments, she assigns Codex agents narrower goals and relies more on their ability to write, modify and test code for control, analysis and simulation, running several agents on different problems at once while she spends most of her own time interpreting results, designing experiments and planning the agents' next steps.
The write-up also notes limits: experienced researchers may still be able to find the best calibration settings faster than current AI models, and interpreting ambiguous or noisy physical results remains a challenge for the agents. The advantage described is time freed up from constant monitoring, not a claim that the agent outperforms human judgment.
Key facts
- Beatriz Yankelevich, an MIT EQuS graduate student, connected GPT-5.6 Sol, harnessed to Codex, to her lab's control software to run superconducting-qubit calibration measurements.
- On an uncalibrated six-qubit chip, the agent completed a standard calibration sequence largely unsupervised: finding transition frequencies, calibrating control and readout pulses, and measuring how long qubits retained quantum information.
- Characterizing a standard EQuS chip normally takes a researcher several days; the agent's autonomy let Yankelevich run measurements overnight or during cleanroom work while checking in from her phone.
- Performance dropped when signals were weak or noisy: the agent took longer to find suitable parameters and sometimes needed guidance from Yankelevich.
- OpenAI notes experienced researchers may still find the best calibration settings faster than current AI models; the time saved goes toward experiment design and analysis, not toward outperforming human judgment.
Why it matters
Calibrating a superconducting qubit chip is unglamorous prerequisite work: hundreds to thousands of preliminary measurements, stretched over months, before any real experiment can begin. That workload is also entirely software-mediated once a chip is cooled, which made it a natural fit for an AI agent rather than a stretch application. Handing routine calibration to GPT-5.6 Sol did not just save time on the measurements themselves; it shifted where Yankelevich spends her attention, from monitoring every step to designing experiments and analyzing results.
Who it affects
Directly, Beatriz Yankelevich and the MIT Engineering Quantum Systems Group, which now uses agents regularly for routine chip characterization. More broadly, the case points at other labs running superconducting-qubit or similar hardware experiments that involve the same cycle of software-controlled measurement, analysis and adjustment, and at researchers weighing whether Codex-style agents can take on well-defined experimental physics tasks beyond software engineering.
How to use it
Yankelevich connected GPT-5.6 Sol, through Codex, directly to the lab software that coordinates her experiments, then supplied measurement-specific skills explaining how to run and evaluate each experiment along with the chip's design targets. From there the agent chose measurement parameters, operated the hardware, analyzed the data and decided whether to refine a measurement or move on. For non-routine experiments, she narrows the agent's goals and leans on its ability to write and modify control, analysis and simulation code. The source gives no pricing, licensing or release details for GPT-5.6 Sol or Codex.
How solid is it
This is a single case study published by OpenAI itself, describing one researcher's experience with one lab's chip; it is not an independent or peer-reviewed evaluation. No accuracy figures, success rates or head-to-head timing comparisons against human researchers are given, only the qualitative account that clear-signal measurements went smoothly and noisy ones did not. No other EQuS researcher or lab lead is named or quoted.
Risks and caveats
The agent's success was limited to clearly defined, well-understood measurement workflows; when signals were weak or noisy it took longer and sometimes needed an experienced researcher to step in. The source itself cautions that experienced researchers may still be able to identify the best calibration settings faster than current AI models can. No time-saved figure, in hours or as a percentage, is given beyond Yankelevich's general statement that agents can run for many hours unattended.
“I can have agents running measurements for many hours overnight or while I'm working in the cleanroom. I can check in from my phone, see what they've done, and steer them if something needs fixing or if I want to explore a different direction.”
— Beatriz Yankelevich, graduate student, MIT Engineering Quantum Systems Group