OutYet reporting
An MIT quantum-lab case study shows where GPT-5.6 Sol agents still need supervision
OpenAI describes a bounded laboratory workflow in which GPT-5.6 Sol operated measurement software through Codex. The useful result is not unsupervised science, but a clearer picture of which repetitive experimental loops can be delegated and where expert judgment remains necessary.
OpenAI's September 8 case study says that MIT graduate researcher Beatriz Yankelevich connected GPT-5.6 Sol, through Codex, to software controlling an uncalibrated six-qubit chip. Given measurement-specific skills and chip design targets, the system selected parameters, operated the hardware, analyzed the resulting data, and either refined a measurement or saved a result for the next step. The account documents a constrained experimental workflow; it does not present a general benchmark of autonomous scientific research.
The setting is consequential because qubit calibration is an interdependent loop: a result from one measurement determines the next, while physical drift and inconsistent signals can change the path. OpenAI says the agent completed standard sequences with little intervention when signals were clear, including identifying transition frequencies and calibrating control and readout pulses. The lab's standard chip characterization can take several days, so automating routine steps could reduce monitoring work without removing the scientist from experiment design.
The comparison with the prior GPT-5 tier is clearest in the API documentation, not in a head-to-head laboratory result. OpenAI describes GPT-5.6 Sol as the flagship GPT-5.6 tier and the target of the gpt-5.6 alias, with configurable reasoning effort. Its documentation lists a 1,050,000-token context window, 128,000 maximum output tokens, and input pricing of $4 per million tokens versus $5 for GPT-5.5. The same documentation does not establish that Sol outperforms GPT-5.5 on quantum calibration.
The practical boundary is as important as the successful run. OpenAI reports that weak or noisy signals made the agent slower to find suitable parameters and sometimes required guidance from an experienced researcher; it also says experienced researchers may still identify the best settings faster. For technical teams, that supports using an agent for repeatable, observable loops with narrow skills, checkpoints, and a human able to intervene. It does not support treating the case study as proof that an agent can independently interpret ambiguous physical results.
Related models
Sources
- How GPT-5.6 Sol helps run quantum computing experiments · OpenAI
- GPT-5.6 Sol Model · OpenAI