A firmware repo had failed the same board check for two days, and the log was long enough that I kept losing the failing line. That is the job OpenAI is pointing GPT-5.5 at. They introduced it on April 23, 2026. This note is the next day, so it includes the April 24 update: GPT-5.5 and GPT-5.5 Pro are in the API, and the system card was updated to describe additional safeguards.
GPT-5.5 is rolling out to Plus, Pro, Business, and Enterprise in ChatGPT and Codex. GPT-5.5 Pro is rolling out to Pro, Business, and Enterprise in ChatGPT. A token is a piece of text the bill counts, often a word or part of one. The context window is how much material fits in one request. Posted API rates, written before the switch and in force after the April 24 note, are $5 per million input tokens and $30 per million output tokens for gpt-5.5, with a 1 million token window. gpt-5.5-pro is $30 and $180 per million. The April 23 text said the API was coming very soon. The April 24 update says it is open. I did not run the model for this piece.
What actually changed?
GPT-5.4 is the model these numbers sit next to. OpenAI says GPT-5.5 sees what you are trying to do sooner and can carry more of a messy task: code, debugging, research, documents, spreadsheets, and the software on the machine. An agent, in their sense, plans, uses tools, checks its work, and keeps going across that chore.
Per-token latency, the wait for each token, matches GPT-5.4 in their serving tests, and the same Codex tasks use fewer tokens. Codex is their coding product. The price per token is higher than GPT-5.4. The older dollar rate is not printed here.
Safeguards ship with the model. OpenAI cites a preparedness review, red-team testing, extra cyber and biology checks, and nearly 200 early partners. The April 24 line says the system card now describes additional safeguards. The update sentence does not list them.
How does the new piece work?
Reasoning effort is how hard the model thinks. OpenAI ran the evals below at xhigh, in a research setup they say can differ slightly from production ChatGPT. Read the percentages as theirs.

They call GPT-5.5 their strongest agentic coding model so far. Terminal-Bench 2.0, hard command-line work, is 82.7% for GPT-5.5 and 75.1% for GPT-5.4. SWE-Bench Pro, real GitHub issues, is 58.6%. On that same OpenAI table, GPT-5.4 is 57.7%, Claude Opus 4.7 is 64.3%, and Gemini 3.1 Pro is 54.2%. Those columns are OpenAI's, not a bake-off I ran. A footnote says labs have noted memorization on this eval. Expert-SWE, an internal test whose tasks have a median estimated human time of 20 hours, is 73.1% versus 68.5% for GPT-5.4. OpenAI says all three coding tests rose while token use fell. Dan Shipper, of Every, is quoted calling it "the first coding model I've used that has serious conceptual clarity."
Computer use means the model runs software: screen, click, type, next program. OSWorld-Verified is 78.7% for GPT-5.5, 75.0% for GPT-5.4, and 78.0% for Claude Opus 4.7 on OpenAI's table. The Gemini cell there is empty, and I am leaving it empty. GDPval, wins or ties on knowledge work, is 84.9% versus 83.0% for GPT-5.4, with Opus 4.7 at 80.3% and Gemini 3.1 Pro at 67.3%, again OpenAI's table. CyberGym is 81.8%, against 79.0% for GPT-5.4 and 73.1% for Opus 4.7.
OpenAI is treating biological and chemical capability, and cybersecurity, as High in its Preparedness Framework, not Critical. They say cyber skill is still a step up from GPT-5.4, with stricter classifiers than GPT-5.4 on risky requests and on repeated misuse. Some users, they warn, will find the classifiers annoying while they are tuned. Trusted Access for Cyber is the verified path, starting in Codex. I am not describing an exploit or a biological method.
What does this look like on a real project?
I would point it at that failing firmware check: a small repo, the serial log, a datasheet PDF already in the tree, and a test command I trust. One task. Find the failure, change the code, run the command, and name what else the edit touches. Then I read the diff, the changed lines, on a board I can probe.

In Codex, GPT-5.5 is on Plus, Pro, Business, Enterprise, Edu, and Go, with a 400K context window, meaning 400,000 tokens. That is sized for a small firmware tree plus a few PDFs. I would still watch the counter. The API window for gpt-5.5 is 1 million tokens. Fast mode makes tokens 1.5 times faster at 2.5 times the cost. I would use it on a compile-and-log loop, and leave it off while I read a plan.
I would ask for the failing lines before any edit. Pietro Schirano, via the post, saw a large merge finish in about 20 minutes. That is his report. I still want the test output, and the board still has to pass. A refused question about a library flaw may be the tighter cyber classifier. Verified defensive work is what Trusted Access for Cyber is for. I would not try to route around a refusal. A file path is easier to audit than a click in a programming tool.
How does it compare with the previous version?
The wide gap versus GPT-5.4 on OpenAI's table is Terminal-Bench 2.0, 75.1% to 82.7%. SWE-Bench Pro barely moves, 57.7% to 58.6%, with the memorization note, and with Opus 4.7 higher at 64.3% in OpenAI's column. OSWorld-Verified goes from 75.0% to 78.7%, a hair over the 78.0% they list for Opus 4.7. GDPval goes from 83.0% to 84.9%. CyberGym goes from 79.0% to 81.8%. Internal Expert-SWE goes from 68.5% to 73.1%.
Latency per token matches GPT-5.4, OpenAI says, with fewer tokens on Codex tasks and a higher sticker price. In Codex they say most subscribers should see better results at lower token use than GPT-5.4. GPT-5.5 Pro is the higher-accuracy sibling, priced at $30 and $180, and early testers told OpenAI the answers were more complete than GPT-5.4 Pro. Cyber checks are stricter than on GPT-5.4. The label they publish is High, and they say Critical cybersecurity was not reached.
Where does it sit next to other tools a maker already uses?
The editor, the serial terminal, the programmer, and the PDF viewer stay. Codex is where this model meets a repo. ChatGPT's GPT-5.5 Thinking variant covers Plus through Enterprise. A script uses gpt-5.5 or gpt-5.5-pro. Claude Opus 4.7 and Gemini 3.1 Pro appear only as OpenAI's table columns: Opus ahead on SWE-Bench Pro, GPT-5.5 ahead on Terminal-Bench 2.0. If the log and a pin disagree, I believe the pin.
What does it cost, and who can use it today?
ChatGPT: GPT-5.5 for Plus, Pro, Business, and Enterprise. GPT-5.5 Pro for Pro, Business, and Enterprise. Codex adds Edu and Go, at a 400K context window, with Fast mode at 1.5 times speed and 2.5 times cost.
API, as of the April 24 update: gpt-5.5 at $5 and $30 per million tokens, 1 million token context. gpt-5.5-pro at $30 and $180. Batch and Flex are half the standard rate. Priority is 2.5 times the standard rate. A long computer-use session can spend tokens even when the note at the end is short. Whether fewer tokens offset the higher rate is a measurement, not a slogan in the post.
What is still unproven?
I have not run GPT-5.5 or Pro on a board. The scores are OpenAI's, at xhigh, in a setup that may differ from the chat product. SWE-Bench Pro is a soft comparison once the memorization footnote and the Opus column are in view. The April 24 safeguard sentence does not enumerate the new controls. Biology and chemistry are classified High. I am not unpacking that suite. The bench question is narrower: does the patch compile, pass the command I named, and match the pin I can probe?
Disclosure
Disclosure: The author is a paying subscriber to ChatGPT Plus, Claude Pro, and SuperGrok and uses all three services on a daily basis. The Makers Workbench is not affiliated with OpenAI, Anthropic, xAI, Google, or any of the other major AI companies covered in our reporting. No company receives favorable editorial treatment based on the author's personal subscriptions.
Sources and image credits
- OpenAI announcement (openai.com)
- Images: published by OpenAI with the announcement.
