Grok 4.6 Takes On the Long Coding Job You Usually Have to Watch

SpaceXAI's Grok 4.6 title card

A test-bench logger does not stay one Python file. It picks up a serial reader, a small page, a pinout note, and bugs you will not leave alone. Serial is the text stream from a microcontroller, the small computer on the device you are measuring. You approve each file, paste the error back, and remind the model of the last filename. That babysitting is what SpaceXAI says Grok 4.6 is for.

On August 12, 2026, SpaceXAI released Grok 4.6. That is the company name on the announcement. It builds on Grok 4.5, aimed at long-running agents and more ambitious interactive and visual work. An agent plans steps, calls tools, and keeps going toward a goal, instead of answering once and stopping. It is in Cursor and Grok Build, on the API, and at OpenRouter, Vercel, and Cloudflare the same day. The API, the application programming interface, is the door other software uses to call the model.

Pricing starts at $2 per million input tokens and $6 per million output tokens. A token is a chunk of text the bill counts, often a word or part of one. A fast variant is twice that, $4 and $12. For the first week, included usage in Grok Build and Cursor is doubled. The post never says what one share of "included" is.

What actually changed?

SpaceXAI says it stays on multi-step work: research, analysis, a codebase, or a first version of an app. The supplemental run was longer than Grok 4.5's. It used model-generated reasoning data, engineering data, and an improved optimizer, the procedure that nudges the model during training.

SFT, supervised fine-tuning, trains on examples whose steps are already written. SpaceXAI used Grok 4.5 to regenerate those examples across reasoning efforts, agent harnesses, and domains such as STEM (science, technology, engineering, and mathematics), software, and knowledge work. A harness is the wrapper that lets the model call tools. Bad traces were filtered by model checks. RL, reinforcement learning, scores attempts and pushes toward better scores, including coding, web work, and computer-aided design (CAD, software for drawing physical parts). There is no parameter count. I will not invent one.

On their trials, a broad idea could come back as a first working version, then take revisions. Longer runs started to show more self-testing. Visual first passes were stronger than they typically saw from Grok 4.5. I did not run one of those projects on August 12.

Safeguards, checks that block or reshape a risky request, were recalibrated. SpaceXAI says the stack should stay useful for vulnerability patching, a faster design cycle, and AI research, after their widest pre-deployment tests plus later third-party tests. This article does not describe a vulnerability or a patch.

How does the new piece work?

Context is what the model can see on a run: files, notes, tool results, and earlier steps. An agent keeps adding to that set. You do not configure the training in Cursor. You pick Grok 4.6, point it at the repo, and state the goal. The API, OpenRouter, Vercel, and Cloudflare are the other doors. Grok Build offers a free start. The doubled included usage lasts the first week, in Grok Build and Cursor. The post does not say it continues.

SpaceXAI's Grok 4.6 title card, uncropped
The same official title card, without the wide crop. The words "Grok 4.6" sit in the middle of a blue and white gradient. The score table on the post is HTML, so it is not this file.

What does this look like on a real project?

The logger is the repo I end up watching. Read voltage and temperature from a serial port. Show the last hour. Refuse to start if the port name is wrong. Keep a note of which connector wire is which, because a swapped sense line looks like a dead sensor.

The attended version is a loop of tracebacks, invented port names, and a pin note the model forgets after coffee. The work is ordinary. The attendance is the cost.

I would put the repo in Cursor or Grok Build and state the goal once, including both rules. Let it find a serial library, lay out the files, write the reader and the page, and check itself before calling the job done. SpaceXAI says longer runs started doing that check, and that a broad idea can return as a real first version. I would still read the diff. A self-check can check the wrong thing. The doubled first week is the time to learn what a long run costs. I have not run this logger on Grok 4.6. One revision I would ask for: large type, a stale-data state when the port goes quiet, and the raw number still visible. Then I would try the hardware. Fake serial lines are not a thermometer. SpaceXAI published one announcement picture for this launch. The score table sits in the page as HTML.

How does it compare with the previous version?

Columns below are Grok 4.6 High, Grok 4.5 High, GPT-5.6 Sol Max, and Fable 5 Max, from SpaceXAI's table. Competitor figures come from those developers' cards or leaderboards, they say, and third-party scores are the best of self-reported or public results. This is not one lab rerunning every model on August 12.

On the Artificial Analysis Intelligence Index, a composite of nine benchmarks, Grok 4.6 and GPT-5.6 Sol are both 61. Grok 4.5 is 56. Fable 5 Max is 62. The headline match with Sol is SpaceXAI's reading of that row. Fable 5 Max is one point higher on the same row.

Other rows, same order: GDPVal-AA v2 is 1753, 1526, 1728, 1741. CursorBench v3.2 is 69.9%, 66.7%, 67.2%, 70.5%. DeepSWE v1.1 is 65.9%, 54%, 73%, 70%. FrontierCode v1.1 Extended is 61.3%, 56.6%, 60.6%, 63.6%. APEX-Agents is 57.5%, 47.1%, 56.7%, 59.2%. Terminal-Bench v3.0 is 26%, 15.7%, 34.6%, 34.1%. APEX-SWE is 56.4%, 53.6%, no Sol score listed, and 58.8% for Fable 5 Max. AA-Briefcase is 1577, 1313, 1502, 1574. Harvey LAB (Vals) is 15.8%, 12.9%, 2.5%, 11.3%.

Grok 4.6 High leads Grok 4.5 High on every printed row. Against the other columns the chart is mixed. Several rows meet or pass Sol and still trail Fable 5 Max. DeepSWE and Terminal-Bench trail both by a wide gap. Harvey LAB leads, and the page does not explain that test. They also say visual first passes beat what they usually saw from 4.5. No parameter count, and no price comparison with 4.5, is in the post.

Where does it sit next to other tools a maker already uses?

Cursor already holds a lot of firmware repos. Grok 4.6 is selectable there on day one, with the first-week doubling, and the same offer is in Grok Build. API callers can use OpenRouter, Vercel, or Cloudflare. CAD shows up as training, and the post names no mechanical file format. The published comparison is the table above. I would judge the logger by whether I could leave the room.

What does it cost, and who can use it today?

Cursor and Grok Build can select Grok 4.6 on August 12. Included usage there is doubled for a week. The base amount is unstated. The API and the three partners are open the same day. Rates start at $2 and $6 per million tokens, input and output. The fast variant is $4 and $12. A long run spends output on files it writes and input on context it rereads, so it can cost more than a short chat. The doubled week is when that appetite is cheapest to measure. No seat price appears in the post.

What is still unproven?

The index tie is SpaceXAI's: 61 and 61, with Fable 5 Max at 62. Other competitor cells are the best public or self-reported numbers, by their own note. DeepSWE and Terminal-Bench on the same chart do not flatter 4.6. A composite tie can still lose the bench that looks like your job. "Started to see" self-testing is not a measured catch rate. There is no context-window size and no definition of the usage being doubled. I did not run the logger for this piece. If you try it, use a disposable repo, put the pin note in the prompt, and read the check.

Disclosure

Disclosure: The author is a paying subscriber to ChatGPT Plus, Claude Pro, and SuperGrok and uses all three services on a daily basis. The Makers Workbench is not affiliated with OpenAI, Anthropic, xAI, Google, or any of the other major AI companies covered in our reporting. No company receives favorable editorial treatment based on the author's personal subscriptions.

Sources and image credits

Sub-Category

Add new comment

Restricted HTML

  • You can align images (data-align="center"), but also videos, blockquotes, and so on.
  • You can caption images (data-caption="Text"), but also videos, blockquotes, and so on.