SpaceXAI Releases Grok 4.7 for Long Bench Jobs

SpaceXAI's published graphic for the Grok 4.7 announcement.

The regulator on this board sags when the radio keys up, the connector pinout in the email does not match the silkscreen, and the firmware hangs before the serial console. A regulator holds a supply voltage steady. That is an electrical question and a long coding job. On September 21, 2026, SpaceXAI released Grok 4.7 and called it the company's most capable model for coding and knowledge work. It is available the same day in Cursor, Grok Build, the Grok API (application programming interface, a program's door to the model), and, in SpaceXAI's words, other coding harnesses, routers, and cloud platforms.

The starting price matches Grok 4.6: $2 per million input tokens and $6 per million output tokens. A token is a small piece of text, often a word or part of a word. A fast variant costs twice as much and, SpaceXAI says, writes output twice as fast. I would use the base rate on the firmware trace. I would not use the fast rate to decode a resistor.

What actually changed?

SpaceXAI says Grok 4.7 uses a larger base model than Grok 4.6. The post states no parameter count, so I will not import one from secondary sites. The training note that matters on a bench is a longer reinforcement-learning run, weighted toward tasks that take many hours. Reinforcement learning means the model practices on tasks and is scored on the result, not only trained to predict the next word. SpaceXAI also says the model checks its own work more carefully and handles longer context. Context is the text it can consider at once. No window length is printed.

The headline says "twice as fast, at half the price of comparable models." The body says Grok 4.7 is served at the same price and speed as Grok 4.6, and that the fast variant is the doubled-speed, doubled-price option. The table is the part I can read without choosing a slogan. Input is $2 per million for Grok 4.7 and Grok 4.6, $4 for GPT-5.6 Sol, and $10 for Fable 5.1. Output is $6, $6, $20, and $50 in that order. SpaceXAI also says it trained the model to understand the Grok Bot harness for conversation as well as code.

How does the new piece work?

The failure I know on a long coding job is a confident patch that never looks at the log. SpaceXAI's bet is that more of the training score came from multi-hour tasks, plus stronger self-checks. On a firmware trace that means: find the power-good pin, see who samples it, and do not declare the board up because a comment says so. That is still a description of training. I did not watch a trace on September 21.

The published table labels the columns Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max. These are SpaceXAI's numbers. CursorBench 4.0, longer coding tasks, reads 46.3%, 40.4%, 41.7%, and 51.8%. EEBench, the electrical-engineering set, reads 64.0%, 53.0%, 39.4%, and 56.4%. Grok leads that row. Terminal-Bench 4.0 reads 37.6%, 20.3%, 37.3%, and 57.9%. DeepSWE v1.1 is 71.0% for Grok 4.7 with an asterisk SpaceXAI marks as high effort, then 65.2%, 72.7%, and 70.0%. AA Briefcase v1.1 is 1,657, 1,546, 1,487, and 1,678. A second chart caption names GPT-6 Astra beside Grok 4.6 and Fable 5.1 for GDPval and related scores. The prose does not print a GDPval number, so I will not invent one.

The same Grok 4.7 title card, uncropped.
The announcement card again, without the wide crop. The score table on the post is HTML, not a separate image file.

What does this look like on a real project?

I have not run Grok 4.7 for this article. The first half is the electrical question, in my own words, plus the datasheet pages I actually have. The 5-volt regulator droops when the radio transmits. I want a checklist: input capacitor, dropout at the current I measured, the ground path, and whether the pin I think is ground is the pin on the drawing. A 64% EEBench score is a reason to let the model draft the list. If the model and the meter disagree, the meter wins.

The second half is Grok Build, which the post offers as a place to try the model. I would ask for a firmware change that blocks boot until power-good has been stable for a few milliseconds, with a serial line if the wait times out. The model can read several files and propose a diff, the line-by-line edit. I compile it. I watch the rail with a scope while the radio keys. The fast variant is for a session where I am stuck waiting. An overnight read of the repo can stay on $2 and $6.

How does it compare with the previous version?

Against Grok 4.6, the same starting price is tied to a larger base model and higher company scores on the rows above. CursorBench moves from 40.4% to 46.3%. EEBench moves from 53.0% to 64.0%. Terminal-Bench 4.0 moves from 20.3% to 37.6%, a large jump that still leaves the score under 40%. Speed is the muddy part. The body says same price and same speed as Grok 4.6, with "faster" as a separate option at twice the money. The headline talks about comparable models, and the table does show a lower price than GPT-5.6 Sol and Fable 5.1. I would budget the base tier unless a session is painfully slow.

Where does it sit next to other tools a maker already uses?

Cursor is already where a lot of firmware edits happen, and SpaceXAI says Grok 4.7 is there today. Grok Build is the company's own bench. The API is there if a jig script needs to call it. On EEBench, SpaceXAI's numbers put Grok 4.7 above GPT-5.6 Sol and Fable 5.1. On CursorBench and Terminal-Bench 4.0, Fable 5.1 is ahead in that same table. HealthBench Professional, which I did not dwell on above, is 56.7% for Grok 4.7, behind 60.5% and 62.1% for the other two named models, and ahead of 48.5% for Grok 4.6. Pick the row that matches the afternoon. I am not relabeling the GPT-5.6 Sol column. A meter and the vendor datasheet can still contradict the model.

What does it cost, and who can use it today?

On September 21, 2026 the model is in Cursor, Grok Build, and the API, plus the other platforms named only as a class. Starting price is $2 per million input tokens and $6 per million output tokens. The fast variant is twice the output speed at twice the price. Applied to those rates, that is $4 and $12 per million. SpaceXAI states the multiple. The doubled figures are arithmetic, labeled as such.

Safeguards are part of the same post. SpaceXAI says a new stack is its strongest yet on refusals. On LatchBio's biosafety benchmark it reports 62.4%. On HackerBench v0.3 it says 3.3% of risky cyber prompts get through, while ordinary security work is rarely blocked. A few partners have invite-only access for defense research. None of that is a procedure. A dual-use question can be refused, and a refusal is the outcome I want on that kind of prompt.

What is still unproven?

I did not run the model on publication day. CursorBench at 46.3% still misses more than half of that set. The DeepSWE asterisk means 71.0% is a high-effort score. Context length is unpublished here. A headline about speed and a body about the same speed as Grok 4.6 need a stopwatch. EEBench will not know the layout error on my connector. Ask for the checklist, measure the rail, and let a Grok Build pass propose the power-good wait. If the patch compiles and the scope agrees, the $2 and $6 rate was earned.

Disclosure: The author is a paying subscriber to ChatGPT Plus, Claude Pro, and SuperGrok and uses all three services on a daily basis. The Makers Workbench is not affiliated with OpenAI, Anthropic, xAI, Google, or any of the other major AI companies covered in our reporting. No company receives favorable editorial treatment based on the author's personal subscriptions.

Sources and image credits

Sub-Category

Add new comment

Restricted HTML

  • You can align images (data-align="center"), but also videos, blockquotes, and so on.
  • You can caption images (data-caption="Text"), but also videos, blockquotes, and so on.