SpaceXAI Launches Grok 4.5 as the Default in Grok Build

SpaceXAI's Grok 4.5 title card

The median bug in my temperature-logger repo is the dull kind. The function sorts the analog buffer, then picks an index that can fall between two samples, so the bench display jitters while the sensor sits still. On July 16, 2026, SpaceXAI launched Grok 4.5 for that sort of job and made it the default model in Grok Build.

SpaceXAI prices it at $2 per million input tokens and $6 per million output tokens. A token is a chunk of text the model counts, usually a piece of a word. The company says the service writes about 80 tokens per second, and that the model has about twice the token efficiency of comparable leading models on the same tasks. Tokens per second is how many of those chunks come back each second.

You can use it on July 16 in Grok Build, in Cursor on all plans, and from the SpaceXAI console. SpaceXAI is also offering Grok 4.5 free for a limited time in Grok Build and in Cursor. The post does not say when that window ends. For a maker, the reason to care is a repo you already have open.

What actually changed?

SpaceXAI calls Grok 4.5 its smartest model, built for coding, agentic tasks, and knowledge work, and its strongest model so far. An agentic task is a job with more than one step: the model reads files, calls tools, and checks what came back before it stops. The company says it was trained alongside Cursor and on data spanning coding, science, engineering, and math.

SpaceXAI says Grok 4.5 solves tasks in under half the number of steps of comparable leading models. On one chart of average output tokens per SWE-Bench Pro task, a software-repair benchmark, it reports 15,954 tokens for Grok 4.5 and 67,020 for Opus 4.8 at max, about 4.2 times fewer. "About twice" is the general claim. The 4.2 times figure is one chart against one model.

Inside Grok Build, this model is the default. The post says it can build Excel workbooks with web research, multi-sheet formulas, and notes, use PowerPoint's own shapes for diagrams, and write the prose in Word.

How does the new piece work?

SpaceXAI says the data work was deduplication, quality scoring, and domain-focused selection. Training ran across tens of thousands of NVIDIA GB300 graphics processing units, the parallel chips these runs use. Reinforcement learning, which the post does describe, is the later stage where the model attempts a task, gets a grade, and is pushed toward attempts that graded better. SpaceXAI says that stage covers hundreds of thousands of tasks, centered on multi-step software engineering, with automated grading and model-based grading. The stack is asynchronous, so a long agent rollout can keep going for many hours while learning continues. The stated focus was per-token intelligence: more useful work from each token.

SpaceXAI's Grok 4.5 title card, uncropped
The same official title card, without the wide crop. The words "Grok 4.5" sit in the middle of a dark gradient. The comparison charts on the post are HTML, so they are not this file.

What does this look like on a real project?

I would use Grok 4.5 inside Grok Build on the logger repo. I did not run that test for this article. The post's own API sample asks the model named grok-4.5 to fix a median function that sorts an array and indexes it at half the length. In JavaScript that division can land between elements.

Open Grok Build during the limited free period that also covers Cursor. The post prints an install curl pointed at x.ai/cli/install.sh. On the API, the sample posts to the responses endpoint with a bearer token and the model name grok-4.5. I would ask for three things: find the median bug, add tests for an odd-length buffer and an even-length buffer, and explain the fix in a comment. Then I run the tests myself. The new test should fail on the old function. If the model only rewrites the comment, the test catches it.

SpaceXAI published one announcement picture for this launch. The score tables sit in the page as HTML.

How does it compare with the previous version?

The post calls this SpaceXAI's strongest model so far, then compares it with other labs. The chart note says competitor figures come from those developers' system cards or from benchmark leaderboards. There is no older-Grok column on the page.

DeepSWE 1.0, created by Datacurve and run by AA with each provider's harness, reads Fable at max 66.1 percent, GPT 5.5 at xhigh 64.31 percent, Grok 4.5 at 62.0 percent, Opus 4.8 at max 55.75 percent, and Opus 4.7 at max 40.12 percent. A harness is the wrapper of tools and instructions around the model. AA is SpaceXAI's label for the group that ran that set. DeepSWE 1.1, a mini-swe-agent harness run by Datacurve, reads Fable at max 70 percent, GPT 5.5 at xhigh 67 percent, Opus 4.8 at max 59 percent, Grok 4.5 at 53 percent, and GLM 5.2 at 44 percent.

SWE Marathon resolution rate at pass@1, one attempt and no retries, is the row SpaceXAI leads here: Grok 4.5 at 29.0 percent, Opus 4.8 at max 26.0 percent, Fable at max 24.0 percent, Opus 4.7 at max 16.0 percent. Terminal Bench 2.1 reads Fable at max 84.3 percent, GPT 5.5 at xhigh 83.4 percent, Grok 4.5 at 83.3 percent, and Opus 4.8 and Opus 4.7 at max both 78.9 percent. SWE-Bench Pro resolve rate reads Fable at max 80.4 percent, Opus 4.8 at max 69.2 percent, Grok 4.5 at 64.7 percent, Opus 4.7 at max 64.3 percent, GLM 5.2 at 62.1 percent, and GPT 5.5 at xhigh 58.6 percent. Read together, Grok 4.5 leads SWE Marathon and trails Fable on DeepSWE and SWE-Bench Pro.

Where does it sit next to other tools a maker already uses?

If Cursor is already the editor on this bench, Grok 4.5 is on all plans there. I would try the logger bug in that editor during the free window and keep my own test command. Grok Build is the door where this model is the default. The console is the door for a script. SpaceXAI also links marketplace plugins for Word, PowerPoint, Excel, and Outlook, for the days when the deliverable is a build report.

What does it cost, and who can use it today?

On the July 16 post, input is $2 per million tokens and output is $6 per million tokens. The post states no parameter count. It also prints no context-window size and no cached-input rate. A context window is how much counted input one request can hold. Cached input is a repeated prefix a service can bill at a lower rate when a cache hits. This launch post does not give either number.

As the post writes availability on July 16, the list is Grok Build, Cursor on all plans, and the SpaceXAI console with the model name grok-4.5. The free offer is limited to Grok Build and Cursor, for a limited time, with no end date printed.

What is still unproven?

I have not run Grok 4.5 on the logger repo. The two DeepSWE rows do not agree about rank, and a harness change can move a model by more than the gap between neighbors. The 4.2 times token figure is one average against Opus 4.8 at max on SWE-Bench Pro. "Under half the steps" is SpaceXAI's summary, not a number I recomputed. Eighty tokens per second will not feel smooth on a run that stops to read a file or wait on a test. The free window has no published end. The July 16 post states no parameter count, no context-window size, and no cached-input rate.

Disclosure: The author is a paying subscriber to ChatGPT Plus, Claude Pro, and SuperGrok and uses all three services on a daily basis. The Makers Workbench is not affiliated with OpenAI, Anthropic, xAI, Google, or any of the other major AI companies covered in our reporting. No company receives favorable editorial treatment based on the author's personal subscriptions.

Sources and image credits

Sub-Category

Add new comment

Restricted HTML

  • You can align images (data-align="center"), but also videos, blockquotes, and so on.
  • You can caption images (data-caption="Text"), but also videos, blockquotes, and so on.