AI NewsModels & agentsAnnouncement
SpaceXAI released Grok 4.7 on Sunday and says it scores 38 percent on Terminal-Bench 4.0, up from 20 percent for Grok 4.6, after training weighted toward multi-hour tasks
SpaceXAI released Grok 4.7 on 21 September, priced at $2 per million input tokens and $6 per million output tokens, and says the model scores 38.0 percent on Terminal-Bench 4.0 against Grok 4.6's 20.3 percent, after a longer reinforcement learning run weighted toward tasks that take many hours to finish.
Why it mattersA model trained on multi-hour tasks changes what a team can hand to an unattended agent, and if the score gap holds up in reproduction, the cost of testing longer runs falls at the same time the price per token comes down.
SpaceXAI released Grok 4.7 on 21 September 2026 and positioned it as a coding and knowledge-work model built for long-running tasks. Its own announcement calls the release "twice as fast, at half the price of comparable models", and prices the API at $2 per million input tokens and $6 per million output tokens. A faster variant serves the same model at double the price with roughly double the output speed.
What the training changed
SpaceXAI says Grok 4.7 is the same base model as Grok 4.6, retrained through a longer reinforcement learning run that deliberately weighted harder problems, including tasks the company describes as taking "many hours" to complete. The company credits the gain to two capabilities: self-verification, so the agent catches its own wrong turns before building on them, and long-context management, so a growing interaction history stays usable.
SpaceXAI has not explained how either capability was measured, or whether context handling comes from architecture, summarisation, retrieval or something else. Amanda Caswell at The New Stack noted on 21 September that "the company did not disclose whether the context gains came from architectural changes, summarization, retrieval, or better retention across long sequences".
The scores it is claiming
On the vendor's own numbers, Grok 4.7 scores 38.0 percent on Terminal-Bench 4.0, up from 20.3 percent for Grok 4.6. On CursorBench 4.0, an in-editor multi-step coding benchmark, it moves from 40.4 to 46.3 percent. On AA Briefcase v1.1, a multi-hour professional work test, it goes from 1,546 to 1,657. SpaceXAI also reports 71.0 percent on DeepSWE v1.1, 64.0 percent on EEBench and 19.6 percent on the Harvey Legal Agent Benchmark.
For a reference point outside SpaceXAI's own numbers, The New Stack cites the independent Terminal-Bench leaderboard, on which Anthropic's Claude Fable 5.1 scores 57.9 percent. Grok 4.7 still trails Fable 5.1 by nearly 20 points on that test.
The harness is now part of the model
SpaceXAI trained Grok 4.7 to natively understand the Grok Bot harness, its terminal coding agent that ships context management, tool schemas and execution results back into the model. Training on the harness rather than only through prompting at runtime is meant to reduce overhead in tool use and multi-step execution.
That direction mirrors OpenAI, which opened its Codex harness as the Agents API last week to move the surrounding infrastructure into a managed service. When a model is trained around a specific tool format and execution environment, swapping models without sacrificing agent performance becomes harder, and the decision about which harness to adopt starts to lock in a set of models along with it.
The score jump on Terminal-Bench looks larger than the release paper explains
A jump from 20.3 to 38.0 percent on the same benchmark, from the same base model, is close to a doubling, and SpaceXAI has attributed it to two vague-sounding levers with no numbers behind either. A separate recent benchmark of private codebases had the best model failing more than 60 percent of the time, so the reliability floor for unattended agents is still low. Read the deltas as vendor claims that need independent reproduction.
For a team weighing an unattended overnight run, the practical question is whether Grok 4.7 clears the threshold at which the run is worth kicking off. At $6 per million output tokens with self-verification claims that can be tested against a public benchmark, that is a decision worth an afternoon of setup.
Source
SpaceXAI, Grok 4.7, 21 September 2026. Reporting and independent Terminal-Bench context from Amanda Caswell, The New Stack, Grok 4.7 was built to work for hours. It still fails most of the time., 21 September 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.



