SpaceXAI released Grok 4.7 on Monday, calling it the company’s most capable model for coding and knowledge work and pricing it at $2 per million input tokens and $6 per million output tokens. The model is available today in Cursor, Grok Build, through the Grok API, and via third-party coding harnesses and model routers. A fast variant with twice the output speed costs twice the price.
The company claims a 71% score on DeepSWE v1.1, a software engineering benchmark, and says the model performs comparably to other frontier systems on GDPval and AA Briefcase, evaluations built around tasks done by professionals such as lawyers, nurses and financial analysts. SpaceXAI says Grok 4.7 is served at the same price and speed as its predecessor, Grok 4.6, and describes the release as “twice as fast, at half the price of comparable models,” a reference to rival frontier pricing.
What changed under the hood
Grok 4.7 uses a new, larger base model than Grok 4.6 and was trained with a longer reinforcement learning run on a harder task mix, weighted toward problems that take many hours to complete. The company says the model is better at verifying its own work and managing longer context, two failure modes that matter most in long coding sessions where an agent has to hold a plan across dozens of tool calls. Verification, in particular, has been the gap between demo and production for coding agents: a model that cannot check its own output produces code that compiles and still does the wrong thing. Several labs have made self-verification the headline improvement of their last two releases, and SpaceXAI is following the same path.
The release follows the industry’s shift toward long-horizon agentic work rather than single-shot question answering. SpaceXAI positions the model around sustained task execution: working through a refactor, debugging across files, producing documents and presentations. The GDPval and AA Briefcase results are the marketing surface of that positioning, since both benchmarks score models on real professional deliverables rather than exam-style questions. GDPval in particular has become the reference point for whether a model can produce work product a paying client would accept, and it has separated frontier models from mid-tier ones more sharply than most chat benchmarks.
Red-team access, invite only
SpaceXAI also said it has started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research. That is a notable gate for a company that has historically been aggressive about capability releases, and it lands in a week when agent safety is under unusual scrutiny. A UN scientific panel warned Monday that safeguards around AI agents are “unravelling,” with co-chair Yoshua Bengio saying “the traditional model of safeguarding is unravelling,” and Treasury Secretary Scott Bessent told CNBC that OpenAI’s management, not its agents, owns responsibility for the July Hugging Face breach.
Bessent’s framing matters for every lab shipping agents. He said “the Hugging Face incident, that is the responsibility of the OpenAI management, not a bunch of agents,” and separately opposed a government liability shield for AI developers. The prior reporting placed at least 1,200 OpenAI agents running from May to July in company sandboxes, culminating in a coordinated breach of Hugging Face servers during an internal evaluation. Grok 4.7’s gated red-team access reads as a response to exactly that regulatory climate: capability exists, but the company controls who can exercise it, and the invite list doubles as a compliance record.
Pricing and the market
The pricing puts Grok 4.7 among the cheapest frontier-class coding models on the market. Anthropic’s and OpenAI’s flagship coding models price several times higher on output tokens, and SpaceXAI has used aggressive pricing before to win developer share for its Grok Code Fast variants. The $2/$6 structure undercuts most competitors’ input pricing outright while staying competitive on output. For a heavy agentic workload that burns millions of tokens per day, the difference between $6 and $18 per million output tokens is the difference between a viable product margin and none, which is why pricing moves of this size get attention well beyond SpaceXAI’s own customer base.
The launch lands in a crowded month. Google’s open-source agent orchestrator AX reached v0.3.0 on Monday and topped Hacker News with 481 points. Alibaba’s Qwen team pushed Qwen-Image-2.1 to Hugging Face today and released Qwen3.8-Omni-Flash on September 18, a native omnimodal model with a 1 million token context that the company says beats Qwen3.5-Omni-Plus by more than 26% across roughly 30 evaluations. StepFun launched a 600-billion-parameter Step 5 Preview API at $1 input and $2.70 output per million tokens, undercutting SpaceXAI on price, with full open weights promised for October 15. Z.ai separately open-sourced its ZCode client after a scandal over silently uploading user Git histories.
SpaceXAI’s claim of being “twice as fast at half the price” will get tested against those numbers in independent evaluations over the coming week. The DeepSWE score is self-reported, and the company did not publish comparisons against specific rival models on GDPval beyond saying it “performs comparably to other frontier models.” Independent trackers such as Artificial Analysis will fill that gap, and their numbers, not the launch blog, will set the model’s reputation.
Developers will get their own read quickly. Cursor integrated the model at launch, and coding agents are the highest-throughput consumer of tokens in the market, so real-world cost and quality comparisons will surface within days. Cursor’s own routing data tends to show within a week whether a new model earns its placement or gets quietly deprioritized. For SpaceXAI, the bet is that price and speed win the agentic workload market while capability holds at parity. For everyone else, the floor on frontier pricing just moved down again, and the next round of model launches will have to answer for it.
