xAI releases Grok 4.7 through coding tools and APIs
xAI released Grok 4.7 through Grok Build, APIs, Cursor, and other gateways. Early evaluations report stronger coding and knowledge-work results than Grok 4.6, with mixed results across individual coding benchmarks.

TL;DR
- xAI shipped Grok 4.7 into Cursor, Grok Build, APIs, and third-party agent surfaces, with ericzakariasson's launch post naming the main availability channels.
- The release uses a new, larger base model and a longer reinforcement-learning run, while keeping the $2/$6 per-million-token starting price from Grok 4.6, according to ericzakariasson's model card post.
- First-party coding results moved up: CursorBench went from 40.4% to 46.3%, +5.9 points, while Terminal-Bench went from 20.3% to 38.0%, +17.7 points, in ericzakariasson's comparison.
- Independent testing put Grok 4.7 at 46 versus Grok 4.6's 44 on the Artificial Analysis Intelligence Index, +2 points, but measured 81k output tokens per task versus 36k, according to ArtificialAnlys's analysis.
- The rollout already spans OpenRouter, Devin, Droid, Cline, Vercel AI Gateway, v0, Venice, and Agent Arena, with Cognition's Devin announcement and Vercel's rollout post showing how quickly it moved into coding products.
The official announcement pairs the model card with a 500k context window, a Fast tier, and claims about self-verification on multi-hour tasks. Artificial Analysis found a large reasoning-token bill behind the gains. A Cursor community thread also records a four-tier price table that makes long-context and Fast requests materially more expensive.
What shipped
- Availability: Grok 4.7 launched in Cursor, Grok Build, and the Grok API, then appeared in third-party coding harnesses, model routers, and cloud platforms, as ericzakariasson's launch post describes.
- API access: xAI exposes the model through its developer platform, while koltregaskes's release notice and ericzakariasson's follow-up confirm the public rollout.
- Reasoning modes: low, medium, high, and xhigh are supported, with Artificial Analysis evaluating xhigh in ArtificialAnlys's analysis.
- Starting price: $2 per million input tokens and $6 per million output tokens, matching Grok 4.6 in ericzakariasson's model card post.
- Fast variant: twice the output speed at twice the starting price, according to the official release.
- Vercel promotion: AI Gateway, fx, and eve offered 40% off through September 27, Vercel Dev's availability post said.
- Rollout timing: Grok 4.7 briefly appeared on OpenCode before the formal launch, according to kimmonismus's pre-release note, and AILeaksAndNews's release notice called it the first model out in that release wave.
Benchmarks that moved
First-party
- CursorBench 4.0: 40.4% → 46.3%, +5.9 points, per ericzakariasson's model card post.
- DeepSWE v1.1: 65.2% → 71.0%, +5.8 points, at high effort in ericzakariasson's model card post.
- Terminal-Bench 4.0: 20.3% → 38.0%, +17.7 points, per ericzakariasson's model card post.
- AA Briefcase v1.1: 1,546 → 1,657 Elo, +111 Elo, in ericzakariasson's model card post.
- Harvey Legal Agent: 15.8% → 19.6%, +3.8 points, in ericzakariasson's model card post.
- EEBench: 53.0% → 64.0%, +11.0 points, in ericzakariasson's model card post.
Third-party evaluators
- Artificial Analysis Intelligence Index: 44 → 46, +2 points, per ArtificialAnlys's analysis.
- Artificial Analysis Coding Agent Index: 47 → 56, +9 points, with Grok Build in ArtificialAnlys's analysis.
- AA-Briefcase: 1,546 → 1,657 Elo, +111 Elo, per ArtificialAnlys's deeper breakdown.
- GDPval-AA: 1,605 → 1,695 Elo, +90 Elo, per ArtificialAnlys's deeper breakdown.
- Grok Build DeepSWE v1.1: 65% → 73%, +8 points, in ArtificialAnlys's native-harness evaluation.
- AA-Omniscience hallucination rate: 34% → 29%, -5 points, while accuracy moved from 48% to 47%, per ArtificialAnlys's analysis.
- Long Horizon Browser Use Benchmark v2: 31.2% → 39.9%, +8.7 points, according to browser_use's benchmark result.
Customer-reported
- FrontierCode 1.1 Extended: Grok 4.6 scored higher, with the numeric baseline undisclosed → Grok 4.7 at 59.4%, delta not reported, in Cognition's FrontierCode result.
Where it regressed
Grok 4.7's overall movement depends heavily on the benchmark and effort setting.
- Vals Index: 59.2% → 54.2%, -5.0 points, moving Grok 4.6 from rank 14 to Grok 4.7 at rank 24, per ValsAI's benchmark result.
- AA-LCR v1.1: Grok 4.6 → Grok 4.7, -3.7 points, in ArtificialAnlys's breakdown.
- AutomationBench-AA: Grok 4.6 → Grok 4.7, -1.1 points, in ArtificialAnlys's breakdown.
- CursorBench at low effort: 33.4% → 33.1%, -0.3 points, while steps rose from 32 to 40, according to chetaslua's effort-matched comparison.
- Vals cost per test: $4.34 → $4.78, +$0.44, even as the headline score fell, per ValsAI's benchmark result.
The official model card also places Grok 4.7 below Fable 5.1 on Terminal-Bench, HealthBench Professional, and AA Briefcase, despite its gains over Grok 4.6. The comparison table gives Terminal-Bench as 38.0% for Grok 4.7 versus 57.9% for Fable 5.1, and HealthBench as 56.7% versus 62.1%, in ericzakariasson's model card post.
Under the hood
xAI describes Grok 4.7 as a new base-model generation rather than another post-training pass on Grok 4.6. The official announcement says the model uses a larger base model, a longer reinforcement-learning run, and a harder task mixture weighted toward problems that take hours to complete.
The training target is visible in the product behavior xAI names:
- More persistent work on difficult tasks.
- More self-verification.
- Better long-context management.
- Native understanding of the Grok Bot harness for conversation and knowledge work.
The context window remains 500k tokens, unchanged from Grok 4.6, while reasoning effort spans low through xhigh, according to ArtificialAnlys's analysis. The public release material does not publish a parameter count or tokenizer change.
Pricing has more layers than the $2/$6 headline. A Cursor pricing clarification lists standard Grok 4.7 at $2 input and $6 output, 500k-context requests at $4/$12, Fast at $4/$12, and 500k Fast at $6/$18 per million tokens.
The same screenshot shows a 200k threshold in practice: 34 calls crossed it during a CursorBench session, and the user recorded $39.07 in total usage, including Fast-mode charges. davis7's usage report also breaks out $17.34 attributed to the Fast premium.
Contested claims
Claim: Grok 4.7 keeps the cost advantage of Grok 4.6 in real workloads.
Cited by: xAI presents identical starting prices for the two models in its official comparison, and ericzakariasson's model card post shows both at $2/$6.
Counter: Artificial Analysis measured 81k output tokens per Intelligence Index task for Grok 4.7 versus 36k for Grok 4.6, while ValsAI's result puts cost per test at $4.78 versus $4.34. theo's cost chart shows Grok 4.7 at $3.74 on its aggregate comparison against GPT-6 Astra at $3.26.
Evidence so far: The token price is unchanged at the base tier, but generated-token volume, Fast mode, and long-context pricing change the completed-task bill.
Claim: Grok 4.7's larger token budget is buying more reliable work rather than simply longer answers.
Cited by: xAI says the longer RL run improves persistence and self-checking, while ArtificialAnlys's analysis reports higher AA-Briefcase analytical quality, from 1,690 to 1,994 Elo.
Counter: kunchenguid's day-one report describes the model as slower and visibly more expensive than Grok 4.5, and theo's review says real-world costs were more than twice Grok 4.6 in some runs.
Evidence so far: The independent results show gains on long-horizon knowledge work alongside higher output-token use, while the hands-on cost reports remain workload-dependent.
Claim: Terminal-Bench's coding improvement is a single comparable result.
Cited by: xAI's model card reports 20.3% → 38.0%, +17.7 points, in ericzakariasson's model card post.
Counter: Artificial Analysis reports 18% → 33%, +15 points, for its Grok Build component, and daniel_mac8's benchmark question asks whether its Terminal-Bench result differs from the official figure.
Evidence so far: The sources use different evaluation presentations and effort or harness labels, so the two deltas establish improvement without being interchangeable scores.
Vibe Check
The first hands-on reports split between speed and steadiness on one side, and token burn plus weak visual work on the other.
- System-prompt adherence: kunchenguid reports that Grok 4.7 followed Firstmate's CI-bypass instructions closely enough to ask for named checks and reject a blanket “yolo” instruction.
- Behavior: kunchenguid describes it as stable and conservative, asking for confirmation before ambiguous actions rather than producing the “spiky” behavior associated with Astra.
- Speed: davis7's hands-on test calls the new Fast mode extremely fast and lists subagents, long-running tasks, and looping as usable capabilities.
- Reverse engineering: rauchg's test says Grok 4.7 solved a hard running-binary reverse-engineering problem quickly.
- 3D pipeline: ai_for_success's test takes a floor plan through Blender and into an interactive Three.js site.
- Frontend and 3D work: theo's review reports poor frontend and design performance, nonexistent 3D capability, and occasional loops.
- Daily use: mattshumer_'s hands-on report calls it a major step up from Grok 4.6 and a strong daily driver.
Where it shows up
Grok 4.7's rollout is unusually broad for a same-day model release.
- Cursor and Grok Build: xAI listed both as day-one surfaces, and TestingCatalog's release report also identified the API as available immediately.
- AI Gateway, fx, and eve: Vercel Dev's availability post announced the model with a 40% discount, while Vercel's rollout post documented the shared model slug and gateway path.
- OpenRouter: OpenRouter's launch post made Grok 4.7 available as a routed coding and knowledge-work model.
- Devin: Cognition's availability post added Grok 4.7 to Devin CLI and Desktop, with a 59.4% FrontierCode 1.1 Extended result.
- Droid: FactoryAI's rollout note reported strength across engineering, debugging, data, and infrastructure work, with Medium as its default effort and High helping on legacy code.
- Cline: Cline's release post shipped the model in its desktop app with a 40% promotion and described it as roughly eight times cheaper than frontier alternatives on DeepSWE.
- v0: v0's availability post added Grok 4.7 to its application-building workflow.
- Venice: AskVenice's rollout post made the model privately available through Venice.
- Agent Arena: Arena's launch post put Grok 4.7 into Battle Mode and Agent Mode with web search, filesystem, and terminal tools.
- Next.js evals: nextjs's leaderboard post placed Grok 4.7 in the top three, at 94% success alongside Claude Opus 5.