Skip to content
AI Primer
release

xAI releases Grok 4.7 through coding tools and APIs

xAI released Grok 4.7 through Grok Build, APIs, Cursor, and other gateways. Early evaluations report stronger coding and knowledge-work results than Grok 4.6, with mixed results across individual coding benchmarks.

9 min read
xAI releases Grok 4.7 through coding tools and APIs
xAI releases Grok 4.7 through coding tools and APIs

TL;DR

  • xAI shipped Grok 4.7 into Cursor, Grok Build, APIs, and third-party agent surfaces, with ericzakariasson's launch post naming the main availability channels.
  • The release uses a new, larger base model and a longer reinforcement-learning run, while keeping the $2/$6 per-million-token starting price from Grok 4.6, according to ericzakariasson's model card post.
  • First-party coding results moved up: CursorBench went from 40.4% to 46.3%, +5.9 points, while Terminal-Bench went from 20.3% to 38.0%, +17.7 points, in ericzakariasson's comparison.
  • Independent testing put Grok 4.7 at 46 versus Grok 4.6's 44 on the Artificial Analysis Intelligence Index, +2 points, but measured 81k output tokens per task versus 36k, according to ArtificialAnlys's analysis.
  • The rollout already spans OpenRouter, Devin, Droid, Cline, Vercel AI Gateway, v0, Venice, and Agent Arena, with Cognition's Devin announcement and Vercel's rollout post showing how quickly it moved into coding products.

The official announcement pairs the model card with a 500k context window, a Fast tier, and claims about self-verification on multi-hour tasks. Artificial Analysis found a large reasoning-token bill behind the gains. A Cursor community thread also records a four-tier price table that makes long-context and Fast requests materially more expensive.

What shipped

Benchmarks that moved

First-party

Third-party evaluators

Customer-reported

  • FrontierCode 1.1 Extended: Grok 4.6 scored higher, with the numeric baseline undisclosed → Grok 4.7 at 59.4%, delta not reported, in Cognition's FrontierCode result.

Where it regressed

Grok 4.7's overall movement depends heavily on the benchmark and effort setting.

The official model card also places Grok 4.7 below Fable 5.1 on Terminal-Bench, HealthBench Professional, and AA Briefcase, despite its gains over Grok 4.6. The comparison table gives Terminal-Bench as 38.0% for Grok 4.7 versus 57.9% for Fable 5.1, and HealthBench as 56.7% versus 62.1%, in ericzakariasson's model card post.

Under the hood

xAI describes Grok 4.7 as a new base-model generation rather than another post-training pass on Grok 4.6. The official announcement says the model uses a larger base model, a longer reinforcement-learning run, and a harder task mixture weighted toward problems that take hours to complete.

The training target is visible in the product behavior xAI names:

  • More persistent work on difficult tasks.
  • More self-verification.
  • Better long-context management.
  • Native understanding of the Grok Bot harness for conversation and knowledge work.

The context window remains 500k tokens, unchanged from Grok 4.6, while reasoning effort spans low through xhigh, according to ArtificialAnlys's analysis. The public release material does not publish a parameter count or tokenizer change.

Pricing has more layers than the $2/$6 headline. A Cursor pricing clarification lists standard Grok 4.7 at $2 input and $6 output, 500k-context requests at $4/$12, Fast at $4/$12, and 500k Fast at $6/$18 per million tokens.

The same screenshot shows a 200k threshold in practice: 34 calls crossed it during a CursorBench session, and the user recorded $39.07 in total usage, including Fast-mode charges. davis7's usage report also breaks out $17.34 attributed to the Fast premium.

Contested claims

Claim: Grok 4.7 keeps the cost advantage of Grok 4.6 in real workloads.

Cited by: xAI presents identical starting prices for the two models in its official comparison, and ericzakariasson's model card post shows both at $2/$6.

Counter: Artificial Analysis measured 81k output tokens per Intelligence Index task for Grok 4.7 versus 36k for Grok 4.6, while ValsAI's result puts cost per test at $4.78 versus $4.34. theo's cost chart shows Grok 4.7 at $3.74 on its aggregate comparison against GPT-6 Astra at $3.26.

Evidence so far: The token price is unchanged at the base tier, but generated-token volume, Fast mode, and long-context pricing change the completed-task bill.

Claim: Grok 4.7's larger token budget is buying more reliable work rather than simply longer answers.

Cited by: xAI says the longer RL run improves persistence and self-checking, while ArtificialAnlys's analysis reports higher AA-Briefcase analytical quality, from 1,690 to 1,994 Elo.

Counter: kunchenguid's day-one report describes the model as slower and visibly more expensive than Grok 4.5, and theo's review says real-world costs were more than twice Grok 4.6 in some runs.

Evidence so far: The independent results show gains on long-horizon knowledge work alongside higher output-token use, while the hands-on cost reports remain workload-dependent.

Claim: Terminal-Bench's coding improvement is a single comparable result.

Cited by: xAI's model card reports 20.3% → 38.0%, +17.7 points, in ericzakariasson's model card post.

Counter: Artificial Analysis reports 18% → 33%, +15 points, for its Grok Build component, and daniel_mac8's benchmark question asks whether its Terminal-Bench result differs from the official figure.

Evidence so far: The sources use different evaluation presentations and effort or harness labels, so the two deltas establish improvement without being interchangeable scores.

Vibe Check

The first hands-on reports split between speed and steadiness on one side, and token burn plus weak visual work on the other.

  • System-prompt adherence: kunchenguid reports that Grok 4.7 followed Firstmate's CI-bypass instructions closely enough to ask for named checks and reject a blanket “yolo” instruction.
  • Behavior: kunchenguid describes it as stable and conservative, asking for confirmation before ambiguous actions rather than producing the “spiky” behavior associated with Astra.
  • Speed: davis7's hands-on test calls the new Fast mode extremely fast and lists subagents, long-running tasks, and looping as usable capabilities.
  • Reverse engineering: rauchg's test says Grok 4.7 solved a hard running-binary reverse-engineering problem quickly.
  • 3D pipeline: ai_for_success's test takes a floor plan through Blender and into an interactive Three.js site.
  • Frontend and 3D work: theo's review reports poor frontend and design performance, nonexistent 3D capability, and occasional loops.
  • Daily use: mattshumer_'s hands-on report calls it a major step up from Grok 4.6 and a strong daily driver.

Where it shows up

Grok 4.7's rollout is unusually broad for a same-day model release.

tweetGroups

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 8 threads
TL;DR4 posts
What shipped7 posts
Benchmarks that moved4 posts
Where it regressed2 posts
Under the hood1 post
Contested claims4 posts
Vibe Check5 posts
Where it shows up11 posts
Share on X