xAI released Grok 4.7 today: a bigger step on coding agents, long professional jobs, and self-checking work. The useful question for a business is the same one we asked for Sol and Astra. Where does this remove a real constraint, and where does it just raise the bill?
A better long-horizon agent is useful. A production system is still the job.
What shipped (FACT)
Grok 4.7 extends Grok 4.6 for coding, engineering, and office-style knowledge work. SpaceXAI’s model card calls it their most capable model to date, with more autonomy on longer tasks and stronger self-verification. The company news post says it works longer on difficult tasks, checks its own work more carefully, and ships with their best-calibrated safeguards to date, at the same list price and speed as Grok 4.6.
Verified product facts (model card, official news post, and named third-party writeups):
- Context window: 500,000 tokens (same class as Grok 4.6)
- Input: text and images; output: text
- Reasoning effort: low, medium, high (default), and xhigh
- Tools called out in launch coverage: function calling, web search, X search, code execution
- List pricing matching Grok 4.6: $2 / $6 per million input / output tokens, with cheaper cached input; a fast variant at twice the output speed and twice the price
- Builder surfaces today: xAI API (
grok-4.7), Cursor, Grok Build, third-party coding harnesses, and model routers / cloud platforms - GitHub Copilot: rolling out to Pro / Pro+ / Max / Business / Enterprise (gradual), billed at provider list pricing
- Consumer Grok surfaces (web, mobile, Grok-in-X): planned later per the model card, not the Day 1 default
The model card also notes supplemental training on anonymized Cursor workflow data for coding and agentic performance. Treat that as a harness hint, not a promise that every Cursor session gets free magic.
What the independent numbers say
Artificial Analysis evaluated Grok 4.7 at xhigh reasoning effort on Sept 21, 2026:
- Intelligence Index: 46 (+2 vs Grok 4.6)
- AA-Briefcase (long-horizon agentic knowledge work): 1657 Elo (+111 vs Grok 4.6 high), just behind Claude Opus 5 and Claude Fable 5.1
- GDPval-AA: 1695 Elo (+90 vs Grok 4.6 high)
- Coding Agent Index with Grok Build: 56 (+9 vs Grok 4.6 xhigh). Among models in their native harnesses, that is 4th behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5
- Hallucination rate (AA-Omniscience): 29% vs 34% for Grok 4.6 high
The catch is cost shape, not sticker price. Artificial Analysis measured roughly 81k output tokens per Intelligence Index task for Grok 4.7 (xhigh), versus about 36k for Grok 4.6 (high) and about 27k for GPT-6 Astra (max). Same list price can still mean a fatter receipt if the model thinks longer.
We will not invent Day 1 win rates for your backlog. If a number is not on the model card, the official news post, or a named eval writeup, it does not belong in this post.
What this changes for operators
1. Longer jobs got cheaper to attempt, not free to own. Grok 4.7 is aimed at multi-step coding and professional deliverables (docs, spreadsheets, analysis). That is the Astra-class job shape. You still need an owner for merge, send, spend, and deploy.
2. Harness matters as much as the base model. The Coding Agent Index jump is measured with Grok Build. Native harness results are not the same as “paste into a blank chat.” If you run agents in Cursor, Copilot, or your own stack, measure in that harness with your fixtures.
3. Self-verification is a feature, not a gate. Better checking of its own work reduces babysitting. It does not replace a human on irreversible actions. Drafts, not sent. PRs, not merged. Campaigns, not published.
4. Token burn is part of the product decision. If Grok 4.7 wins a hard agent task by spending 2x the tokens, your unit economics change even when $/1M tokens stay flat. Put a success number on the first workflow before you make it the default model everywhere.
5. Availability is split. Builders can start today in API / Cursor / Grok Build / Copilot rollout. Consumer Grok surfaces come later. Do not plan a company process around a surface that is not live yet.
What Viking Labs is doing with this
We already run Grok inside a gated fleet: specialist bots, a CoS send gate, human merge and deploy, thin v1. Grok 4.7 is a candidate engine upgrade for long coding and research-week jobs where the harness is ours and the success metric is written down first.
Practical next step for us (and for you):
- Pick one painful workflow with a clear done state.
- Run Grok 4.7 and your current default side by side on the same fixtures.
- Score quality, steps, tokens, and human edits.
- Keep the human gate on send, merge, publish, and spend.
We are not flipping every bot to 4.7 on launch day. We are testing where the long-horizon gains beat the token burn.
What we are not copying from launch day
- Leaderboard theater without a fixture from your own work
- “Autonomous company” claims that skip ownership
- Defaulting every seat to xhigh reasoning without a cost cap
- Pretending consumer Grok already has 4.7 if the model card says later
- Invented benchmark spreads or param counts we cannot source
Grok 4.7 is a better long-horizon agent. It still needs an owner, a harness, and a success number.
If you want help wrapping agents around one workflow you can measure, start at vikinglabs.com. For approve-before-write research weeks: researcherflow.com.
Leave a Reply