xAI ships Grok 4.6, matches GPT-5.6 Sol on benchmark index

xAI released Grok 4.6 today, building on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. The company says the model stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application. xAI states that Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks, and that it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite score built from nine benchmarks; no individual benchmark scores are given for either model.
On the training side, xAI says Grok 4.6 went through a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe, which it describes as producing a stronger foundation for the SFT and RL stages that followed. Grok 4.5 was then used to regenerate SFT trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, with problematic traces filtered out by model-based checks. Reinforcement learning covered a wide range of agentic tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and computer-aided design.
xAI reports that in internal testing on projects meant to stretch the model's range, Grok 4.6 was especially strong at turning a broad product idea into a working first version: researching an unfamiliar domain, structuring the application, implementing core interactions, and refining the result through several rounds of feedback. On longer trajectories, xAI says it also started seeing more self-testing and verification, with the model checking its own work before moving on. The company also says Grok 4.6 produces stronger first passes on visual and interactive projects than Grok 4.5 typically did, establishing structure and visual language for an application in one pass.
xAI says Grok 4.6's safeguards have been improved and calibrated to match the model's capabilities, with its widest-ever suite of pre-deployment testing plus extensive post-deployment and third-party testing, aimed at keeping the model helpful and safe in areas such as vulnerability patching, engineering design work, and AI research.
Grok 4.6 is available today in Cursor and Grok Build, as well as through the API and partners including OpenRouter, Vercel, and Cloudflare. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a faster variant priced at twice the standard rate. For the first week, xAI is offering 2x included usage inside Grok Build and Cursor.
Key facts
- xAI released Grok 4.6, built on Grok 4.5, with a stated focus on long-running agent tasks and more ambitious visual and interactive work.
- xAI says Grok 4.6 matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks; no individual scores are disclosed.
- Training included a longer supplemental run than Grok 4.5, curated model-generated data, an improved optimizer, SFT trajectories regenerated via Grok 4.5, and agentic RL covering coding, kernel optimization, web development, and CAD.
- API pricing starts at $2 per million input tokens and $6 per million output tokens; a faster variant costs twice as much.
- Grok 4.6 is live today in Cursor, Grok Build, the API, and partners OpenRouter, Vercel, and Cloudflare, with 2x included usage in Grok Build and Cursor for the first week.
Why it matters
Grok 4.6 marks xAI's push to keep pace at the frontier while shifting emphasis toward agents that sustain work over many steps and toward one-pass visual and interactive output, rather than chasing a single headline benchmark number. The claimed parity with GPT-5.6 Sol on a nine-benchmark composite index positions Grok 4.6 as a frontier-tier model in xAI's telling, though the comparison rests on one aggregate score rather than a benchmark-by-benchmark breakdown.
Who it affects
Developers building agentic coding tools and AI-assisted applications, since Grok 4.6 is already wired into Cursor and xAI's own Grok Build. Teams evaluating API providers for long-running research, codebase work, or turning product ideas into working applications are the other direct audience, alongside anyone accessing the model through OpenRouter, Vercel, or Cloudflare.
How to use it
Grok 4.6 is available today in Cursor and Grok Build, through xAI's API, and via partners OpenRouter, Vercel, and Cloudflare. API pricing starts at $2 per million input tokens and $6 per million output tokens; a faster variant is priced at twice the standard rate. For the first week, xAI is offering 2x included usage inside Grok Build and Cursor so users can try the model at reduced effective cost.
How solid is it
The claims come entirely from xAI's own release announcement. The central benchmark claim, matching GPT-5.6 Sol on the Artificial Analysis Intelligence Index, is reported as a single composite outcome; the nine underlying benchmarks are not named and no individual scores are given for either model, so the comparison cannot be checked against a public leaderboard from the announcement text alone. Claims about training methodology, self-testing behavior on long trajectories, and safety calibration are likewise self-reported, with no third-party figures cited in the source.
Risks and caveats
Every performance and safety claim in the announcement is xAI describing its own model, with no independent benchmark results, no named third-party evaluator, and no specific date given beyond "today." The source does not say how much longer the supplemental training run was compared with Grok 4.5's, nor what specifically makes the fast API variant faster beyond its price. The promotional 2x included usage offer has no stated end date beyond "the first week."
“It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks.”
— xAI, Grok 4.6 release announcement