SpaceXAI Releases Grok 4.6 to Challenge GPT-5.6 Sol
SpaceXAI has launched Grok 4.6, a highly capable vision-language model that rivals top-tier competitors at a lower cost, signaling a major shift in the economics of agentic AI development.

SpaceXAI has released Grok 4.6, a vision-language model with roughly 1.5 trillion parameters designed for agentic tasks. Developed with Cursor, which SpaceXAI acquired in a 60 billion dollar all-stock deal on August 14, the model is available via API, Cursor, Grok Build, Microsoft Office, and GitHub Copilot. It supports up to 500,000 input tokens of text and images, and outputs text at 58.4 tokens per second. The API costs 2.00, 0.50, and 6.00 dollars per million input, cached, and output tokens, with higher rates beyond 200,000 tokens. A fast variant doubles these prices to 4.00, 1.00, and 12.00 dollars.
The model achieves highly competitive scores across major benchmarks. On Artificial Analysis’ Intelligence Index, Grok 4.6 set to high reasoning scored 61 at 0.84 dollars per task, tying for third with GPT-5.6 Sol set to max reasoning at 1.23 dollars, and beating Kimi K3 set to max reasoning at 60 and 0.84 dollars. It sits just behind Claude Fable 5 set to max reasoning fallback at 62 and 3.14 dollars, while improving on Grok 4.5, which scored 56 at 0.36 dollars and lifted Grok 4.3 from 38 points. Grok 4.6 also scored 94.9 percent on GPQA Diamond. On Terminal-Bench 2.1, it scored 88.4 percent, trailing GPT-5.6 Sol set to xhigh reasoning at 89.5 percent and Claude Opus 5 set to max reasoning at 89.1 percent.
On the agentic benchmark AA-Briefcase, Grok 4.6 set to high reasoning scored 1,577 Elo, trailing Claude Opus 5 set to max reasoning at 1,715 Elo but beating Claude Fable 5 set to max reasoning fallback at 1,574 Elo. Grok 4.6 finished in half the turns and used one quarter of the input tokens of Claude Opus 5. On GDPval-AA v2, it scored 1,746 Elo, trailing Claude Opus 5 at 1,849 Elo. On τ³-Bench Banking, it scored 50.7 percent, just behind Qwen3.8-Max at 51.3 percent. SpaceXAI trained the model using synthetic reasoning data generated by Grok 4.5, Cursor coding data, and reinforcement learning on tasks like writing low-level GPU code, building websites, and computer-aided design.
For practitioners, these developments represent a significant shift in the economics of building AI agents. Because Grok 4.6 can complete complex, multi-turn projects in far fewer steps than rivals like Claude Opus 5, developers can deploy highly capable agents at a fraction of the operational cost. This efficiency makes previously cost-prohibitive, long-running agentic workflows commercially viable.
This is our own summary of reporting by The Batch


