Search

Cookies

We use cookies to improve your experience. By continuing, you accept our use of cookies.

Technology

SpaceXAI Unveils Grok 4.6 AI Model, Claims Parity with GPT-5.6 Sol & Fable 5

· · 3 min read

SpaceXAI has launched Grok 4.6, its new AI model designed for complex tasks like software engineering and research. The company claims it achieves benchmark performance comparable to GPT-5.6 Sol and Fable 5, citing upgraded training and agentic capabilities.

SpaceXAI has officially introduced Grok 4.6, its latest artificial intelligence model, positioning it as a powerful tool for demanding AI workloads. The company asserts that Grok 4.6 demonstrates competitive benchmark performance against leading models such as GPT-5.6 Sol and Fable 5, marking a significant advancement over its predecessor, Grok 4.5.

Targeting Complex AI Workflows

Grok 4.6 is engineered to excel in a range of sophisticated applications, including long-running agentic workflows, software engineering, complex visual projects, and knowledge-intensive tasks. SpaceXAI highlights the model's ability to maintain focus across multi-step processes, navigate unfamiliar codebases, and critically assess its own outputs, making it suitable for intricate problem-solving scenarios.

Enhanced Training and Reinforcement Learning

The development of Grok 4.6 involved an extended and refined training phase compared to Grok 4.5. This process incorporated an upgraded optimizer, advanced training techniques, and the integration of both engineering data and meticulously curated synthetic datasets. These datasets were specifically designed to bolster the model's grasp of complex technical concepts and improve its reasoning capabilities.

During its supervised fine-tuning (SFT), trajectories were regenerated across diverse domains like reasoning problems, agent environments, software engineering, STEM fields, and general knowledge work. Evaluator models then filtered out weaker trajectories, creating an SFT checkpoint aimed at optimizing the model's behavior and overall performance.

Further enhancing its capabilities, Grok 4.6 underwent extensive training in agentic reinforcement-learning environments. These environments covered areas such as knowledge tasks, software development, computer-aided design (CAD), web development, and kernel optimization, contributing to its robust and versatile performance.

Benchmark Performance Claims

SpaceXAI has released benchmark results indicating Grok 4.6's competitive standing. On the AA Intelligence Index, Grok 4.6 scored 61, matching GPT-5.6 Sol Max, while Fable 5 Max achieved 62. The new Grok model reportedly led its competitors on GDPVal-AA v2, scoring 1,753, and on AA-Briefcase, with a score of 1,577.

However, the model did trail in some specific tests, notably Terminal-Bench v3.0, where it scored 26% compared to 34.6% for GPT-5.6 Sol Max and 34.1% for Fable 5 Max. These results suggest a strong overall performance with specific areas for further development.

Availability and Pricing

Grok 4.6 is now accessible through various platforms, including Grok Build, Cursor, the SpaceXAI API, and partner platforms such as Cloudflare, OpenRouter, and Vercel. The base API pricing is set at $2 per million input tokens and $6 per million output tokens. A high-speed variant is also available at double the base price.

During its launch week, subscribers to Grok Build and Cursor will benefit from double their usual Grok 4.6 usage limits, providing an incentive for early adoption.

Related