Grok’s Cursor Alliance Pays Off: Grok 4.6 rivals Claude Opus 5 and GPT-5.6 Sol at a lower price

Once a lab that produced mid-tier models, SpaceXAI has steadily improved. It just built one of the most capable models in the world while keeping prices relatively low.

Share
The table compares model metrics; Grok 4.6 excels in GDPVal-AA v2 and CursorBench v3.2.
Loading the Elevenlabs Text to Speech AudioNative Player...

Once a lab that produced mid-tier models, SpaceXAI has steadily improved. It just built one of the most capable models in the world while keeping prices relatively low.

What’s new: SpaceXAI introduced Grok 4.6, a vision-language model developed with Cursor and aimed at long-running agentic work. It’s available to developers now via the API, in Grok Build and Cursor, and is due in the consumer Grok apps later.

  • Input/output: Text and images in (up to 500,000 tokens), text out (no limit, 58.4 tokens per second)
  • Knowledge cutoff: February 1, 2026
  • Features: Adjustable reasoning levels (low, medium, high, xhigh — defaults to high reasoning), function calling, web search, X search, sandboxed code execution, a fast variant at double price
  • Architecture: Roughly 1.5 trillion parameters
  • Performance: Tied for third on Artificial Analysis’ Intelligence Index (61), second on GDPval-AA v2 and AA-Briefcase (1,577 Elo), top score on GPQA Diamond (94.9 percent)
  • Availability/price: Via Cursor, Grok Build coding agent, Microsoft Office add-ins, GitHub Copilot, via API at $2.00/$0.50/$6.00 per million input/cached/output tokens with higher rates for requests beyond 200,000 tokens, fast mode $4.00/$1.00/$12.00 per million input/cached/output tokens
  • Undisclosed: Architecture details, active parameter count, details of training data and methods 

How it works: Grok 4.6 is the latest model in SpaceXAI’s 1.5-trillion-parameter model family, building on Grok 4.5. SpaceXAI credits gains in performance to longer training on curated data, followed by fine-tuning on data generated by Grok 4.5 and reinforcement learning on agentic tasks. The training data included anonymized coding-agent data from Cursor, which included use of non-Grok models.

  • The company pretrained the model on publicly available, internal, and licensed data. A further round of supplemental pretraining ran longer than Grok 4.5’s equivalent stage and used what the company calls an improved optimizer and training recipe. This data included synthetic data selected for reasoning, advanced technical concepts, and software engineering data to establish a stronger foundation for later fine-tuning.
  • The company used Grok 4.5 to generate new training examples for fine-tuning, including transcripts of a model working through tasks. Grok 4.5 generated these transcripts for reasoning traces, agent harnesses, and tasks in STEM, software engineering, and knowledge work. Model-based filters removed flawed examples before Grok 4.6 trained on them.
  • The company used human and synthetic reward signals for further reinforcement learning to fine-tune the model on knowledge work and coding tasks. It also used reinforcement learning to train the model in simulated environments for writing low-level GPU code (kernel optimization), building websites, and computer-aided design. 

Performance: Grok 4.6 improved its performance on both self-reported and independently measured benchmarks, rising to near the top of the leaderboards. On many benchmarks, Grok 4.6 rivals Claude Opus 5 and GPT-5.6 Sol and does so at a lower cost per task.

  • On Artificial Analysis’ Intelligence Index, a composite of nine evaluations of economically useful tasks, Grok 4.6 set to high reasoning (61, $0.84 per task) ties for third place with GPT-5.6 Sol set to max reasoning ($1.23 per task).It jumps 5 points but more than doubles the price per task of its predecessor, Grok 4.5 set to high reasoning (56, $0.36 per task). Grok 4.6 ranks just ahead of Kimi K3 set to max reasoning (60, $0.84 per task) and just behind Claude Fable 5 set to max reasoning fallback (62, $3.14 per task).
  • On GPQA Diamond, a test of graduate-level biology, physics, and chemistry questions, Grok 4.6 achieved 94.9 percent, the highest score among models that Artificial Analysis has tested.
  • On Terminal-Bench 2.1, a test of command-line coding tasks, Grok 4.6 (88.4 percent) ranked third behind GPT-5.6 Sol set to xhigh reasoning (89.5 percent) and Claude Opus 5 set to max reasoning (89.1 percent).
  • On AA-Briefcase, Artificial Analysis’ private benchmark of four multi-week knowledge-work projects that require an agent to use thousands of files across many turns to generate deliverables such as spreadsheets, presentations, and memos, Grok 4.6 set to high reasoning (1,577 Elo) trailed only Claude Opus 5 set to max reasoning (1,715 Elo), and it edged past Claude Fable 5 set to max reasoning with fallback (1,574 Elo). Grok 4.6 reached its results in about half the turns and one quarter of the input tokens as Claude Opus 5. Likewise, on GDPval-AA v2, which tests a model’s ability to generate a single deliverable such as a document or spreadsheet, Grok 4.6 set to high reasoning (1,746 Elo) trailed only Claude Opus 5 set to max reasoning (1,849 Elo).
  • On τ³-Bench Banking, a test in which an agent resolves customer requests by searching policy documents and using tools, Grok 4.6 set to high reasoning (50.7 percent of tasks passed) trailed only Qwen3.8-Max (51.3 percent).

Behind the news: Grok 4.6 is the second model to come out of a partnership that led to an acquisition. In April, Cursor agreed to train its models on SpaceX’s Colossus supercomputer, a deal that gave SpaceX an option to buy the company. Cursor’s coding-agent data and SpaceXAI’s computation yielded results almost immediately: Grok 4.5, jointly trained with Cursor and introduced in July, lifted Grok 4.3 from 38 points on Artificial Analysis’ Intelligence Index to 56 points. SpaceX exercised its option in June, and the roughly $60 billion all-stock acquisition closed on August 14, days after Grok 4.6 launched. Three days later, Cursor introduced Origin, a code hosting service comparable to GitHub designed to handle the higher volume of code that agents generate.

Why it matters: Model makers used to tout benchmark scores at launch. Increasingly, they also publicize cost and steps per task. Grok 4.6's clearest advantage over its near competitors is completing long-running work with fewer turns. At the same price per token and task, an agent that finishes in half the turns costs around half as much, which affects what applications are feasible to build with that model.

We’re thinking: The Cursor team brought data and technical expertise to SpaceXAI, and deserves credit for supporting Grok's rapid rise in model capability. Let’s hope Grok’s continued progress and aggressive pricing makes other top labs follow suit to keep per-token prices in check.