Warp Brings xAI's Grok 4.6 to Its Terminal, Matching GPT-5.6 on Key Benchmarks

Grok 4.6, xAI's frontier-class agentic model, lands in Warp Terminal and the new Warp Agent CLI via X Premium or SuperGrok

·
·
AuthorWarp
Read2 min
  • Grok 4.6 is now available in Warp Terminal and the Warp Agent CLI via /connect-grok using an X Premium or SuperGrok subscription.
  • Grok 4.6 matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index (score: 61) and scores 69.9% on CursorBench v3.2.
  • The model keeps the same 1.5T parameter base as Grok 4.5 but adds a longer SFT and RL post-training run focused on agentic and multi-step tasks.
  • It features a 500K token context window and is priced at $2/M input tokens and $6/M output tokens via the SpaceXAI API.
  • The Warp Agent CLI uses a tmux-like multiplexing architecture, letting the agent drive interactive apps like sqlite, vim, and Python REPLs natively.
  • Grok 4.6 is also available in Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare, with 2x usage included in Cursor and Grok Build for the first week.

Warp just added Grok 4.6 to both its terminal app and the newly launched Warp Agent CLI. If you have an X Premium or SuperGrok subscription, you can connect right now by running /connect-grok inside Warp -- no separate API key needed.

What Grok 4.6 actually is

Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application.

The distinguishing bet in Grok 4.6 is capability gains through post-training, not scale. Where most frontier releases lean on a bigger base model, xAI held the 1.5T V9 foundation constant and poured the delta into supervised fine-tuning and reinforcement learning (RL). Supervised fine-tuning (SFT) shapes the model's behavior using curated examples; RL then pushes it further by rewarding good outcomes on hard tasks rather than just imitating examples.

Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe.

How it benchmarks

Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks, matching GPT-5.6 Sol on the Artificial Analysis Intelligence Index -- a composite score of nine benchmarks. Here is how it stacks up on the key evals:

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves