xAI's Grok 4.7 Beats Claude at Enterprise Analysis but Doubles Your Bill
xAI's Grok 4.7 jumps 111 Elo on Artificial Analysis's agentic knowledge work benchmark, closing the gap with Anthropic at roughly half the cost per task.
- Grok 4.7 hits 1657 Elo on AA-Briefcase, +111 over Grok 4.6, just behind Claude Opus 5.
- Analytical Quality Elo jumps 1690 to 1994; Presentation Elo dips slightly from 1519 to 1499.
- Cost per task roughly $8 (xhigh) vs $4.40 for Grok 4.6, but half of Opus 5.
- Uses ~81,000 output tokens per Intelligence Index task, 2x Grok 4.6, 3x GPT-6 Astra.
- Coding Agent Index rises 47 to 56; Terminal-Bench nearly doubles from 18% to 33%.
- Pricing unchanged: $2/M input, $6/M output, 500k context window on the xAI API.
Grok 4.7 improves enterprise analysis at a higher token cost
Grok 4.7 scored 1,657 Elo on AA-Briefcase, 111 points above Grok 4.6 (high) and just behind Claude Opus 5 and Claude Fable 5.1. The benchmark covers professional deliverables such as market models and acquisition-target assessments. Grok 4.7 completed those tasks at roughly half the cost of Anthropic’s flagship, according to Artificial Analysis.
AA-Briefcase evaluates multi-step knowledge work that ends in a document, spreadsheet, or presentation. Its Elo score reflects relative performance in head-to-head comparisons, so rankings can change as the model pool and evaluations evolve. AA-Briefcase-Lite provides public examples of the underlying tasks.
Analysis rises while polish slips
Grok 4.7’s improvement came from analytical quality, which rose by 304 Elo points. Its presentation score fell by 20 points, indicating stronger reasoning without a corresponding gain in formatting or visual delivery.
| Metric | Grok 4.6 (high) | Grok 4.7 | Change |
|---|---|---|---|
| Overall Elo | 1,546 | 1,657 | +111 |
| Analytical quality | 1,690 | 1,994 | +304 |
| Presentation quality | 1,519 | 1,499 | -20 |
What changed in the sample tasks
- Valuation chain: A private-equity template required comparable-company benchmarking. Grok 4.6 relied on the deal partner’s shorthand estimate, while Grok 4.7 calculated the valuation independently and identified a discrepancy.
- Asset profile: Grok 4.6 used only the latest annual figures and omitted trend and currency analysis. Grok 4.7 examined three years of data, found that currency depreciation had erased revenue growth, and changed the resulting acquisition recommendation.
Deeper reasoning expands the bill
Artificial Analysis measured an API cost of about $8 for an example deck generated with Grok 4.7 (xhigh), compared with $4.40 for Grok 4.6 (xhigh). That 82% increase tracks the model’s much larger output.
| Measurement | Grok 4.6 | Grok 4.7 | Difference |
|---|---|---|---|
| Example deck cost at xhigh | About $4.40 | About $8 | About 82% higher |
| Output tokens per Intelligence Index task | About 36,000 at high | About 81,000 at xhigh | About 125% higher |
GPT-6 Astra (max) used about 27,000 output tokens on the same index, making Grok 4.7’s output roughly three times as large. Artificial Analysis reports the difference as approximately 196%.
Grok’s list pricing remains $2 per million input tokens and $6 per million output tokens, with cached input priced at $0.50 per million. The model retains a 500,000-token context window. Its low token prices keep the completed-task cost below Opus 5 in this benchmark, although higher output volume absorbs part of that advantage.
The published comparisons use different reasoning settings in some rows, including high for the Grok 4.6 token baseline and xhigh for Grok 4.7. Teams should treat setting-matched results as the cleaner measure of model-to-model improvement and reproduce cost tests with their intended configuration.
Coding agents gain at the terminal
Grok Build running Grok 4.7 (xhigh) scored 56 on the Artificial Analysis Coding Agent Index, up from 47 with Grok 4.6 (xhigh). All three components improved.
| Benchmark | Grok 4.6 (xhigh) | Grok 4.7 (xhigh) | Change |
|---|---|---|---|
| DeepSWE v1.1 | 65% | 73% | +8 points |
| Terminal-Bench 4.0 | 18% | 33% | +15 points |
| SWE-Atlas-QnA | 58% | 63% | +5 points |
Terminal-Bench produced the largest gain. It tests whether an agent can complete command-line tasks in a controlled environment, making it relevant to systems that edit code, run tools, inspect failures, and iterate without step-by-step human direction.
Results across the broader Intelligence Index were mixed. Grok 4.7 improved on Terminal-Bench 4.0 by 4.5 percentage points and GDP.pdf by 3 points, while AA-LCR fell by 3.7 points and AutomationBench-AA fell by 1.1 points. The AA-LCR decline warrants workload-specific testing for applications that depend on retrieving information from long contexts.
High throughput still produces seven-minute runs
Artificial Analysis measured output throughput of about 188 tokens per second on long prompts, while an Intelligence Index task took an average of 7.1 minutes. The combination of high generation speed and long completion time follows from the model’s roughly 81,000 output tokens per task.
Where Grok 4.7 fits
Grok 4.7 aligns most closely with document-heavy agents that must inspect evidence, challenge supplied assumptions, perform multi-step calculations, and produce a deliverable. Due-diligence memos, financial analysis, market sizing, and acquisition research resemble the tasks on which its analytical score improved.
- Favor Grok 4.7 in evaluations when analytical depth, a 500,000-token context window, and lower benchmark cost per completed task carry more weight than presentation polish.
- Compare Anthropic models directly when final-document quality is a primary acceptance criterion or when the workflow depends heavily on long-context retrieval.
- Rebuild token budgets before migrating from Grok 4.6 because prior per-task estimates will understate Grok 4.7’s output volume.
- Set latency and output limits for interactive agents, where seven-minute average runs may exceed product requirements.
Production evaluation should also measure tool-call reliability, retry rates, permission handling, citation accuracy, structured-output compliance, and the amount of human revision required. AA-Briefcase captures deliverable quality, while those operational concerns determine whether an agent works reliably inside an enterprise system.
xAI’s claims meet a mixed presentation result
xAI release notes describe GDPval and AA-Briefcase as evaluations modeled on work performed by professionals including lawyers, nurses, and financial analysts. The company says Grok 4.7 improves on Grok 4.6 across both benchmarks and performs comparably to other frontier models.
Artificial Analysis supports the claim of stronger professional reasoning but adds an important qualification: presentation quality slipped even as analytical quality climbed. Procurement teams should therefore compare cost per accepted deliverable rather than list token prices alone, accounting for output volume, latency, formatting quality, and human review.