For most of the last two years, Anthropic owned the high ground in coding. Opus models set the bar for reliability on hard engineering work, and more recently Claude Fable 5 has been widely treated as the most accurate model available for serious coding and agentic tasks. That hierarchy is no longer stable. On 8 July 2026, xAI shipped Grok 4.5 — built explicitly for coding, agentic workflows, and knowledge work — and the story is not that it is marginally better on one leaderboard. The story is cost-performance at frontier quality, delivered extremely fast.

The shift in one line

Grok 4.5 sits in the same performance band as Opus-class and near-Fable coding systems on several key agentic benchmarks, while pricing at roughly $2 per million input tokens and $6 per million output — versus Fable 5 at $10 / $50. When quality is close enough and price diverges by that much, buying behaviour changes.

What actually shipped

Grok 4.5 is xAI's first model framed as a coding- and agent-first release rather than a general chat upgrade. Independent coverage describes a large mixture-of-experts system trained for software engineering and long-running agentic work, now default in Grok Build and available through the xAI API, Cursor, and related tooling. Elon Musk has publicly positioned it as an "Opus-class" model — language that would have sounded ambitious a year ago, and that independent benchmarks now make harder to dismiss.

On Artificial Analysis's Intelligence Index, Grok 4.5 ranks fourth overall — behind Fable 5, GPT-5.5, and Claude Opus 4.8 — after a large jump from Grok 4.3. That is not "number one." It is close enough to the frontier that price and speed become decisive for most production workloads.

$2 / $6
Grok 4.5 per 1M input / output tokens
$10 / $50
Fable 5 per 1M input / output tokens
~5×
Cheaper input than Fable 5 at sticker rate

The numbers that matter for builders

Raw leaderboard rank is less interesting than what teams actually pay per completed task. On the Coding Agent Index, Grok 4.5 in Grok Build scores 76 — matching GPT-5.5 in Codex and trailing Fable 5 in Claude Code by a single point — while average cost per task lands around $2.49 versus roughly $11.80 for Fable 5. Token volume per task is also much lower: about 1.9M tokens for Grok 4.5 against 7.2M for Fable 5 on the same suite. xAI further claims ~80 tokens per second and roughly 4.2× fewer tokens than Opus 4.8 on SWE Bench Pro-style work.

ModelInput / 1MOutput / 1MDeepSWE 1.1Terminal Bench 2.1SWE Bench Pro
Fable 5 (max)$10$5070%84.3%80.4%
GPT-5.5 (xhigh)$5$3067%83.4%58.6%
Opus 4.8 (max)$5$2559%78.9%69.2%
Grok 4.5$2$653%83.3%64.7%

Read that table carefully. Fable 5 still leads on the hardest pure coding accuracy measures — DeepSWE and SWE Bench Pro. Grok is not "better at everything." It is extremely close on Terminal Bench, competitive on agentic coding cost-efficiency, and dramatically cheaper. For many engineering organisations, that combination is more disruptive than a clean win on every benchmark, because most production agent loops are budget- and latency-constrained long before they are accuracy-maxed.

When a model is Opus-level enough, extremely fast, and five times cheaper on input, the competitive question stops being "who is smartest?" and becomes "who is still worth the premium?"

What this does to the industry

Three dynamics shift at once.

01

Frontier quality becomes a price war

Chinese open and semi-open vendors already proved that "close enough + radically cheaper" forces closed labs to respond. Grok 4.5 imports that logic into the closed frontier, from a well-capitalised Western lab with distribution through Grok Build and Cursor. The floor under coding-model pricing has dropped.

02

Agentic throughput starts to matter as much as peak IQ

Multi-agent systems, long coding sessions, and tool-using loops burn tokens. A model that is slightly behind Fable 5 on a hard GitHub-issue benchmark but uses far fewer tokens and returns answers faster can still win on total cost of ownership and developer experience. Speed is not a vanity metric here — it is loop velocity.

03

Vendor lock-in gets more expensive to defend

Teams that hard-wired themselves to a single premium model now face a visible arbitrage. That is exactly why model-routing layers and unified interfaces matter: the winning architecture treats the model as a swappable component, not the product.

The Fable 5 problem

Anthropic spent years building trust as the coding lab of record. Opus models did the heavy lifting; Fable 5 then claimed the accuracy crown for the hardest software work. That positioning still holds on several pure-accuracy measures. The problem is commercial, not technical.

Fable 5 was always priced like a flagship: $10 / $50 per million tokens on the API. For a stretch, Anthropic soft-landed that premium by including Fable 5 inside Pro, Max, Team, and some Enterprise subscriptions as a promotional window — first after earlier access disruptions, then through successive extensions. The scheduled move off subscription and fully onto usage credits was widely expected around early July. That hard cutoff has been delayed more than once; as of mid-July, Anthropic is still managing the transition carefully rather than ripping the model out of the plans that make Claude feel like a complete product.

The reason is not hard to see. If Fable 5 leaves the subscription while Grok 4.5 offers near-Opus coding performance at a fraction of the cost — and is extremely fast in day-to-day use — the subscription itself loses part of its reason to exist for power users. Coding-heavy teams do not need a philosophical argument to migrate; they need a weekend of evaluation and a routing change. Keeping Fable 5 "inside the tent" a little longer is a way of buying time: time to reprice, time to differentiate on Claude Code and ecosystem, time to decide whether the flagship stays a subscription magnet or becomes a pure metered luxury.

Caveat, not a coronation

Grok 4.5 is not free of trade-offs. Independent analysis has flagged higher hallucination rates even as knowledge scores rose — more confident when wrong is a real risk in production. Fable 5 still leads on the hardest coding accuracy suites. The point is not that every team should switch tomorrow; it is that the default assumption "Anthropic is the only serious coding stack" no longer holds without a cost model to match.

What this means if you are buying or building AI systems

For consulting clients and internal platform teams, the practical response is not "pick a winner." It is to stop treating model choice as a one-time brand decision.

01

Route by task, not by logo

Reserve Fable 5 (or whatever remains your accuracy ceiling) for the minority of tasks where the last few points of correctness are worth 5× the token bill. Route high-volume coding agents, refactors, and exploratory loops to the cheapest model that clears your quality bar — which may now be Grok 4.5 or a mid-tier Claude.

02

Measure cost per completed task, not cost per million tokens

Grok's advantage is larger once you include tokens-per-task and latency. Instrument your agent harnesses so finance and engineering share one number: dollars and minutes per successful outcome.

03

Keep the harness portable

Tools, guardrails, evals, and orchestration should outlive any single model release. If swapping Grok in for a slice of your Claude traffic requires a rewrite, the system around the model is under-built — and that is where durable value actually lives.

04

Re-evaluate subscriptions as product bundles, not model access

Claude's subscription value may increasingly rest on Claude Code, integrations, and workflow features rather than unlimited flagship model access. Budget accordingly, and do not assume promotional inclusion of Fable 5 is a permanent feature of the plan.

The executive takeaway

Grok 4.5 is a genuine industry event because it compresses the gap between "best coding model" and "good enough coding model you can run all day." Anthropic still leads on several hard accuracy measures with Fable 5, and that lead is real. What has changed is the price of defending that lead — and the risk that subscription economics collapse if the flagship is both the main reason people pay and the first thing they can replace with a faster, cheaper alternative.

The teams that win the next twelve months will not be the ones that marry a single lab. They will be the ones whose systems can move with the market: accuracy where it pays, cost-performance everywhere else, and a harness that does not care which model is briefly on top.

Key takeaways

  • Grok 4.5 delivers near-frontier coding and agentic performance at ~$2 / $6 per million tokens — a fraction of Fable 5's $10 / $50 rates.
  • On agentic coding cost-efficiency, Grok is within a point of Fable 5 on some indices while costing roughly one-fifth per task.
  • Fable 5 still leads on the hardest pure coding accuracy benchmarks; the disruption is commercial and operational, not a total technical dethroning.
  • Anthropic's careful handling of Fable 5's exit from subscriptions reflects a real risk: remove the flagship from the plan, and power users have a clear path into the Grok ecosystem.
  • The durable response is model routing, portable harnesses, and cost-per-task measurement — not loyalty to a single provider.

Revisiting your model stack after Grok 4.5?

We help teams design provider-agnostic agent systems that route for accuracy, cost, and speed — without rewriting the product every release cycle.

Talk to us →