s/mple

LLM model selection: cost per token is an input, not a metric

Dark data-centre corridor with three server racks lit at descending intensities

OpenAI released GPT-5.6 on 9 July as three tiers. Sol costs $5 per million input tokens and $30 per million output. Terra costs $2.50 and $15. Luna costs $1 and $6. A fixed five-times spread from top to bottom, per token, locked.

On SWE-Bench Pro, Luna scores 62.7 and Sol scores 64.6. Luna gets 97% of the flagship’s score for a fifth of the token price.

On long-context recall, the eight-needle test at 256K to 512K, Luna scores 41.3 and Sol scores 91.5. Same token price gap. Luna gets 45%.

The price ladder is fixed, but the score gap is not, and it moves depending on what you ask the model to do. Everything below comes out of that one observation.

Three translations, three places to lose money

Before the argument, the caveat that carries it.

A benchmark score is not capability. Capability is not business value. And a price per token is not what a task costs you.

Three translations sit between the number on a pricing page and the number on your P&L. Most of the decisions I see skip all three, which is the actual problem. Benchmark tables are useful precisely because they are the first thing people look at and the last thing they interrogate.

So read what follows as scores, because that is what they are.

Cost 1: price per token

The number on the pricing page. It is necessary for forecasting and it is not a metric, it’s an input. Comparing two models on it tells you very little about what either one will cost you.

Cost 2: price per attempt

Price per token multiplied by the tokens the model actually burns, plus cache and tool costs. This is where the first surprise lives, because token appetite varies enormously between models sitting at similar price points.

Artificial Analysis publishes the token counts from its own Intelligence Index runs. Grok 4.5 at high reasoning generated around 60 million output tokens completing the index. DeepSeek V4 Pro at max effort generated around 180 million completing the same index, and Artificial Analysis flags it as very verbose.

DeepSeek’s output tokens list at $0.87 per million. Grok’s list at $6. That is roughly a seven-times advantage on the pricing page.

It burns three times the tokens to do the same work. Most of the headline advantage disappears into verbosity before you have made a single business decision.

Cheapest per token is not cheapest per job.

A model that costs seven times less and talks three times more is not seven times cheaper. It is roughly twice as cheap.

Cost 3: price per accepted output

Cost 2, plus the human time spent reviewing it, plus the rework. These add rather than multiply, and together they are the only figure that touches your P&L.

Vendors can publish cost per task, and the better ones now do. Artificial Analysis measured Grok 4.5 at roughly $0.49 per completed GDPval task and called it nearly ninety per cent cheaper than the models ranked above it. That is genuinely useful and it is the direction the whole market needs to move in.

What no vendor can publish is your acceptance rate, because acceptance is your standard, not theirs. The last translation, the one that determines whether any of this saved you money, is the one only you can do.

The variable nobody names

Here is where the obvious conclusion breaks.

The received wisdom is to route cheap work to cheap models and reserve the expensive model for work that matters. That holds only if checking the output is cheap.

If you have a fast check available, an automated test, a second model acting as critic, a junior who can validate against a written standard, then routing cheap and verifying is the right structure. The check costs less than the upgrade.

If the only person in the business who can reliably tell good work from bad is you, and your time is the scarcest thing you own, then the expensive model is your review process. You are not buying intelligence. You are buying back your own attention, and that may well be the cheaper trade.

That is a legitimate reason to run the flagship on everything. It is just worth knowing that is the purchase you are making, rather than discovering it on the invoice.

Where marketing sits

Marketing is not uniquely damaged by any of this. It is unusually exposed, and the reasons are structural.

There is no leaderboard to defer to. GDPval, OpenAI’s evaluation of real knowledge-work deliverables, covers 44 occupations across the nine largest sectors of the US economy. It includes sales managers, editors, producers and journalists. It does not include a marketing strategy or positioning role, and no widely accepted equivalent of SWE-Bench exists for that work. So the decision falls to whoever in the room understands how these systems actually behave.

And the check is expensive, which is what matters given everything above. Code has executable ground truth available to it. Copy is judged by whoever reads it, and in most teams nobody has ever written down what good looks like. The verification loop, the thing that makes cheap routing safe, is the exact thing marketing has never built.

Reliability does not track price

A tempting shortcut is to assume the cheap model is the unreliable one. The data does not support that, and the assumption is expensive.

Artificial Analysis measured Grok 4.5’s accuracy on its AA-Omniscience test rising from 35% to 52% over the previous model, while its hallucination rate rose from 25% to 54%. Their metric counts incorrect answers as a share of all non-correct responses, including abstentions, so that is not 54% of everything it says. It is still a sharp move in the wrong direction from a model that got measurably smarter.

Grok 4.5 sits near the frontier on the Intelligence Index and is among the cheapest models in its performance tier. Cheap and unreliable are separate axes. Check both, per model, on your own work, because nothing on the pricing page tells you which way either one runs.

Uber had perfect visibility and no attribution

Uber rolled Claude Code out to roughly 5,000 engineers from December 2025 and ranked teams on internal usage leaderboards, according to The Information’s reporting. By April the annual budget was gone, and CTO Praveen Neppalli Naga told The Information he was back to the drawing board. Average spend ran $150 to $250 per engineer per month, with heavy users between $500 and $2,000. Bloomberg later reported a $1,500 monthly cap per employee per tool, with an exceptions process.

The overspend is not the interesting part.

In May, Uber’s COO Andrew Macdonald said on the Rapid Response podcast that it is very hard to draw a line between the AI usage statistics and “producing 25% more useful consumer features.”

Uber had complete visibility into consumption and no established link between consumption and value. That is an attribution problem rather than a measurement one, and attribution problems are considerably harder to fix.

Uber is among the most heavily instrumented companies on earth. If their analytics organisation could not draw that line, the odds that yours has are not good.

The decision, and the three questions that make it

1. How large is the score gap on this specific task?

2. What does being wrong cost, and can you undo it?

3. How expensive is it for you to check the work?

The third question is the one that changes the answer, and it is the one nobody asks.

Cheap to verifyExpensive to verify
Small score gapCheapest tier. Straightforward.The contested box. The expensive model is functionally a purchase of verification. Legitimate, if you know that is the purchase.
Large score gapMid tier. Iterate.Top tier, and build the check anyway. The gap is where the errors hide.

Where this argument breaks

If your AI spend is a few hundred pounds a month, optimising it is procrastination dressed as rigour. As a rule of thumb, the leverage is in what you are doing, not what you are paying for it.

Routing badly is worse than not routing. RouteLLM, from UC Berkeley and Anyscale, published at ICLR 2025, reports cost reductions of over two times on public benchmarks without sacrificing response quality. That is peer-reviewed, and it is also a GPT-4-era model pairing. Take it as evidence the mechanism works, not as a number you will hit. If your routing layer misjudges and sends hard prompts to the small model, the savings vanish into retries and quality regression that you find out about when a client does.

And the benchmark numbers themselves deserve suspicion. In April, a UC Berkeley team led by Dawn Song built an agent that scored close to perfect on eight major agent benchmarks, including 100% on SWE-bench Verified, by exploiting how the scores are computed. It solved zero tasks. That does not make every published figure fiction, but it does mean the harness matters as much as the model, and it is one more reason to test on your own work rather than trusting a table you did not build.

What to actually do

Take one task you run repeatedly. Pull five to ten real examples of it, not clean ones.

Run each through the cheapest tier and the flagship with identical instructions. Strip the model names, shuffle the order, and have someone who is not you score them against a written standard of what good means.

Writing that standard down is most of the value here, and it is the step almost nobody takes. If you cannot articulate what a good output looks like, you cannot route to a model, you cannot evaluate one, and you were never going to be able to delegate the work to a human either. The AI question is just where that gap finally becomes visible.

Log the tokens, the API cost, the review time and how often you sent it back. Then decide.

Then run it again when the next model ships, which will be in about a month.

Who this is not for

Anyone who wants a model name. There isn’t one. A recommendation with no workload attached to it is worth nothing, regardless of who is giving it.

Sources

GPT-5.6 tiers, pricing and benchmark tables. OpenAI, 9 July 2026. https://openai.com/index/gpt-5-6/ Used for: Sol, Terra and Luna pricing; SWE-Bench Pro (64.6 / 63.4 / 62.7); OpenAI MRCR v2 eight-needle 256K to 512K (91.5 / 89.6 / 41.3); GDPval-AA v2 Elo (1,747.8 / 1,593.0 / 1,591.8). Within any single row the three tiers were run under matched conditions. Rows are not necessarily comparable to each other, and no comparison in this piece requires them to be.

Grok 4.5 measurements. Artificial Analysis, accessed 11 July 2026. https://artificialanalysis.ai/models/grok-4-5 Used for: Intelligence Index score of 54; roughly 60M output tokens to complete the Intelligence Index; $2 / $6 per million token pricing.

DeepSeek V4 Pro measurements. Artificial Analysis, accessed 11 July 2026. https://artificialanalysis.ai/models/deepseek-v4-pro Used for: roughly 180M output tokens to complete the same Intelligence Index, flagged “very verbose”; $0.43 / $0.87 per million token pricing.

AA-Omniscience methodology. Artificial Analysis. https://artificialanalysis.ai/evaluations/omniscience Used for: the definition of hallucination rate as incorrect responses over all non-correct responses. Required context for the 54% figure.

Grok 4.5 cost per completed task. VentureBeat, 8 July 2026, reporting Artificial Analysis measurements. https://venturebeat.com/technology/spacexs-grok-4-5-launches-at-half-the-price-of-rivals-heres-why-that-could-rattle-anthropic-and-openai Used for: roughly $0.49 per completed GDPval task; “nearly 90% cheaper than the models ahead of it.”

GDPval occupation coverage. OpenAI, arXiv:2510.04374. https://openai.com/index/gdpval/ Used for: 44 occupations, nine sectors, and the absence of a marketing strategy or positioning role.

Uber AI budget overrun and spend caps. Praveen Neppalli Naga to The Information, April 2026; caps reported by Bloomberg, June 2026, via TechCrunch. https://techcrunch.com/2026/06/02/uber-caps-employee-ai-spending-after-blowing-through-budget-in-four-months/ Used for: roughly 5,000 engineers; budget exhausted by April; $150 to $250 average and $500 to $2,000 heavy-user monthly spend; $1,500 cap per employee per tool with exceptions process.

Andrew Macdonald on AI attribution. Rapid Response podcast, 22 May 2026. Used for: the “25% more useful consumer features” quote.

RouteLLM. Ong et al., UC Berkeley and Anyscale, ICLR 2025. arXiv:2406.18665 Used for: over two times cost reduction on public benchmarks without sacrificing response quality, on a GPT-4-1106-preview and Mixtral 8x7B pairing. This is the paper’s own abstract claim, used deliberately in place of the higher figures quoted in the authors’ project blog.

Agent benchmark exploitation. Wang, Mang, Cheung, Sen and Song, UC Berkeley Center for Responsible, Decentralized Intelligence, April 2026. https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/ Used for: near-perfect scores across eight agent benchmarks, including 100% on SWE-bench Verified, without solving any tasks.

Get started

Ready to make great things happen?

Campaigns

Brand Positioning

Content Social

Inbound Marketing

Fractional

Founder Led Marketing

Personal Brand

Portfolio

Contact

home