Aug 5, 202612 min read/2026/08/05/a-barrel-is-a-barrel-a-token-is-not/

A Barrel Is a Barrel. A Token Is Not a Token.

Forget language models for a few minutes. I want to start with corn.

What a commodity actually is

A commodity is a good whose units are interchangeable regardless of who produced them. A barrel of West Texas Intermediate is a barrel of WTI. The buyer does not care which well it came from, and more importantly, cannot tell. Same for a bushel of No. 2 yellow corn or a troy ounce of .995 gold.

The usual definition — "a raw material" — is wrong, or at best accidental. Raw materials tend to be commodities because nature doesn't brand things, but cement, A4 paper, DRAM and shipping containers are all manufactured and thoroughly commoditized. The test isn't whether it came out of the ground. It's whether the buyer can swap one seller's unit for another's without caring.

Everything else follows from that one property. If units are interchangeable, buyers choose on price alone, so you don't set your price — you accept the market's and decide only how much to produce at it. You're a price taker. Margins compress toward the cost of capital, because any producer earning excess returns invites entry until the return is just enough to keep capital in the business. Your strategy collapses to two levers: sit lower on the cost curve than the marginal producer, or hedge better. Marketing does nothing for you. You cannot be the premium seller of something indistinguishable.

Now the part most people skip, and the part this whole post turns on.

Commodities are manufactured, not found. WTI exists as a tradeable thing because someone wrote a specification: 1,000 barrels, stated gravity and sulfur content, delivered at Cushing, Oklahoma. Brent is a different document. "Grade A large" is a written standard for eggs.

The grading standard is what creates the fungibility. Before it, you inspect every lot. After it, you trade paper. A commodity market is not a fact about goods — it's a legal and institutional achievement, and the document is load-bearing.

Which also means fungibility is never universal. WTI and Brent aren't interchangeable with each other; they're separate grades that trade at a spread. What a commodity market gives you is interchangeability within a defined grade.

Hold onto that. It's the whole argument.

I ran a substitution test without knowing that's what it was

Two posts ago I raced Claude Opus 4.8 against DeepSeek-V4-Pro on 36 office-document tasks — same prompt builder, same grader, same tasks, one field changed. I thought I was benchmarking two models. I was doing something older than that.

Format Opus 4.8 DeepSeek-V4-Pro
Excel 94% (15/16) 94% (15/16)
Word 90% (9/10) 90% (9/10)
PowerPoint 100% (10/10) 70% (7/10)

Read that as an economist rather than an engineer. On Excel and Word, two different suppliers' output passed the same inspection at the same rate. Within that grade of work, the units were interchangeable. On PowerPoint they weren't — which is the same finding stated the other way: that task was a different grade, and those two suppliers were not substitutes for it.

That is a substitution test, and it is the only way anyone has ever established that a good is a commodity. Not by arguing about it. By swapping suppliers and checking whether the output still passes inspection.

Your eval is the grading standard

Here's the thing I didn't see at the time.

If the grading document is what turns a heap of grain into No. 2 yellow corn, then an eval harness is that document for model output. It is the assay. It's the mechanism that converts "some text a language model produced" into a unit with a stated specification, one a buyer can accept without inspecting it personally.

And look at what the harness had to do before it was allowed to grade anything. The reference solver — the known-good answer — has to score 100%, or the standard is rejecting good grain. The no-op solver, which just copies the input to the output, has to score 0%, or the standard is passing empty sacks. Only after both pass is the thing permitted to certify anything.

That's not engineering hygiene. That's assay calibration, and grain exchanges have been doing it since the 1850s.

So the line I put in that post — an eval you haven't validated is a random number generator with a spreadsheet attached — has an economic translation I like better. A market with no grading authority isn't a market. Every trade requires personal inspection, nothing is fungible, and you're left with a pile of bilateral deals where the buyer takes the seller's word for quality.

Which is, precisely, how most teams pick a model. They're buying on reputation because nobody ever wrote the spec.

The container standard

The other half of a commodity market is plumbing that makes switching cheap, and here the parallel is almost too on-the-nose.

The OpenAI-compatible /v1/chat/completions endpoint is the shipping container of this industry. A mediocre interface that won by being everywhere — which is exactly how container standards always win. My gateway takes a URL, a bearer token, and this:

MODEL = "deepseek-v4-pro"

Change that one string and a different company's hardware, in a different country, runs the job. Nothing else in the harness moves.

A market where you change supplier by editing one field and altering nothing else is the definition of the thing. Grain traders spent a century building the institutions to get there. We got it because everyone cloned one company's JSON schema.

The market is splitting in two, on schedule

If the grades are real and switching costs one line, the textbook consequences should show up in the price data. They have, and in a shape sharper than I expected.

According to BenchLM's Token Price Index, as reported here, the median across 130 tracked models as of 24 July 2026 sits at $1.00 per million input tokens and $4.00 per million output. But the median hides the interesting part:

Frontier prices rose 36.4% year over year. Mid-tier prices fell 35.8% over the same period.

Nearly mirror-image moves in opposite directions. That is not a market drifting in one direction — it's a market coming apart into two markets.

And it's exactly what the theory predicts. The frontier is not commoditized. A lab holding the best reasoning model has genuine pricing power, because for the hardest grade of work there is no substitute — that's monopolistic competition with a real quality difference behind it, and prices rise. Everything the frontier has already passed falls into the commodity zone, competes on the cost curve, and deflates.

The commodity forms behind the moving edge. And since the edge keeps moving, the commodity zone keeps expanding upward, swallowing work that used to require the expensive supplier.

My 36-task result is a snapshot of exactly that boundary on one particular afternoon: two formats had already fallen behind the frontier into the commodity zone, one hadn't yet.

Three places the analogy breaks

A metaphor that admits its edges is more useful than one pushed until it snaps. Here's where this one stops working.

1. The priced unit isn't standardized. A barrel is 42 gallons at every supplier on earth. A token is not a token — providers use different tokenizers, so the same English paragraph bills as a different quantity depending on who you send it to. We are pricing in a unit with no cross-supplier definition, which would be intolerable in any mature commodity market. Imagine oil where each producer used their own barrel.

The honest denominator is cost per completed task. The only instrument that measures that is a grader. That's the third time in this post that the harness turns out to be the load-bearing part, and I promise I didn't plan it that way.

2. The price doesn't oscillate, it only falls. Commodity volatility comes from inelastic short-run supply meeting inelastic short-run demand — a copper mine takes seven to ten years to build, so a 2% shortfall can move price 20%. Tokens don't work like that. Nothing is depleted, the cost floor keeps dropping from hardware and algorithmic improvement, and demand is highly elastic: cheaper tokens create new applications rather than just cheaper versions of old ones. The cost structure looks like semiconductors or airlines — enormous sunk fixed cost, near-zero marginal cost — not agriculture.

Which is why there's no token futures market and shouldn't be. You can't store a token, and nobody needs to hedge a price that moves one way. The hedge people actually buy is availability — reserved throughput, provisioned capacity — because the real scarcity is rate limits, not price.

3. Grades don't hold still. No. 2 yellow corn meant the same thing in 1970 as it does today. Token grades slide continuously; this year's frontier is next year's bulk tier.

That has a sharp practical consequence: a substitution test has a shelf life. My July result is true as of July, and the drift runs toward more substitutable, not less. Which doesn't weaken the argument for measuring — it's the argument for measuring. A one-time vendor evaluation is a depreciating asset. The harness is the part that holds its value, because you can re-run it.

There's a fourth, smaller crack worth naming: identical inputs don't reliably produce identical outputs even from a single supplier, thanks to batching and hardware non-determinism. Corn doesn't do that. It's why grading here has to be statistical — three variants per task, all must pass — rather than a single inspection.

Where the money goes when the middle commoditizes

Christensen's law of conservation of attractive profits, from The Innovator's Solution (2003), says that when profits disappear at one stage of a value chain because that stage became modular and commoditized, attractive profits emerge at an adjacent stage. Value doesn't evaporate. It migrates toward whoever controls the current bottleneck.

The model layer is getting squeezed from both directions at once. Upward, value migrates to power, fabs, and datacenter capacity — the genuinely inelastic inputs that nobody can conjure on demand. Downward, it migrates to the harness, the evals, the routing, the proprietary context. The middle is where margin leaves.

I have one number for this, and it's the number I'd keep if I had to throw away everything else in this series.

Going from the cheap supplier to the expensive one bought me 3/10 on PowerPoint. Handing the cheap supplier its own Python traceback and letting it retry once bought me 2/10 — for about forty lines of harness code.

The scarce thing wasn't the tokens. It was never the tokens.

The counter-move, and how to recognize it

If you sell something that's commoditizing, you try to make it non-fungible again, usually by attaching properties the buyer can't verify by inspection and has to take on trust. Coffee does this with single-origin and fair-trade labels: physically the same bean, economically no longer interchangeable, because the certification is now part of the good.

In this market that's compliance certifications, data residency, safety guarantees, proprietary tool ecosystems, fine-tunes.

The sharpest one deserves to be named for what it is. Prompt caching is a switching cost dressed as an optimization. It's sold as a way to cut your bill, and it genuinely is one — I use it. It is also a mechanism that strands your context at one provider, where it doesn't travel. That's a grain elevator you don't own and can't move.

Perfectly legitimate. Genuinely useful. Just know which of the two things you're buying, and price the second one.

This one wants re-checking every six months

Most things I write here are true for years. This one has an expiry date printed on it, and I'd rather say so than have it quietly rot.

Every claim above is a measurement of a moving boundary: how far the commodity zone has crept up behind the frontier, and how wide the spread between the two grades has opened. Both numbers change constantly, and only in one direction. So this is a post I intend to re-run rather than merely re-read, and I'd suggest you do the same with your own version of it.

Twice a year is about the right cadence — often enough to catch a grade crossing over, rare enough that you're not chasing noise. Three things to look at:

  • Re-run the substitution test. Same harness, same tasks, swap the supplier, and see which grades of work have fallen into the interchangeable zone since last time. This is the only one that's actually about your workload, and it's the only one that can tell you to change a routing config.
  • Watch the spread, not the median. The headline "tokens got cheaper" hides the interesting motion. Frontier up 36.4% against mid-tier down 35.8% is two markets separating; a narrowing spread would mean something quite different.
  • Check what's still stuck. My PowerPoint tasks were the holdout, and then most of that gap turned out to be a missing retry loop rather than a missing capability. The residue after you've fixed your harness is the real frontier premium — and it's the only part you should still be paying for.

If the pattern holds, each time you run this the answer moves the same way: more of what you're paying a premium for turns out to be purchasable at commodity prices, and the thing that stays scarce is the apparatus you built to tell the difference.

I'll post the numbers again when I re-run mine.

So

Tokens are becoming a commodity one grade at a time, behind a frontier that keeps moving. The eval is the grading standard that makes substitution legible. The unit we all price in is the wrong unit, and only a grader gives you the right one. And the margin has already left the model layer — upward to the power plant, downward to the harness.

Three posts ago I was going to distill a model. I ended up with a config change, a retry loop, and an argument about corn.

Build the ruler first. It's still the only part that tells you the truth — and it turns out that's not just an engineering opinion. It's what a market needs before it can have prices at all.