We built a small model. 80 GPU hours until it was usable. Meta published around 30 million GPU hours for an open model with 400 billion parameters. That is a factor of 400,000.
Two numbers
I like those two numbers side by side because they show something you otherwise don't see.
Our model has a clearly bounded task and is good at it. The big model can do almost anything, reasonably well. Both make sense. But they are so far apart that you can't measure them with the same yardstick.
And the frontier models we all use through APIs sit above that again. How far above, nobody says exactly.
The price per token is not the price
Whoever uses AI today sees one number: dollars per million tokens. It's small. And it gets smaller every year.
That's not wrong. But it's like looking at the electricity price and concluding what a power plant costs.
Behind the token price sit the 30 million GPU hours. The data someone collected. The people who aligned and tested the model. The data centres that run it. And the capital betting that it will pay off some day.
What a percentage point costs
With our small model we saw it precisely. The first 80 percent of accuracy came fast. After that, every percentage point was more expensive than the one before. More data, more runs, more evaluation, more people looking.
Now multiply that by 400,000.
A percentage point on a frontier model is not a weekend. That's months, teams, data centres. When a provider says "our new model is three percent better", behind it is an effort that hardly anyone outside those companies can imagine.
That's why I take benchmarks seriously. And that's why I don't take them too seriously. Three percent on a benchmark doesn't automatically mean three percent in my use case.
Why this changes decisions
The feeling for these numbers changes how you plan projects.
First: you stop building everything on the largest model. For 900 categories you don't need a frontier model. A small, well-trained one is enough, is faster, and costs almost nothing to run.
Second: you understand why token prices can't fall forever. Fact is: at some point someone has to pay for the 30 million GPU hours. Today competition subsidises that. It won't stay that way.
Third: you stop saying "the AI will learn that". Learning costs. For us 80 hours. For the big ones 30 million. None of it is free.
You have to have felt it
I don't think you learn this from an article. Not from this one either.
You learn it when you start a training run for the thirtieth time, wait two hours, and the model got 0.4 percent better. Then you understand what factor 400,000 means.
Whoever has never experienced that can still use AI well. But they should know they don't know the price. And leave the decisions that are about price to someone who does.
The token price is the electricity bill. You don't see the power plant behind it. Whoever has built one themselves, on a small scale, suddenly understands the big ones.
Written by me. The thoughts, the values, the learnings, the mistakes: all mine. Grammar and spelling are corrected by our own twin model, trained on my texts. Sometimes a stumble stays in. That is mine too.