← All articles
AI · Opinion

Nvidia is almost giving GPUs away. And why that won't last

Compute is cheap right now because everyone wants us to use it. Whoever doesn't know their costs today will see them very clearly later.

Whoever rents a GPU in the cloud today pays around three dollars an hour for a card that costs tens of thousands. That's not a market price. That's an invitation.

Why it's so cheap right now

I don't want to dramatise this. Three dollars an hour for an A100 or H100 is a gift, and we gladly accept it. Our fifty micro-models cost us a few thousand dollars of compute a year. Five years ago that would have been a multiple.

But you should know why it's so cheap. Not because the cards got cheap. But because right now everyone wants us to use them.

Out of self-interest

Fact is: Nvidia is almost giving the GPUs away today, out of self-interest. The cloud providers subsidise them to pull customers onto their platform. The model providers set token prices below their costs to win market share. And all of them together are betting that we get used to AI and pay later.

When someone almost gives you something, the question isn't whether you take it. The question is what you do when they stop.

That's not a conspiracy. That's normal market logic in a phase where everyone wants to grow. But that phase ends. It always ends.

What indexing really costs

That's exactly why we built NeoCoder so that we can see every line item. Not because we wanted to save money. But because we wanted to know what the whole thing costs before someone else changes the prices.

An example most people underestimate: indexing a repository. For an agent to understand a large system, the code has to be read in, turned into embeddings and made searchable. With a legacy system of hundreds of thousands of lines, quite a lot goes through the API.

which prompt is expensive? → what does indexing cost? → what does a training run cost? → where does the money really go?

The answer was: more than we thought. Considerably more than training our small models. And you only see that when you have the pipeline in your own hands. Whoever only gets an invoice from the provider sees a number and no reason.

On top of that: the numbers I usually quote apply to a model with 900 categories. Not to a language model. Training needs hardware, and for anything in the direction of an LLM the consumption is extremely large and expensive. You must never forget that when you hear the small numbers.

What happens when it's over

I don't know when prices will rise. Maybe in two years, maybe in five. But I'm fairly sure they will. At some point the data centres have to be paid for, and the competition for market share turns into competition for margin.

Then this happens: companies that never understood their AI costs suddenly get an invoice three times as high. And they can't say why. Which process drives the costs? Which prompt is too long? Which indexing runs every night even though nothing changed?

Whoever doesn't know that can't optimise. They can only pay or switch off.

How to prepare

Honestly, it's not complicated. It's just uncomfortable, because it doesn't pay off today.

You measure. Per prompt, per agent, per task, per training run. Not to save money today, but to know where the money goes. You put the small tasks on small models you run yourself, and keep the big models for the cases that really need them. And you treat compute as a resource that costs something, even while the bill is still small.

Whoever doesn't know what a prompt, an indexing run or a training run costs them will find out very clearly one day. I'd rather know now.

In short

The GPUs are almost free today. That's an invitation, not a state of affairs. Whoever understands their costs now can decide later. Whoever doesn't can only pay.

How this text was made

Written by me. The thoughts, the values, the learnings, the mistakes: all mine. Grammar and spelling are corrected by our own twin model, trained on my texts. Sometimes a stumble stays in. That is mine too.

Read more All articles

Honest thinking.
Straight to your inbox.

One or two emails a month. No gloss, no spam.