I still remember the first weeks in which we seriously worked with language models. It was exhausting. Not because they were stupid. But because they were so convincing.
The first months
One model recommended a library to me that did not exist. With a version number, with sample code, with an explanation of why it was better than the alternative. I searched for twenty minutes before I understood that all of it was made up.
Another described an API response to me that never comes back that way. A third reported a test as passed that it had never run. Always in the same tone: calm, confident, helpful.
That was the hard part. A person who does not know something hesitates. He says "I think" or "we would have to check". The models back then did not hesitate. They asserted. And when you pushed back, they apologised and made the next assertion.
There were days when I spent more time checking answers than I would have needed to do the thing myself. A few people on the team wanted to stop. Honestly, I almost would have understood.
The internet had the same problem
At some point I noticed that I had been through this before.
When the internet was new, everything was on there too. Correct things, wrong things, invented things, all in the same layout, all phrased with the same confidence. If you had a question back then, you got ten answers and had to find out for yourself which one was right.
We did not abolish the internet because of that. We learned to deal with it. Check sources. Get a second opinion. Know which sites to take seriously and which not. That took years, and it was never perfect. But at some point it was normal.
With the models it is the same, only faster. What took ten years with the internet happens here in two or three.
And so did I
And then there is a third comparison, one I like less.
At the start of my career I was exactly the same. Young, fast, convinced. I claimed things in meetings that I did not know for sure, because showing uncertainty felt weak. I gave estimates that were more like wishes. And I apologised and moved on when they turned out wrong.
What got me out of that was not lectures about honesty. It was people who asked follow-up questions. Who said: Show me. How do you know that? What happens if it is wrong?
That is exactly what we do with the models today. Not because we distrust them. But because asking is the only way that claims turn into knowledge. With people and with machines.
How we deal with it today
A few things have taken hold with us, and they are all unspectacular.
A model may not report anything as done that has not verifiably run. Tests, builds, queries: We look at the result, not at the statement about it.
Claims about facts are checked against something solid. A piece of documentation, a dataset, a second model with a different context. Never against the model itself.
We give the models the option to say "I don't know", and we reward it. An agent that says "I can't find a suitable function" is worth more than one that invents one.
Fact is: the models have become considerably better in the meantime. They assert less, they hesitate more often, they show their sources. The reflex to ask has stayed anyway. Good.
What stays
When someone tells me today that they cannot work with AI because it sometimes talks nonsense, I understand. I have been there.
But I also tell them: The internet talks nonsense too. So do your colleagues sometimes. So do you, more often than you would like. The problem is not that someone is wrong. The problem is when nobody asks.
The models lied in the beginning, and I believed them too often. What helped was not more trust and not less. What helped was asking. Like with the internet. Like with me.
Written by me. The thoughts, the values, the learnings, the mistakes: all mine. Grammar and spelling are corrected by our own twin model, trained on my texts. Sometimes a stumble stays in. That is mine too.