Zero is not the same as nothing

A few weeks ago, a client opened an Excel export from our platform and found a column of zeros. Every person in the file had “0” under training completed, including people we knew had completed plenty. The other columns were right. That one wasn’t.

The client caught it. We didn’t. I keep coming back to that.

The right numbers existed. They were in our data, in another table, under another name. The system just went to the wrong place, found nothing there, and wrote down what it found: nothing, as a zero.

What actually happened

Our platform has more than one way to deliver training. This client only uses one of them. The table for the other one exists in their dataset, because every tenant gets the same schema, but it has never had a single row.

The AI assistant that builds these exports gets a short notice before each conversation telling it what data this client has. That notice was built by checking whether a table exists. Not whether it has anything in it. So the model was told, in effect, “this client has training programmes”, went looking for completions, counted an empty table, and got zero. The export step then checked that the column existed, which it did, and wrote the file.

Every step was correct on its own terms. Counting an empty set and getting zero is not a bug. It’s arithmetic. That’s what makes this kind of failure so easy to miss: nothing throws an error, and nothing looks wrong unless you already know the answer.

Why something so basic costs so much

The honest description of this bug is almost embarrassing. We confused “there is nothing here” with “the answer is zero”. It’s the kind of distinction you would explain to someone in their first week of working with data. Tony Hoare called his invention of the null reference a “billion-dollar mistake”, but null at least exists to say “I don’t know”. We had that information and flattened it into a number.

What I didn’t fully appreciate until this happened is how unevenly trust works. A file with ten correct columns and one wrong column is probably not seen as 91% correct. I suspect it’s seen as a file you can’t trust, from a system you now have to check. And the person who has to check it is the HR manager who asked for the export so they wouldn’t have to.

A crash is honest. It says “I couldn’t do this”. A zero is confident. It says “I did this, and here is the answer”. In people data, a confident wrong answer is worse than no answer, because someone might act on it. Zero completions could read as nobody taking the training seriously. That’s a conclusion about people, drawn from a table that was never in use.

The same failure, in the other direction

Around the same time, we were auditing our AI course generator. It takes a client’s document and turns it into a course. A large part of the recent work on it was teaching it to say what it couldn’t do: a chapter it couldn’t match to the source, a list that got lost along the way, fewer slides than planned, a video search that found nothing relevant. Each step of the generation can attach a note saying so.

Then we looked at what reaches the person building the course. The layer between the generator and the interface passes on each step’s name, status, error and duration. Not the note. On one course we checked, the logs showed two warnings written by the generator. The interface showed every step in green.

So the system knew. It did the expensive part: it noticed its own limits and wrote them down. And the signal died one layer before the person who could have acted on it.

Is it the same problem?

I think so, mostly. In both cases the interface claimed a certainty the system didn’t have. A zero and a green tick are both the default you get when nothing says otherwise. And in both cases, nothing said otherwise.

The difference is where the information was lost. In the export, the system never knew it didn’t know: the doubt was never produced. In the course generator, the doubt was produced and then dropped in transit. I’m not sure which is worse. My instinct says the second, because we paid to produce the doubt and then threw it away. But the first is probably more common, because it doesn’t require anyone to have thought about it at all.

If there’s a general lesson, it might be this: an AI system’s honesty is only as good as the narrowest pipe between it and the person who decides. You can make the model careful. If the API contract between two services has no field for “careful”, the care stops there.

What we got wrong while fixing it

Our first fix for the export made a new mistake. We changed the assistant so that, when the table was empty, it would tell the client the module “is not in use by this client”. It sounded right. For this client, it was.

Then I asked the question I should have asked first: what about a new client, who turned the module on last week and simply hasn’t had anyone complete anything yet? For them, “zero completions” was the true answer, and the old behaviour had been right. We were about to tell a client in onboarding that they don’t use a product they had just bought.

The evidence was “this table has no rows”. “The client doesn’t use this module” is an inference on top of it, and an empty table can’t tell those two worlds apart. We had fixed the case that broke by quietly breaking one that hadn’t, without anyone arguing for the trade. The rule we took from it is simple to say and harder to follow: say what the evidence shows, and don’t give a reason for the absence.

What we changed

  • Three states instead of two. A table is now absent, present with data, or present and empty. The assistant is told which, and an empty table gets its own line in the notice rather than being folded into “this client has”.
  • The export checks rows, not columns. An empty table says yes to every “does this column exist?” question. That check was the last gate before the file reached the client, and it couldn’t fail.
  • The new check is allowed to fail quietly. If counting rows fails, we fall back to the old behaviour rather than blocking the export. A safety check that takes the service down is a trade we didn’t want to make.
  • Empty is said out loud. In our dashboards, when an organisation has no data for something, we don’t hide the card. It says the organisation has no data of that kind. It makes no claim about the person, and it doesn’t pretend to be a zero.

The question I’m left with

We spend a lot of time, as an industry, on whether AI gives the right answer. I think we spend much less on whether the system around it can carry the answer “I don’t know” from one end to the other without turning it into a zero or a green tick along the way.

I don’t have a framework for this. What I have is a habit we’re trying to build: whenever a number looks plausible, ask what else would produce the same number. An empty table produces zero. So does a team that never started. So does a module that was switched on yesterday. If you can’t tell those apart, you probably shouldn’t show the number as if you could.

The client who caught our zeros did us a favour. I’d rather the next one doesn’t have to.