Can You Test an LLM Like Normal Software?
If you’ve even considered using AI for business (or really production software of any kind), you should probably be asking this question.
Because at this point, we have decades upon decades of accumulated knowledge about how to safely and effectively use software in business. And yet we still mess it up more often than anyone would probably like to admit.
So with the current AI boom producing increasingly… interesting decisions about what should and should not be automated, it seems worth asking a very basic question:
How do we know whether any of this actually works?
Or, more specifically:
Can you test an LLM the way you test normal software?
If the rules of headline journalism are anything to go by, you may already know where this is heading.
No.
At least, not in the same way. But that difference has as much to do with how we’re using the technology right now as with the technology itself.
Software kinda doesn’t work without boring inputs producing boring outputs
“Given the same conditions, the software should behave the same way every time.”
Most conventional software relies on that assumption.
Because of this property, software teams can create thousands of automated tests that say, in effect, “The software we shipped on Tuesday still behaves like the software we tested on Monday.”
That’s important even just building something by yourself in your spare time. When you’re building software so large it’s put together by hundreds or thousands of people and too big to review by hand, that’s really important.
That predictability is one of the foundations modern computing is built on. Unfortunately, an LLM does not naturally give you that same guarantee.
What is “passing” when there’s no one correct answer?
Ask a normal software function the same question ten thousand times and, assuming nothing else changes, you generally expect the same answer ten thousand times (barring the odd cosmic particle link to the Mario particle video).
Ask an LLM the same thing ten thousand times and you can produce something different each time… but still linguistically right. Kind of the whole magic of it.
If you ask it to help you brainstorm five different taglines, you probably don’t want the exact same five lines every time. It would stop being useful pretty quickly.
If you ask it to summarize a paragraph, there may be dozens of equally valid summaries. In fact, the contrast between two summaries may be just as interesting as the content of one of the summaries itself.
This is, quite genuinely, one of the things that makes LLMs useful, and yet it also makes conventional software testing much harder.
If your whole system is an LLM (or many LLMs), you can no longer simply ask:
Did the system produce the correct output?
Now you have to ask things like:
Was the output acceptable? Did it misunderstand the request? Would another reasonable person judge this output differently?
And half of those are so open-ended that then you wind up asking:
What even is acceptable output?
This is where conversations about AI often turn toward benchmarks, evaluations, stress testing, and percentages. I mean, who wants to be in a three hour meeting debating the finer points of what’s subjectively “acceptable” in your workflows? Of course, not debating it might be worse, because then your entire workflow is based on what Bob in engineering decided 20 years ago on a Friday right before clocking out.
But here’s the problem. Benchmarks for a non-deterministic system end up taking the form of percentages, and we say, “It gets this fixed test ‘right’ 98% of the time.”
Is that a good score? Well, kind of depends on the context.
Getting a good match for a fuzzy search in an untagged database with no metadata? That’s awesome!
Issuing paychecks correctly 98% of the time? Not so much.
Which means that the important question is not, “How accurate is the AI?” As I mentioned recently, I’m not even sure that’s a meaningful question at all link to previous hallucination article.
The far more important question is:
“What happens when it isn’t?”
The answer is usually “add more software for the AI to use” currently
Suppose we want to use an LLM somewhere important.
Right now, the fad is to just throw an “agent” into the system, give it the needed access and… quickly discover that simply giving the model a prompt and trusting whatever comes back is not particularly effective.
So we start adding things around it. Validation, constraints, logs, escalations. MORE PROMPTS!
Maybe we give up and just route uncertain cases to people. But who’s deciding what’s uncertain? Is it an AI? Now the process has started again to make sure the AI watching the AI isn’t missing things.
Before you say “but the prompts probably just aren’t right!”, let’s talk about that.
There’s an odd tendency to place a huge amount of responsibility onto prompts. Which, I guess if you’re viewing LLMs as fledgling synthetic intelligences, could make sense. Except they’re not really, and dumping information on your “employee” and assuming they’ll figure it out would be terrible management even if they were.
But the more seriously we try to use an LLM, the more conventional software we tend to build around it to get it to do what we want. So if “using an LLM” ends up meaning “building a system around it” then why are we building the system as an afterthought?
Maybe the LLM isn’t supposed to be the whole application
Like I said, we’ve talked about this before link to previous article. But if you slow down a minute and step back, this points toward a different way to think about the technology. And that way dramatically changes the answer to the testing problem.
Maybe the problem is not simply that LLMs are difficult to test.
Maybe we just keep trying to test them at the wrong level.