In my recent essay about testing LLMs, I asked whether we keep trying to test them at the wrong level, and I ended up in an interesting discussion earlier this week about the specifics.
So part of it, I think, is that the answer isn’t to begin with an LLM, discover that it behaves unpredictably, and build increasingly elaborate machinery around it until the result resembles a product (or an “intelligence”).
Begin with the system you actually want to build. What does it need to do for a person? What behavior state does it need to preserve? Which decisions have rules? How are errors surfaced to the user without exception?
Once those questions have answers, you can ask where the system would benefit from a semiotic transform: taking messy or fuzzy human expression and turning it into another useful representation of meaning. Maybe that’s code. Maybe it’s dialogue. An LLM can do that operation without owning the workflow, the rules, or the durable state.
Imagine a support system receiving: “I think I got billed twice, but one of those charges might be from last month.” A useful transform might identify a possible duplicate charge, the customer’s uncertainty, and the need to compare transactions. That’s an interpretation of language. It isn’t a refund decision, a database update, or a promise to the customer.
The system can hold the account history, decide what evidence a refund requires, and define who may approve one. It may use the transform at one point in that process, then move on. The transform doesn’t need to remember being called, or know how or why its output will be evaluated. The system remembers what matters.
That changes the design question. Instead of asking, “How do we make the AI do this job reliably?” I would ask, “What job are we designing, and is there a specific relationship in language that this tool can help us process?”
If the answer is yes, the model has a place. If the answer is no, why are you even using AI in the final system?