When I describe an LLM to people as a stateless semiotic process, “stateless” can sound like a limitation. How could it interpret anything useful without knowing what happened before?
To which I point out: well, more context doesn’t really equal “knowing” anything.
There’s a key distinction there between receiving context and owning state.
Suppose a practice system asks a learner to respond to a difficult workplace conversation (Hi, Saige). The system knows the scene’s objective, what has happened so far, and which choices should change the scene. It can pass the learner’s reply and the relevant context to a language transform, asking for a bounded interpretation: what is the learner trying to communicate here?
On the next turn, the system can make a new call with new context. The model doesn’t need an independent, growing memory of the scene. The application can keep the history it needs, choose what is relevant, and remain responsible for how the scene advances. Most history in conversation condenses down to very, very little critical information as the conversation progresses, and of that information very little of it is relevant in any given exchange.
This is also why “just put everything in the prompt” misses the point of what an LLM can really do. A long prompt may contain facts about a workflow, but how much of that wall of text even matters at a given moment? Just because context windows are comically huge now, it doesn’t make the model the right place to store the workflow’s state or enforce its rules.
Honestly, I think the context size of frontier models says less about “intelligence” and more about just how stable human language actually is.
The useful boundary to what should go where is honestly simple: the system remembers and decides; the transform interprets the language it is given for the task at hand, from one known input class to one known output class. Context can cross that boundary, but only ever deliberately. The moment you’re loading unneeded information into the context window and hoping the model “thinks” its way to the right solution you’re already in trouble.