
Science
Facts, not memories.
A language model wakes into every question with no memory of the last one, so everything that feels like memory is a note someone handed it. What it takes to make those notes trustworthy - and why the one that a model writes is never allowed to say anything new.
In Memento, Leonard Shelby can't make new memories. Every few minutes the slate is wiped and he comes to in the middle of a conversation he doesn't remember starting. So he builds himself a system. Polaroids. Notes in his own handwriting. The things he can't afford to lose go on as tattoos, where he'll always find them.
It works, mostly. He gets a very long way on it.
The joke the film plays on him is that his system has no way to tell a true note from a false one. If it's in his handwriting, it's true. He spends the whole film acting on a lie he wrote himself, with complete conviction.
A Synthetic User is Leonard. That's a description, not a metaphor. A language model has no memory from one question to the next, so everything that looks like memory is really a note that somebody handed it. Which makes the notes the whole problem.
§ 1Why long conversations slip
The feeling of a conversation comes from re-reading the conversation. Every time a Synthetic User is asked something, everything already said gets handed back to it along with the question. Leonard's coat pocket, basically. Not recollection. A record, read again from scratch every time.
There's a limit to how much it can take in at once. So in a long interview something has to give, and what gives is the oldest part of it. That's not a flaw in one model, or one vendor. It's how all of this works.
In practice you got a Synthetic User that was excellent for a while and then stopped being the same person. Quietly. It never announced that it had lost the thread. It just started agreeing with things, or answered in a way that didn't square with what it had told you forty minutes earlier.
§ 2What we changed
Answers don't drop off the end any more. Each one leaves a short note behind, in the Synthetic User's own words, and the notes travel on with the interview. When the notes themselves get long, the oldest of them are condensed into a shorter record, and that travels too.
The notes are the unglamorous half of this and they do almost all of the work. No model writes them. Nothing gets rephrased or interpreted. A note says what was asked and what the answer was, in the words that were used, and that's it. It can't get anything wrong because it isn't saying anything new. For most interviews that's the entire feature.
Follow-ups get it too. That matters more than it sounds like it should. A follow-up is asked once the interview is over, against all of it, so it's the longest thing the system ever has to hold, and it's the place people are most likely to ask about something from the beginning.
§ 3It may repeat. It may not invent.
The condensed record is the only place a model writes anything, which makes it the only place Leonard's problem can happen. A note nobody checks becomes true just by existing, and after that it gets acted on.
So it can shorten what was said and nothing else. No adding, no merging, no interpreting. A price from one answer and a product name from another can't be brought together into a new sentence, because that sentence is something the Synthetic User never said.
Then every line gets checked back against the original answers. Figures have to match. Names have to match. We also look for the small changes that reverse a sentence while looking like they've changed nothing, a dropped "not", a condition that goes missing. Lines that fail get thrown out. We don't try to repair them.
When in doubt, the Synthetic User forgets rather than guesses. Forgetting is an ordinary interview problem and you can see it in the transcript. Confidently remembering something it never said is a different problem, and you can't.
None of this is visible to the Synthetic User, and none of it reaches you. It isn't told to go and consult its notes. It's answering your question with everything it has already said still available to it, which is all anybody means by remembering. The notes are never written into the conversation either, so your transcript, and every report built on it, has only what was actually said in it.
§ 4How we tested it, and what came back
Memory is easy to claim and hard to check. An interview that has forgotten half of itself reads just as well as one that hasn't, and nothing looks wrong. So we built the test to be failable.
We built interviews to order and ran them through the real interview code, with the limit tightened far below anything you'd run under, so most of each one was pushed out of reach. We planted specific things early on, a price, a tool and the year it was adopted, a headcount, and then asked about them at the far end, where the condensed record was the only place they could still be. We also put false claims to the Synthetic User about things it had never said. That's the trap a model walks into most easily, because agreeing is what models do with a confident statement.
Then we ran it twice, once with the memory and once with the behaviour it replaced. Sixty questions about a planted fact, thirty-six false claims, each asked three times over.
- The fact came back 100% of the time. Without the memory, 0%.
- Every false claim was corrected. 100%. Without the memory, 0%.
- 38% of the time, the old behaviour denied having said it. Not "I don't remember", but "I haven't said where the team is based", to an interviewer who had been told exactly that. With the memory, 0%.
A live study run alongside it came out the same way, and settled the check that could most easily have been quietly false: across every answer it saved, not one line of this machinery showed up.
Those conditions are a stress test and not how studies run. Nineteen of every twenty answers were pushed out of reach, which is far past anything real. Most interviews never come near the limit, and we raised it substantially in the same piece of work. This is for the long ones, and for follow-ups.
A note on the examples: they're illustrative, and drawn from our own test studies. No client material appears anywhere in this post.
§ 5Where this leaves us
One boundary is deliberate. Memory covers one interview and any follow-ups to it. Two interviews with the same Synthetic User share nothing, and the second one starts fresh. A Synthetic User carrying memories between studies stops being someone drawn from the audience you described and becomes a character with a history, and then every study after the first is quietly shaped by studies you weren't running. There's no setting for that, because it shouldn't be one.
Leonard's problem was never that he forgot. It was that what he wrote down could be wrong and he had no way of knowing. He had all of the discipline and none of the verification, which is the worse half to be left with.
The notes are the Synthetic User's own words. They're checked against what it actually said. What was said stays said.
If you want the rest of the story, Nobody is the average covers where a Synthetic User gets its evidence and its differences, and Enriching your Synthetic Users covers how your own documents reach one during an interview.