
Science
Enriching your Synthetic Users
Until now, a client's documents shaped who our Synthetic Users were. Now they are also in the room during the interview, consulted question by question, with a record of what grounded every answer.
When a client attaches their own material to a study like a market report, a survey export, a folder of past interview transcripts they are looking for something specific. They want Synthetic Users whose answers are anchored in that material, not in a language model's general impression of the world.
Until now we delivered that in one way. Today we're adding a second, and the two do different jobs.
§ 1How documents already reach a Synthetic User
Before an audience's Synthetic Users are created, we take each attached document, have a model read it end to end, and write a condensed account of it aimed at that audience. Those accounts feed into the prompt that builds the Synthetic Users themselves. The result is people whose circumstances, vocabulary and priorities are shaped by what the client actually knows about their market.
This works, and it stays. Reading a document whole is how you get a Synthetic Users who hangs together, someone with a coherent history rather than a bundle of retrieved facts. Nothing about the new work replaces it, and we'd have built it the same way again.
What it does, though, is shape who a Synthetic Users is. It happens once, before the conversation starts. Which leaves one thing on the table: during the interview itself, a Synthetic Users had no way to reach back into the material for the particular detail that answered the particular question in front of it. It carried a well-formed impression of the documents. It couldn't consult them.
That's the headroom we went for.
§ 2What we added
Every attached document is now also indexed when it's uploaded, broken into passages of a few hundred words. During an interview, each time a Synthetic User is asked something, we find the passages that genuinely bear on that question and fold them into that one answer.
So the two passes now cover both halves of the problem.
Two things follow from that which weren't possible before:
- Answers can be specific. A Synthetic User can ground a reply in the actual figure, the actual finding, the actual lived detail from the passage that matched, rather than in a summary's worth of general impression.
- We can see what grounded what. Every answer records which passages informed it, and whether retrieval succeeded, found nothing relevant, or failed. That last distinction matters more than it sounds: an interview grounded in nothing reads exactly as well as one grounded in everything, so "it looked fine" is not evidence.
A note on the examples that follow: they're illustrative. No client material appears anywhere in this post.
§ 3Six kinds of document, not one
Indexing documents properly raised a question the earlier design never had to answer: what kind of thing is this document, exactly?
Up to now, everything attached to a study was treated as one category: background material, handled uniformly. So before designing anything more elaborate, we asked the people using the platform what they were actually uploading, and what they expected it to do to a Synthetic User. Six recognizably different kinds of document came out of those conversations:
- General background market reports, product specs, category overviews. Anything without a sharper pattern, and still the sensible default.
- A real individual material about one named, real, findable person. Uncommon, but it happens, and it should belong to exactly one synthetic user and no one else.
- Segment Synthetic Users a fictional named character standing in for a type of customer: an invented name, a written bio, one consistent voice. On paper these look identical to case 2, name, bio, voice, but they need the opposite treatment. Pinning a segment archetype to a single Synthetic User would be exactly wrong. It's meant to inform every Synthetic User who resembles that segment.
- Raw transcripts verbatim quotes from real research participants. Genuinely useful for vocabulary and for the shape of what someone in that position tends to say. Genuinely risky if a Synthetic User starts recounting a specific stranger's specific experience as its own memory.
- Quantitative data crosstabs, demographic breakdowns, survey tables. A population statistic is not a diary entry, and "this is something you Synthetic Userlly lived" is the wrong frame for a percentage.
- Generation-only reference material documents whose whole job is to help build a Synthetic Users: recruitment criteria, screening specs, demographic guides for a target audience. These have no business surfacing mid-interview. They're useful before the Synthetic User exists and should be invisible afterwards.
Each of the six now gets its own handling: how much of it any one Synthetic User is exposed to, how heavily it counts when we're looking for relevant passages, what caveat travels with it when it's used, and whether it may appear during an interview at all. Documents we can't confidently classify fall back to general background rather than guessing. So an uncertain call lands on the safe, familiar behaviour instead of an exotic one.
§ 4Giving each Synthetic User its own evidence
Classification settles what a document is. It doesn't settle who gets to see which part of it.
Left alone, every Synthetic User in an audience would retrieve the same passages for the same question, and a panel of twenty would have less to distinguish them than it should. So each Synthetic Users is now assigned its own slice of each document's material, roughly 60% of it, chosen using the Synthetic User's own identity as the seed. Regenerate the same Synthetic User and you get the same slice; compare two Synthetic Users and the slices diverge. Two Synthetic Users can share a document and still end up grounded in different specifics, because they were never handed the same part of it.
Retrieval is also Synthetic User-aware, not only question-aware. Alongside how well a passage matches the question, we weigh how well it matches the Synthetic User being asked. A segment-Synthetic User document surfaces more readily for the Synthetic Users who actually resemble that segment, rather than surfacing identically for everyone. And the distinction from the list above is honoured here too: segment material is shared broadly across the Synthetic Users it fits, while a real individual's material is pinned to one Synthetic User and withheld from the rest.
§ 5What the Synthetic User never knows
Two rules hold throughout.
Retrieved material never announces itself. It's framed as something the Synthetic User knows from its own life, with an explicit instruction never to mention documents, data, files or research: and if for some reason it conflicts with who the Synthetic User is, the Synthetic User wins.
That framing, and the per-document-type caveats that ride alongside it, are now versioned and editable like any of our other prompts. That sounds like housekeeping and isn't: the exact wording that turns a retrieved passage into something a Synthetic User will speak as lived experience is the most load-bearing text in the whole feature, and it should be something we can improve and evaluate deliberately rather than something buried out of reach.
And it never accumulates. Whatever informs one answer is used for that answer and then dropped: it isn't written back into the conversation the Synthetic User carries forward. It shapes one reply and disappears, so a stray detail can't resurface three questions later as an established fact.
§ 6How we tested it
Retrieval is the kind of feature that's easy to demo and hard to validate: an interview grounded in nothing reads just as fluently as one grounded in everything. So during development we built a harness to check it rather than take it on trust.
It runs the same Synthetic Users through the same questions on a real study twice, once with retrieval switched off, once with it on, through the actual interview pipeline rather than a simulation, and compares the two. Three of the things it assesses are judgements: whether answers are genuinely grounded in the uploaded material, whether Synthetic Users actually differ from one another, and whether anyone breaks immersion by mentioning documents or research. The fourth needs no judgement at all, because it's computed straight from the record of what grounded each answer: whether Synthetic Users draw on a spread of the available material, or retrieval quietly converges on the same one or two documents no matter who's asking.
That fourth check earns its keep. The first three can all pass comfortably while every Synthetic Users in a panel leans on the same source, which is precisely the failure we wanted to catch, and the one you can't see by reading transcripts.
Worth being clear about what this is: a development tool, used to validate the work before shipping it. It isn't a monitor running against live interviews.
§ 7What the numbers show
Across the tests we've run, the mechanical checks come out clean. These are facts about the plumbing rather than judgements about quality, which is exactly why they're worth stating first. They're the claims that could have been quietly false.
Every answer that used retrieval drew on a full set of passages, on every turn. Two-thirds of answers combined more than one passage from the same document, which is only possible if the retrieval unit really is the passage rather than the file. Not one citation reached for generation-only material, so the "invisible once the Synthetic User exists" rule held. Every Synthetic User stayed inside its own assigned slice. And every citation carried the Synthetic User-affinity component, meaning Synthetic User-aware ranking is doing something rather than sitting inert.
On quality, the direction is consistent across independent runs. Answers carry noticeably more of the client's own material per unit of text, and Synthetic Users repeat themselves within an interview measurably less, which is the difference between a panel that circles one idea and a panel that gives you twenty angles on it. Those are the two things retrieval was built to move, and they moved.
§ 8Where this leaves us
The part we'd stand behind most is also the least technical: asking the people who use the platform what they were actually attaching, instead of assuming one category covered it. Treating every document alike wasn't a bad design, it was a reasonable simplification, and it did real work for a long time. What those conversations showed is that some of the material clients bring needs handling close to the opposite of what a single category gave it, and that isn't a conclusion we'd have reached by reasoning about it from our side of the product.
Every claim above about what reached a Synthetic User is something we can point at a record for, rather than something we inferred from reading transcripts and feeling good about them. That trail is part of the feature, not scaffolding around it, and it's what lets us keep tuning retrieval against evidence rather than impressions.
The documents a client uploads no longer just shape who turns up to the interview. They're in the room, available to the Synthetic User at the moment the question is asked, and every answer can show where it came from.