What we are learning about modelling people, and why describing them is not enough.
The usual approach
You could start with what you know about them:
The usual approach
“Pretend you are this person.”
That is roughly how many synthetic-human systems work today.
The problem
Knowing facts about someone is not the same as understanding them.
The evidence
Recent research has tested digital twins built from hundreds of answers about personality, preferences, demographics and cognitive traits.
Giving an AI hundreds of personal answers can improve its predictions, but sometimes only slightly compared with giving it basic demographic information.
A better question
What if we don't need more questions? What if we need a different way to model people?
How knowing actually works
You don't understand your best friend because they filled out a personality test. You understand them because you've watched them over time.
The shift
Don't just describe the person. Learn from what they do.
Models become more useful when they learn from a person's history of actions, and the situations in which those actions happened.
Words vs actions
Those contradictions are not noise. They are part of the person.
From personas to people
Others
Us
Instead of asking the LLM to invent the person from a description, we give it a richer model of who that person is and how they have behaved before.
Step 1
Some parts of a person change slowly. We call this the Core.
One person may be more anxious about risk. Another naturally optimistic. Someone else highly sensitive to rejection, or strongly motivated by status. These tendencies belong in the Core because they influence many different situations.
Step 2
Ask three people: “Would you buy this electric car?”
Same question. Different memories. Different answer. A Synthetic User needs to retrieve the right experiences at the right moment.
Step 3
Persistent
The Core. Who this person is normally.
Temporary
The Current State. How they feel right now.
Step 4
Only after we have the person's Core, their relevant memories, their Current State and the current situation do we ask for a response. And we don't ask a single LLM: we reason across several, so the answer isn't shaped by the bias of any one model.
The models are no longer asked to create the person. They are asked to reason from the person.
A caveat
Maybe personality, values and motivations are useful to us, but not the best internal representation for a machine. A model can also learn its own representation from someone's history.
Human-designed
Machine-learned
The hybrid we are testing
At Synthetic Users, we are testing whether the strongest approach is a combination of both.
Testing the architecture
Testing the architecture
The test that matters
Does this change make the synthetic person predict the real person better?
We ask it of every version. If a layer doesn't improve performance, there is no reason to keep it.
A stronger test
It is useful to ask whether a Synthetic User can answer a survey like this person. It is stronger to ask what this person will do next.
Given a person's history and current situation, predict their next action. Synthetic Users is bringing that idea into research simulations.
An open experiment
Research suggests longer histories tend to produce better representations, and even the order in which things happened carries useful information.
The practical question: how much do we actually need to know about someone before we can model them well?
The ceiling
Ask someone the same question two weeks later and they might answer differently. So 100% accuracy is the wrong goal. We measure both:
The model
The ceiling
The real target isn't perfection. It is getting closer to human-level consistency.
Calibration
Those are very different research findings. For researchers, knowing when not to trust a simulation may be nearly as important as improving the simulation itself.
The bigger idea
A person is not a document. A person is a history.
Things happen to us. We react. We remember. We learn. We repeat ourselves. We contradict ourselves. And we change.