The real problem, and exactly what we did about it.
Every claim on this page is checkable in the actual code, not a summary of a whitepaper. Nothing here is marketing copy for a black box.
Ask an LLM to imagine five customers, and you get one person wearing five names.
This isn't a one-off quirk. Left with no constraints, a language model asked to generate consumer personas defaults to a narrow, agreeable slice of the population — probably in their late 20s or early 30s, probably urban, probably enthusiastic about whatever you're proposing.
The failure has two layers. The first is demographic: personas skew young, urban, educated, and higher-income by default. The second, subtler one is mode collapse — even when personas differ on paper, their actual opinions converge.
Explicit hard rules, not the model's default distribution.
No two personas may share the same combination of age range, income tier, and adoption stance.
The model is explicitly told not to default to young, urban, educated, early-adopter personas, and to actively include older, rural or suburban, lower-income, and skeptical or resistant ones.
At least one persona in every panel must be actively skeptical or resistant to the product category described in the brief.
Each persona needs a genuinely distinct voice: vocabulary, sentence length, and tone vary on purpose, not just the biographical facts.
Personas are allowed to be vague, blunt, or uninterested. Not every persona has to be articulate and agreeable.
A panel can pass every diversity check and still fake disagreement.
Demographic diversity and opinion diversity aren't the same thing. Five personas can have different ages, incomes, and a genuine skeptic in the mix, and still give you functionally the same answer to the one question you actually asked, especially an easy or leading one.
The check itself is simple on purpose: after a question is answered, every response gets embedded and compared pairwise using cosine similarity. We average the pairwise scores across the panel and flag anything above roughly 92% as functionally identical, regardless of how differently worded the answers are on the surface.
92% is a judgment call, not a law of nature — set from looking at enough real panel output to find where genuinely converged answers start clustering versus where personas are still recognizably disagreeing even when they land on a similar conclusion.
Both checks, rendered from the real components.
Not a redrawn illustration of what the diversity check looks like — this is the actual DiversityChart and SimilarityGauge component the product itself uses, fed representative example data.
Answers diverge enough to represent real disagreement.
It means this specific panel skewed too far one way and got caught before you acted on it — that's the check working, not the panel breaking.
Above that, answers are similar enough to suspect the panel is echoing itself rather than genuinely disagreeing.
The baseline note compares this panel's age distribution against real US adult population data, not an arbitrary target.
Deliberately narrower than it could be.
Every panel's age distribution is compared against real U.S. adult population estimates (18-24, 25-44, 45-64, 65+ age brackets) using total variation distance — 0 means identical distributions, 1 means completely disjoint.
We deliberately don't run the same comparison for income tier. Our income categories are qualitative labels an LLM assigns, not tied to a real dollar figure or Census bracket, and pretending otherwise would be exactly the kind of overclaiming this whole exercise exists to avoid.
What none of this proves.
None of these checks make a panel's answer more true. They can't, from internal checks alone. A panel can be genuinely diverse, pass every check on this page, and still be wrong about what real people would say — because it's built from a language model's training data, not from your actual customers.
What these checks do is make a specific, well-documented failure mode visible instead of silent: you get a number telling you whether a panel is actually diverse and whether its answers genuinely disagree, instead of quietly trusting a wall of confident, engaged, remarkably similar-sounding opinions.
A directional gut-check before real research, never a substitute for talking to actual customers.
Before you trust a number.
It's a judgment call set by looking at real panel output to find where genuinely converged answers cluster versus where personas still disagree even when they reach a similar conclusion. It's not a law of nature and it'll move if real usage shows it should.
See the checks run on a real panel.
Every panel you generate ships with this same diversity check, visible before you act on it.
Start a free panel