thepanelist

LLM personas skew young, urban, and agreeable: here's what we built to fight it

thepanelist team
September 24, 2026
Methodology

Ask an LLM to imagine five people who might buy your product, with no other constraints, and you'll get five variations on the same person: probably in their late 20s or early 30s, probably urban, probably articulate, probably enthusiastic about your idea. This isn't a one-off quirk. It's a documented artifact of what these models were trained on, and it shows up reliably enough that it's worth designing against rather than hoping to prompt around.

The failure has two layers, and they need different fixes. The first is demographic: personas skew young, urban, educated, and higher-income by default. The second, subtler one is what we ended up calling mode collapse: even when personas differ on paper (different names, different ages), their actual opinions converge. Five "different" people who all like your idea, expressed in slightly different words, isn't five opinions. It's one opinion wearing five names.

Prompting alone gets you partway. "Make the personas diverse" produces personas that are diverse in decoration (different jobs, different hobbies) while still converging on the same underlying stance toward whatever you're asking about. The demographic slider moves; the actual disagreement doesn't show up unless you ask for it explicitly and specifically.

So the system prompt behind every panel carries

“

Hard rules, not a vague request for diversity.

No two personas may share the same combination of age range, income tier, and adoption stance. At least one persona in every panel must be actively skeptical or resistant to the product category, not just "has concerns," actually resistant. Each persona needs a genuinely distinct voice, specified as sentence length and vocabulary and tone, not just biographical facts. And critically, personas are allowed to be vague, blunt, or uninterested: the instinct to make every character articulate and engaged is itself part of the bias.

None of this is a one-time fix you apply and forget. It's why the product has a visible diversity check on every panel, not just at generation time: a rough measure of whether the panel actually spans a real distribution, and a separate check for whether the answers to a specific question are suspiciously similar to each other regardless of what the demographic spread looks like. A panel can pass the first check and still fail the second: five demographically distinct personas that all happen to answer identically to the one question you actually asked.

We also added a comparison against a real population baseline, not just "does this panel disagree with itself" but "does this panel's age distribution look anything like the real adult population's." That one is deliberately narrower in scope than it could be: we don't attempt the same comparison for income tier, because our income categories are qualitative labels an LLM assigns, not tied to a real dollar figure, and pretending otherwise would be exactly the kind of overclaiming this whole exercise exists to avoid.

“

The honest summary: this doesn't prove a panel's answers match what real customers would say.

It can't, from internal checks alone. What it does do is make the specific, well-documented failure mode visible instead of silent: you get a number telling you whether this panel is actually diverse, instead of quietly trusting a wall of confident, engaged, remarkably similar-sounding opinions.