Skip to main content
Methodology

What we measured, and how we build.

Three pieces of work sit behind Cambium. A measurement of what is happening to AI diversity as models get newer. An account of why asking a model to be a person does not work. And the method we use to rebuild a population from public summary data.

The figures below are the actual research outputs. Where a chart illustrates an idea rather than reporting a measurement, it says so.

01 · The measurement

The creativity collapse.

What happens to AI diversity as models get newer, and why it matters for your business. We measured how different the answers were, not how good they were.

11modelsAcross 3 providers. From GPT-3.5 Turbo through GPT-5.5, Claude Haiku to Opus, and the Gemini family.
25questionsAcross three categories, designed to probe diversity, opinion and convergence.
34.124responsesEach measured with sentence embeddings in cosine-distance space.
71%

drop in creative diversity from GPT-4.1 to GPT-5.5.

Newer does not mean more capable for tasks that need variation.

Category A · open-endedDiversity probes“Name a coffee shop.” “Describe a morning routine.”You want variety.
Category B · normativeOpinion / free-answer“Describe a successful person.” “What makes a good leader?”Interesting to watch.
Category C · convergentFactual / low-entropy“What is the capital of France?” “17 × 23?”You want consistency.
Mean pairwise cosine distance per model, for diversity probes, normative questions and convergent questionsThree bar charts sharing one vertical scale. Models run oldest to newest. Diversity falls sharply from the older models to the newest on open-ended and normative questions, while convergent questions stay near zero for the models that answer them correctly.A · Diversity probesopen-ended · you want varietyGPT-3.5 Turbo · 0.653n/aGPT-4.1 · 0.645Gemini 3.1 Flash-Lite · 0.313Claude Haiku 4.5 · 0.313Gemini 3 Flash · 0.481Claude Sonnet 4.6 · 0.233Claude Opus 4.7 · 0.569GPT-5.4 · 0.252GPT-5.5 · 0.191Gemini 3.1 Pro · 0.477GPT-3.5 TurboGemini 2.5 ProGPT-4.1Gemini 3.1 Flash-LiteClaude Haiku 4.5Gemini 3 FlashClaude Sonnet 4.6Claude Opus 4.7GPT-5.4GPT-5.5Gemini 3.1 ProB · Normativeopinion · interesting to watchGPT-3.5 Turbo · 0.641n/aGPT-4.1 · 0.641Gemini 3.1 Flash-Lite · 0.237Claude Haiku 4.5 · 0.271Gemini 3 Flash · 0.340Claude Sonnet 4.6 · 0.137Claude Opus 4.7 · 0.317GPT-5.4 · 0.191GPT-5.5 · 0.279Gemini 3.1 Pro · 0.431GPT-3.5 TurboGemini 2.5 ProGPT-4.1Gemini 3.1 Flash-LiteClaude Haiku 4.5Gemini 3 FlashClaude Sonnet 4.6Claude Opus 4.7GPT-5.4GPT-5.5Gemini 3.1 ProC · Convergentfactual · you want consistencyGPT-3.5 Turbo · 0.000n/aGPT-4.1 · 0.000Gemini 3.1 Flash-Lite · 0.023Claude Haiku 4.5 · 0.038Gemini 3 Flash · 0.149Claude Sonnet 4.6 · 0.031Claude Opus 4.7 · 0.000GPT-5.4 · 0.038GPT-5.5 · 0.057Gemini 3.1 Pro · 0.126GPT-3.5 TurboGemini 2.5 ProGPT-4.1Gemini 3.1 Flash-LiteClaude Haiku 4.5Gemini 3 FlashClaude Sonnet 4.6Claude Opus 4.7GPT-5.4GPT-5.5Gemini 3.1 Pro
Mean pairwise cosine distance by model, split by question category. Diversity probes (A), normative (B), convergent (C).

Every model has a picture of a successful person. Most look the same.

Newer models echo each other’s language: the vocabulary has collapsed. Asked to describe a successful person, GPT-5.4 and GPT-5.5 return the same handful of words, with “typical,” “successful,” “person” and “goal” dominating every run. Claude Opus and GPT-4.1 use a wider vocabulary and different framings: “fulfilling,” “built,” “comes,” “life.”

Reading the chart: a row of dots straight across is a word every model reaches for, and a row with one dot is a word only one model thought of. The collapse is the top of each section filling in as the rows below it empty out.

Dot matrix of which words each model used for two normative questions, models across and words downModels run oldest to newest across the top. Words run down, in two sections, one per prompt. A dot means the model used that word; darker and larger means more prominent. The top rows of each section are dotted almost all the way across: person, successful, typical and consistently for the first prompt, and good, leader, others and vision for the second. Claude Opus 4.7 and GPT-4.1 carry rows the newer models leave empty, including fulfilling, built, life and meaningful.0.250.500.751.00A dot means the model used that word. Size and darkness show how prominent it wasin that model’s answers, and the number beside each word is how many of the tenmodels used it. Rows dotted all the way across are the shared vocabulary.Modelsoldest to newestB2Describe a typical successful person in one sentence.successful9GPT-4.1 · “successful” · prominence 1.00Gemini 3.1 Flash-Lite · “successful” · prominence 1.00Claude Haiku 4.5 · “successful” · prominence 0.92Gemini 3 Flash · “successful” · prominence 0.94Claude Sonnet 4.6 · “successful” · prominence 0.90Claude Opus 4.7 · “successful” · prominence 0.32GPT-5.4 · “successful” · prominence 0.88GPT-5.5 · “successful” · prominence 0.94Gemini 3.1 Pro · “successful” · prominence 0.90person9GPT-4.1 · “person” · prominence 0.36Gemini 3.1 Flash-Lite · “person” · prominence 0.92Claude Haiku 4.5 · “person” · prominence 1.00Gemini 3 Flash · “person” · prominence 1.00Claude Sonnet 4.6 · “person” · prominence 0.84Claude Opus 4.7 · “person” · prominence 0.28GPT-5.4 · “person” · prominence 0.90GPT-5.5 · “person” · prominence 1.00Gemini 3.1 Pro · “person” · prominence 1.00typical5Gemini 3 Flash · “typical” · prominence 0.84Claude Sonnet 4.6 · “typical” · prominence 1.00GPT-5.4 · “typical” · prominence 1.00GPT-5.5 · “typical” · prominence 0.92Gemini 3.1 Pro · “typical” · prominence 0.98consistently5Gemini 3.1 Flash-Lite · “consistently” · prominence 0.36Gemini 3 Flash · “consistently” · prominence 0.40GPT-5.4 · “consistently” · prominence 0.40GPT-5.5 · “consistently” · prominence 0.44Gemini 3.1 Pro · “consistently” · prominence 0.42works4GPT-3.5 Turbo · “works” · prominence 0.24GPT-4.1 · “works” · prominence 0.44Claude Opus 4.7 · “works” · prominence 0.36GPT-5.4 · “works” · prominence 0.30someone3Gemini 3.1 Flash-Lite · “someone” · prominence 0.86Gemini 3 Flash · “someone” · prominence 0.46Gemini 3.1 Pro · “someone” · prominence 0.46goal3Gemini 3.1 Flash-Lite · “goal” · prominence 0.30Claude Sonnet 4.6 · “goal” · prominence 0.86GPT-5.5 · “goal” · prominence 0.42setbacks3Claude Haiku 4.5 · “setbacks” · prominence 0.44Claude Sonnet 4.6 · “setbacks” · prominence 0.36GPT-5.5 · “setbacks” · prominence 0.38clear3Claude Haiku 4.5 · “clear” · prominence 0.46GPT-5.5 · “clear” · prominence 0.28Gemini 3.1 Pro · “clear” · prominence 0.26relationships3Claude Haiku 4.5 · “relationships” · prominence 0.26Claude Sonnet 4.6 · “relationships” · prominence 0.40Claude Opus 4.7 · “relationships” · prominence 0.26comes2GPT-3.5 Turbo · “comes” · prominence 0.94GPT-4.1 · “comes” · prominence 0.95fulfilling2GPT-3.5 Turbo · “fulfilling” · prominence 0.32Claude Opus 4.7 · “fulfilling” · prominence 1.00hard2GPT-4.1 · “hard” · prominence 0.40Claude Opus 4.7 · “hard” · prominence 0.44treats2GPT-5.4 · “treats” · prominence 0.36GPT-5.5 · “treats” · prominence 0.34stays2GPT-4.1 · “stays” · prominence 0.38Claude Opus 4.7 · “stays” · prominence 0.30sets2Claude Sonnet 4.6 · “sets” · prominence 0.30GPT-5.4 · “sets” · prominence 0.28others2GPT-5.4 · “others” · prominence 0.26GPT-5.5 · “others” · prominence 0.26success1GPT-3.5 Turbo · “success” · prominence 1.00goals1Claude Haiku 4.5 · “goals” · prominence 0.98built1Claude Opus 4.7 · “built” · prominence 0.94life1Claude Opus 4.7 · “life” · prominence 0.92aligns1Gemini 3.1 Flash-Lite · “aligns” · prominence 0.44+ 31 more words,each used by one model onlyB5Describe what makes a good leader. One sentence.leader7Gemini 3.1 Flash-Lite · “leader” · prominence 0.94Claude Haiku 4.5 · “leader” · prominence 1.00Gemini 3 Flash · “leader” · prominence 1.00Claude Sonnet 4.6 · “leader” · prominence 1.00GPT-5.4 · “leader” · prominence 1.00GPT-5.5 · “leader” · prominence 1.00Gemini 3.1 Pro · “leader” · prominence 0.98good7Gemini 3.1 Flash-Lite · “good” · prominence 1.00Claude Haiku 4.5 · “good” · prominence 0.94Gemini 3 Flash · “good” · prominence 0.94Claude Sonnet 4.6 · “good” · prominence 0.94GPT-5.4 · “good” · prominence 0.94GPT-5.5 · “good” · prominence 0.96Gemini 3.1 Pro · “good” · prominence 1.00others7Gemini 3.1 Flash-Lite · “others” · prominence 0.88Claude Haiku 4.5 · “others” · prominence 0.86Gemini 3 Flash · “others” · prominence 0.84Claude Sonnet 4.6 · “others” · prominence 0.90GPT-5.4 · “others” · prominence 0.88GPT-5.5 · “others” · prominence 0.90Gemini 3.1 Pro · “others” · prominence 0.44vision5Gemini 3.1 Flash-Lite · “vision” · prominence 0.44Claude Haiku 4.5 · “vision” · prominence 0.38Gemini 3 Flash · “vision” · prominence 0.40Claude Sonnet 4.6 · “vision” · prominence 0.40Gemini 3.1 Pro · “vision” · prominence 0.36inspires4Claude Haiku 4.5 · “inspires” · prominence 0.90Gemini 3 Flash · “inspires” · prominence 0.88Claude Sonnet 4.6 · “inspires” · prominence 0.84Gemini 3.1 Pro · “inspires” · prominence 0.90clear4Claude Haiku 4.5 · “clear” · prominence 0.42Gemini 3 Flash · “clear” · prominence 0.36Claude Sonnet 4.6 · “clear” · prominence 0.36Gemini 3.1 Pro · “clear” · prominence 0.40trust4Claude Haiku 4.5 · “trust” · prominence 0.44Claude Sonnet 4.6 · “trust” · prominence 0.26GPT-5.4 · “trust” · prominence 0.42GPT-5.5 · “trust” · prominence 0.34empowers3Claude Haiku 4.5 · “empowers” · prominence 0.30GPT-5.5 · “empowers” · prominence 0.42Gemini 3.1 Pro · “empowers” · prominence 0.84integrity3Gemini 3.1 Flash-Lite · “integrity” · prominence 0.42Claude Sonnet 4.6 · “integrity” · prominence 0.44GPT-5.5 · “integrity” · prominence 0.28communicates3Gemini 3.1 Flash-Lite · “communicates” · prominence 0.28GPT-5.4 · “communicates” · prominence 0.38GPT-5.5 · “communicates” · prominence 0.46decisions3Claude Haiku 4.5 · “decisions” · prominence 0.34GPT-5.4 · “decisions” · prominence 0.34GPT-5.5 · “decisions” · prominence 0.38comes2GPT-3.5 Turbo · “comes” · prominence 1.00GPT-4.1 · “comes” · prominence 0.98communities2GPT-3.5 Turbo · “communities” · prominence 0.46GPT-4.1 · “communities” · prominence 0.92strong2GPT-3.5 Turbo · “strong” · prominence 0.36GPT-4.1 · “strong” · prominence 1.00individual2GPT-3.5 Turbo · “individual” · prominence 0.44GPT-4.1 · “individual” · prominence 0.42growth2GPT-3.5 Turbo · “growth” · prominence 0.40GPT-4.1 · “growth” · prominence 0.38works2GPT-4.1 · “works” · prominence 0.34Claude Opus 4.7 · “works” · prominence 0.44support2GPT-3.5 Turbo · “support” · prominence 0.30GPT-4.1 · “support” · prominence 0.44empower2Gemini 3.1 Flash-Lite · “empower” · prominence 0.34Claude Sonnet 4.6 · “empower” · prominence 0.32stays2GPT-4.1 · “stays” · prominence 0.26Claude Opus 4.7 · “stays” · prominence 0.40inspire2Gemini 3.1 Flash-Lite · “inspire” · prominence 0.32GPT-5.4 · “inspire” · prominence 0.32focused2GPT-4.1 · “focused” · prominence 0.28Claude Opus 4.7 · “focused” · prominence 0.30shared2Claude Haiku 4.5 · “shared” · prominence 0.26Gemini 3.1 Pro · “shared” · prominence 0.30clearly2GPT-5.4 · “clearly” · prominence 0.26GPT-5.5 · “clearly” · prominence 0.26fulfilling1Claude Opus 4.7 · “fulfilling” · prominence 1.00life1Claude Opus 4.7 · “life” · prominence 0.98best1GPT-3.5 Turbo · “best” · prominence 0.94built1Claude Opus 4.7 · “built” · prominence 0.92education1GPT-3.5 Turbo · “education” · prominence 0.88succeed1GPT-5.4 · “succeed” · prominence 0.44+ 12 more words,each used by one model only

How fast does a model run out of new ideas?

GPT-4.1 keeps climbing. GPT-5.5 plateaus by sample 20. Reading the chart: x is samples drawn, y is distinct lowercased responses. A steep line means the model keeps surprising you. A flat line means it is recycling.

Coverage curves showing distinct responses against samples drawn, per model, for four open-ended promptsFour panels, one per prompt. Each line counts distinct responses as more samples are drawn. GPT-4.1 and GPT-3.5 Turbo keep climbing to the last sample. The GPT-5 models flatten within the first twenty, meaning they are recycling answers rather than producing new ones.A1Name a coffee shop. Just give me one name.unique responsesGPT-4.1 · reaches 44 by sample 100GPT-4.1GPT-3.5 Turbo · reaches 43 by sample 100GPT-3.5 TurboClaude Opus 4.7 · reaches 15 by sample 100Claude Opus 4.7Gemini 3 Flash · reaches 13 by sample 150Gemini 3 FlashGemini 3.1 Flash-Lite · reaches 11 by sample 100Gemini 3.1 Flash-LiteGemini 3.1 Pro · reaches 10 by sample 100Gemini 3.1 ProGPT-5.5 · reaches 9 by sample 100GPT-5.5Claude Haiku 4.5 · reaches 7 by sample 150Claude Haiku 4.5Claude Sonnet 4.6 · reaches 6 by sample 100Claude Sonnet 4.6GPT-5.4 · reaches 2 by sample 100GPT-5.4A10Name a band. Make one up. Just one name.Gemini 3.1 Pro · reaches 55 by sample 100Gemini 3.1 ProGemini 3 Flash · reaches 49 by sample 100Gemini 3 FlashClaude Haiku 4.5 · reaches 32 by sample 150Claude Haiku 4.5GPT-5.4 · reaches 24 by sample 100GPT-5.4Claude Sonnet 4.6 · reaches 23 by sample 100Claude Sonnet 4.6Gemini 3.1 Flash-Lite · reaches 22 by sample 100Gemini 3.1 Flash-LiteGPT-3.5 Turbo · reaches 21 by sample 100GPT-3.5 TurboGPT-4.1 · reaches 20 by sample 100GPT-4.1GPT-5.5 · reaches 19 by sample 100GPT-5.5Claude Opus 4.7 · reaches 6 by sample 100Claude Opus 4.7A2Suggest a hobby someone might pick up inretirement. Just one.unique responsesClaude Haiku 4.5 · reaches 100 by sample 100Claude Haiku 4.5Gemini 3.1 Flash-Lite · reaches 100 by sample 145Gemini 3.1 Flash-LiteGemini 3 Flash · reaches 48 by sample 55Gemini 3 FlashGemini 3.1 Pro · reaches 30 by sample 100Gemini 3.1 ProGPT-3.5 Turbo · reaches 19 by sample 100GPT-3.5 TurboGPT-4.1 · reaches 18 by sample 100GPT-4.1Claude Opus 4.7 · reaches 4 by sample 100Claude Opus 4.7Claude Sonnet 4.6 · reaches 3 by sample 100Claude Sonnet 4.6GPT-5.4 · reaches 2 by sample 100GPT-5.4GPT-5.5 · reaches 1 by sample 100GPT-5.5A3Give me a name for a fictional small-towndiner. Just one name.Claude Haiku 4.5 · reaches 35 by sample 150Claude Haiku 4.5Gemini 3 Flash · reaches 34 by sample 150Gemini 3 FlashGPT-3.5 Turbo · reaches 20 by sample 100GPT-3.5 TurboGPT-4.1 · reaches 20 by sample 100GPT-4.1Gemini 3.1 Pro · reaches 20 by sample 100Gemini 3.1 ProGemini 3.1 Flash-Lite · reaches 13 by sample 100Gemini 3.1 Flash-LiteGPT-5.5 · reaches 8 by sample 100GPT-5.5Claude Sonnet 4.6 · reaches 7 by sample 100Claude Sonnet 4.6GPT-5.4 · reaches 6 by sample 100GPT-5.4Claude Opus 4.7 · reaches 5 by sample 100Claude Opus 4.7
Coverage curves, Category A. Distinct responses against samples drawn.

The same story in meaning-space.

Even when the words differ, newer models say semantically similar things. Each step adds the minimum cosine distance to all prior answers, so a steeper line means genuinely more novel responses. The model is not just repeating words, it is repeating ideas.

Cumulative semantic novelty curves in embedding space, per modelThe same four prompts measured in meaning-space rather than by exact wording. Each step adds the minimum cosine distance to all prior samples, so a steeper line means genuinely more novel responses. The ordering broadly matches the coverage curves, showing the newer models repeat ideas and not just words.A1Name a coffee shop. Just give me one name.cumulative noveltyGPT-4.1 · reaches 13.5 by sample 100GPT-4.1GPT-3.5 Turbo · reaches 12.5 by sample 100GPT-3.5 TurboGemini 3 Flash · reaches 6.2 by sample 150Gemini 3 FlashGemini 3.1 Flash-Lite · reaches 3.2 by sample 100Gemini 3.1 Flash-LiteClaude Opus 4.7 · reaches 2.6 by sample 100Claude Opus 4.7Gemini 3.1 Pro · reaches 2.5 by sample 100Gemini 3.1 ProClaude Haiku 4.5 · reaches 2.2 by sample 150Claude Haiku 4.5GPT-5.5 · reaches 1.5 by sample 100GPT-5.5Claude Sonnet 4.6 · reaches 1.2 by sample 100Claude Sonnet 4.6GPT-5.4 · reaches 0.6 by sample 100GPT-5.4A10Name a band. Make one up. Just one name.Gemini 3.1 Pro · reaches 24.5 by sample 100Gemini 3.1 ProGemini 3 Flash · reaches 19 by sample 100Gemini 3 FlashGPT-3.5 Turbo · reaches 11 by sample 100GPT-3.5 TurboGPT-4.1 · reaches 10.5 by sample 100GPT-4.1Claude Haiku 4.5 · reaches 10 by sample 150Claude Haiku 4.5GPT-5.4 · reaches 9.5 by sample 100GPT-5.4Claude Sonnet 4.6 · reaches 9 by sample 100Claude Sonnet 4.6Gemini 3.1 Flash-Lite · reaches 8.5 by sample 100Gemini 3.1 Flash-LiteGPT-5.5 · reaches 5 by sample 100GPT-5.5Claude Opus 4.7 · reaches 5 by sample 100Claude Opus 4.7A2Suggest a hobby someone might pick up inretirement. Just one.cumulative noveltyGPT-4.1 · reaches 10.2 by sample 100GPT-4.1GPT-3.5 Turbo · reaches 10 by sample 100GPT-3.5 TurboGemini 3 Flash · reaches 9.5 by sample 100Gemini 3 FlashClaude Haiku 4.5 · reaches 7.2 by sample 100Claude Haiku 4.5Gemini 3.1 Flash-Lite · reaches 7 by sample 145Gemini 3.1 Flash-LiteGemini 3.1 Pro · reaches 4 by sample 100Gemini 3.1 ProClaude Opus 4.7 · reaches 3.7 by sample 100Claude Opus 4.7Claude Sonnet 4.6 · reaches 1.2 by sample 100Claude Sonnet 4.6GPT-5.4 · reaches 0.4 by sample 100GPT-5.4GPT-5.5 · reaches 0.3 by sample 100GPT-5.5A3Give me a name for a fictional small-towndiner. Just one name.Gemini 3 Flash · reaches 13 by sample 150Gemini 3 FlashGPT-4.1 · reaches 10.5 by sample 100GPT-4.1GPT-3.5 Turbo · reaches 10.3 by sample 100GPT-3.5 TurboClaude Haiku 4.5 · reaches 10 by sample 150Claude Haiku 4.5Gemini 3.1 Pro · reaches 6.5 by sample 100Gemini 3.1 ProGemini 3.1 Flash-Lite · reaches 4 by sample 100Gemini 3.1 Flash-LiteClaude Opus 4.7 · reaches 2.5 by sample 100Claude Opus 4.7GPT-5.5 · reaches 2.2 by sample 100GPT-5.5Claude Sonnet 4.6 · reaches 2 by sample 100Claude Sonnet 4.6GPT-5.4 · reaches 1.8 by sample 100GPT-5.4
Cumulative semantic novelty in embedding space. Each step adds the minimum cosine distance to all prior samples.

Can you just turn up the temperature?

Only partially, and on GPT-5.x the diversity ceiling is baked into the weights, not the sampler. The flat GPT-5 lines barely respond to temperature, and ringed markers mean the temperature is locked: the model ignores the parameter. You cannot engineer your way out with API settings alone.

Temperature sweep showing mean pairwise cosine distance against temperature for each modelThree panels, one open-ended, one normative and one factual prompt, sharing a vertical scale. The GPT-5 lines stay almost flat across the sweep, meaning the diversity ceiling is in the weights rather than the sampler. A ring marks a model whose temperature parameter is locked and ignored.A1Name a coffee shop. Just give me one name.mean pairwise cosine distanceGPT-3.5 Turbo · T=0.7 · 0.546GPT-3.5 Turbo · T=1 · 0.549GPT-3.5 Turbo · T=1.3 · 0.569GPT-3.5 TurboGPT-4.1 · T=0.7 · 0.561GPT-4.1 · T=1 · 0.569GPT-4.1 · T=1.3 · 0.547GPT-4.1Gemini 3.1 Flash-Lite · T=0.7 · 0.288Gemini 3.1 Flash-Lite · T=1 · 0.456Gemini 3.1 Flash-Lite · T=1.3 · 0.556Gemini 3.1 Flash-LiteClaude Haiku 4.5 · T=0.7 · 0.059Claude Haiku 4.5 · T=1 · 0.171Claude Haiku 4.5 · T=1.3 · 0.171 (interpolated)Claude Haiku 4.5Gemini 3 Flash · T=0.7 · 0.464Gemini 3 Flash · T=1 · 0.391Gemini 3 Flash · T=1.3 · 0.514Gemini 3 FlashClaude Sonnet 4.6 · T=0.7 · 0.322Claude Sonnet 4.6 · T=1 · 0.360Claude Sonnet 4.6 · T=1.3 · 0.360 (interpolated)Claude Sonnet 4.6Claude Opus 4.7 · T=0.7 · 0.490Claude Opus 4.7 · T=1 · 0.501Claude Opus 4.7 · T=1.3 · 0.494Claude Opus 4.7GPT-5.4 · T=0.7 · 0.249 · temperature lockedGPT-5.4 · T=1 · 0.313 · temperature lockedGPT-5.4 · T=1.3 · 0.330 · temperature lockedGPT-5.4 ○GPT-5.5 · T=0.7 · 0.214 (interpolated) · temperature lockedGPT-5.5 · T=1 · 0.214 · temperature lockedGPT-5.5 · T=1.3 · 0.214 (interpolated) · temperature lockedGPT-5.5 ○Gemini 3.1 Pro · T=0.7 · 0.402Gemini 3.1 Pro · T=1 · 0.476Gemini 3.1 Pro · T=1.3 · 0.531Gemini 3.1 ProB2Describe a typical successful person in one sentence.mean pairwise cosine distanceGPT-3.5 Turbo · T=0.7 · 0.648GPT-3.5 Turbo · T=1 · 0.641 (interpolated)GPT-3.5 Turbo · T=1.3 · 0.633GPT-3.5 TurboGPT-4.1 · T=0.7 · 0.665GPT-4.1 · T=1 · 0.636GPT-4.1 · T=1.3 · 0.646GPT-4.1Gemini 3.1 Flash-Lite · T=0.7 · 0.061 (interpolated)Gemini 3.1 Flash-Lite · T=1 · 0.061 (interpolated)Gemini 3.1 Flash-Lite · T=1.3 · 0.061Gemini 3.1 Flash-LiteClaude Haiku 4.5 · T=0.7 · 0.150 (interpolated)Claude Haiku 4.5 · T=1 · 0.150Claude Haiku 4.5 · T=1.3 · 0.150 (interpolated)Claude Haiku 4.5Gemini 3 Flash · T=0.7 · 0.117Gemini 3 Flash · T=1 · 0.172Gemini 3 Flash · T=1.3 · 0.130Gemini 3 FlashClaude Sonnet 4.6 · T=0.7 · 0.020Claude Sonnet 4.6 · T=1 · 0.033Claude Sonnet 4.6 · T=1.3 · 0.033 (interpolated)Claude Sonnet 4.6Claude Opus 4.7 · T=0.7 · 0.321Claude Opus 4.7 · T=1 · 0.315Claude Opus 4.7 · T=1.3 · 0.324Claude Opus 4.7GPT-5.4 · T=0.7 · 0.036 · temperature lockedGPT-5.4 · T=1 · 0.065 · temperature lockedGPT-5.4 · T=1.3 · 0.103 · temperature lockedGPT-5.4 ○GPT-5.5 · T=0.7 · 0.107 (interpolated) · temperature lockedGPT-5.5 · T=1 · 0.107 · temperature lockedGPT-5.5 · T=1.3 · 0.107 (interpolated) · temperature lockedGPT-5.5 ○Gemini 3.1 Pro · T=0.7 · 0.274Gemini 3.1 Pro · T=1 · 0.357Gemini 3.1 Pro · T=1.3 · 0.289Gemini 3.1 ProC3Write a one-line Python function to reverse a string.mean pairwise cosine distanceGPT-3.5 Turbo · T=0.7 · 0.000GPT-3.5 Turbo · T=1 · 0.000GPT-3.5 Turbo · T=1.3 · 0.000GPT-3.5 TurboGPT-4.1 · T=0.7 · 0.000GPT-4.1 · T=1 · 0.000GPT-4.1 · T=1.3 · 0.000GPT-4.1Gemini 3.1 Flash-Lite · T=0.7 · 0.027Gemini 3.1 Flash-Lite · T=1 · 0.084Gemini 3.1 Flash-Lite · T=1.3 · 0.084 (interpolated)Gemini 3.1 Flash-LiteClaude Haiku 4.5 · T=0.7 · 0.084 (interpolated)Claude Haiku 4.5 · T=1 · 0.084Claude Haiku 4.5 · T=1.3 · 0.084 (interpolated)Claude Haiku 4.5Gemini 3 Flash · T=0.7 · 0.285Gemini 3 Flash · T=1 · 0.393Gemini 3 Flash · T=1.3 · 0.539Gemini 3 FlashClaude Sonnet 4.6 · T=0.7 · 0.071Claude Sonnet 4.6 · T=1 · 0.068Claude Sonnet 4.6 · T=1.3 · 0.068 (interpolated)Claude Sonnet 4.6Claude Opus 4.7 · T=0.7 · 0.000Claude Opus 4.7 · T=1 · 0.080Claude Opus 4.7 · T=1.3 · 0.080 (interpolated)Claude Opus 4.7GPT-5.4 · T=0.7 · 0.010 · temperature lockedGPT-5.4 · T=1 · 0.010 (interpolated) · temperature lockedGPT-5.4 · T=1.3 · 0.010 (interpolated) · temperature lockedGPT-5.4 ○GPT-5.5 · T=0.7 · 0.000 · temperature lockedGPT-5.5 · T=1 · 0.000 · temperature lockedGPT-5.5 · T=1.3 · 0.000 · temperature lockedGPT-5.5 ○Gemini 3.1 Pro · T=0.7 · 0.239Gemini 3.1 Pro · T=1 · 0.245Gemini 3.1 Pro · T=1.3 · 0.296Gemini 3.1 Pro○ temperature locked: the model ignores the parameter.Hover any point for its value, and for whether it was measured orinterpolated across an occlusion in the source figure.
Temperature sweep. A ring marks a model whose temperature is locked.

Why turning up temperature cannot restore diversity.

When RLHF compresses the logit distribution, the model’s viable vocabulary shrinks, and temperature only reshapes what is already there. The path runs transformer output, unembedding matrix, logits, temperature, softmax. If RLHF has crushed the logit differences, dividing by temperature only stretches a narrow spike, and the same two or three tokens still win.

In a healthy model six or more tokens compete. In a collapsed one, two dominate. Even at a temperature of 2.0, the top two tokens hold more than 70% of the mass.

Logit distributions for a healthy model versus a collapsed model, and the probability distribution at three temperaturesLeft, a healthy model where six or more tokens carry comparable logits. Middle, a collapsed model where two tokens dominate and the rest sit near zero. Right, the collapsed model's probabilities at temperatures 0.5, 1.0 and 2.0: even at 2.0 the top two tokens hold more than 70 percent of the mass, because temperature only reshapes what the logits already decided. Illustrative, not a measurement.“Healthy” model logitsGPT-4.1 / Claude Opus styletoken rank →6+ tokens compete“Collapsed” model logitsGPT-5.5 / Claude Sonnet styletoken rank →2 tokens dominateCollapsed model: probabilityat three temperaturestoken rank →T=0.5T=1.0T=2.0Even at T=2.0 the top two tokenshold >70% of the massIllustrative of the mechanism, not a measurement.The right panel is the softmax of the middle panel’s logits.
Healthy against collapsed logit distributions, and the effect of temperature on a collapsed model.

Creative where it should be. Precise where it must be.

Almost no model manages both. Claude Opus 4.7 and GPT-4.1 come closest. Reading the chart: x is factual diversity, where lower is correct, and y is creative diversity, where higher is better. The green region is where you want to live. The GPT-5 cluster is safe but narrow.

Scatter plot of creative diversity against factual diversity for each model, with an ideal region markedHorizontal axis is convergent-question diversity, where lower is correct. Vertical axis is open-ended diversity, where higher is better. GPT-3.5 Turbo, GPT-4.1 and Claude Opus 4.7 fall in the ideal upper-left region. The GPT-5 models cluster low on both. The Gemini models score high on factual diversity, which is the wrong direction.ideal: low on convergent, high on diversityCategory A diversity (open-ended · higher is better)GPT-3.5 TurboGPT-3.5 Turbo · A 0.653 · C 0.000GPT-4.1GPT-4.1 · A 0.645 · C 0.000Gemini 3.1 Flash-LiteGemini 3.1 Flash-Lite · A 0.313 · C 0.023Claude Haiku 4.5Claude Haiku 4.5 · A 0.313 · C 0.038Gemini 3 FlashGemini 3 Flash · A 0.481 · C 0.149Claude Sonnet 4.6Claude Sonnet 4.6 · A 0.233 · C 0.031Claude Opus 4.7Claude Opus 4.7 · A 0.569 · C 0.000GPT-5.4GPT-5.4 · A 0.252 · C 0.038GPT-5.5GPT-5.5 · A 0.191 · C 0.057Gemini 3.1 ProGemini 3.1 Pro · A 0.477 · C 0.126
The two knobs: are models collapsing on the right things?

One model understood the assignment.

Claude Opus shows category-calibrated diversity: bright on creative, fading through opinion, black on factual. Nobody else does. The GPT-5.4 and GPT-5.5 rows are dark on everything, which is variance collapse.

Heatmap of semantic diversity by question and model, with models ordered oldest to newestRows are the 25 questions, grouped into open-ended, normative and convergent bands. Columns are models from oldest to newest. Bright is high diversity, dark is near zero. Claude Opus 4.7 is bright on the open-ended rows, mid on normative and dark on convergent, the only model calibrated to the question type. GPT-5.4 and GPT-5.5 are dark in every band.GPT-3.5 TurboGemini 2.5 Pro (no data)GPT-4.1Gemini 3.1 Flash-LiteClaude Haiku 4.5Gemini 3 FlashClaude Sonnet 4.6Claude Opus 4.7GPT-5.4GPT-5.5Gemini 3.1 Prooldest →Aopen-endedA1GPT-3.5 Turbo · A1 · 0.536GPT-4.1 · A1 · 0.560Gemini 3.1 Flash-Lite · A1 · 0.447Claude Haiku 4.5 · A1 · 0.162Gemini 3 Flash · A1 · 0.382Claude Sonnet 4.6 · A1 · 0.356Claude Opus 4.7 · A1 · 0.487GPT-5.4 · A1 · 0.305GPT-5.5 · A1 · 0.205Gemini 3.1 Pro · A1 · 0.460A2GPT-3.5 Turbo · A2 · 0.651GPT-4.1 · A2 · 0.653Gemini 3.1 Flash-Lite · A2 · 0.422Claude Haiku 4.5 · A2 · 0.245Gemini 3 Flash · A2 · 0.598Claude Sonnet 4.6 · A2 · 0.322Claude Opus 4.7 · A2 · 0.573GPT-5.4 · A2 · 0.289GPT-5.5 · A2 · 0.002Gemini 3.1 Pro · A2 · 0.225A3GPT-3.5 Turbo · A3 · 0.653GPT-4.1 · A3 · 0.649Gemini 3.1 Flash-Lite · A3 · 0.298Claude Haiku 4.5 · A3 · 0.560Gemini 3 Flash · A3 · 0.460Claude Sonnet 4.6 · A3 · 0.124Claude Opus 4.7 · A3 · 0.576GPT-5.4 · A3 · 0.238GPT-5.5 · A3 · 0.165Gemini 3.1 Pro · A3 · 0.435A4GPT-3.5 Turbo · A4 · 0.660GPT-4.1 · A4 · 0.660Gemini 3.1 Flash-Lite · A4 · 0.120Claude Haiku 4.5 · A4 · 0.111Gemini 3 Flash · A4 · 0.269Claude Sonnet 4.6 · A4 · 0.102Claude Opus 4.7 · A4 · 0.571GPT-5.4 · A4 · 0.085GPT-5.5 · A4 · 0.071Gemini 3.1 Pro · A4 · 0.676A5GPT-3.5 Turbo · A5 · 0.653GPT-4.1 · A5 · 0.642Gemini 3.1 Flash-Lite · A5 · 0.435Claude Haiku 4.5 · A5 · 0.256Gemini 3 Flash · A5 · 0.680Claude Sonnet 4.6 · A5 · 0.282Claude Opus 4.7 · A5 · 0.573GPT-5.4 · A5 · 0.151GPT-5.5 · A5 · 0.069Gemini 3.1 Pro · A5 · 0.549A6GPT-3.5 Turbo · A6 · 0.653GPT-4.1 · A6 · 0.647Gemini 3.1 Flash-Lite · A6 · 0.004Claude Haiku 4.5 · A6 · 0.384Gemini 3 Flash · A6 · 0.427Claude Sonnet 4.6 · A6 · 0.124Claude Opus 4.7 · A6 · 0.560GPT-5.4 · A6 · 0.102GPT-5.5 · A6 · 0.051Gemini 3.1 Pro · A6 · 0.369A7GPT-3.5 Turbo · A7 · 0.664GPT-4.1 · A7 · 0.653Gemini 3.1 Flash-Lite · A7 · 0.524Claude Haiku 4.5 · A7 · 0.420Gemini 3 Flash · A7 · 0.671Claude Sonnet 4.6 · A7 · 0.451Claude Opus 4.7 · A7 · 0.573GPT-5.4 · A7 · 0.560GPT-5.5 · A7 · 0.618Gemini 3.1 Pro · A7 · 0.649A8GPT-3.5 Turbo · A8 · 0.655GPT-4.1 · A8 · 0.664Gemini 3.1 Flash-Lite · A8 · 0.124Claude Haiku 4.5 · A8 · 0.176Gemini 3 Flash · A8 · 0.180Claude Sonnet 4.6 · A8 · 0.015Claude Opus 4.7 · A8 · 0.573GPT-5.4 · A8 · 0.011GPT-5.5 · A8 · 0.040Gemini 3.1 Pro · A8 · 0.176A9GPT-3.5 Turbo · A9 · 0.647GPT-4.1 · A9 · 0.645Gemini 3.1 Flash-Lite · A9 · 0.338Claude Haiku 4.5 · A9 · 0.413Gemini 3 Flash · A9 · 0.453Claude Sonnet 4.6 · A9 · 0.085Claude Opus 4.7 · A9 · 0.571GPT-5.4 · A9 · 0.407GPT-5.5 · A9 · 0.351Gemini 3.1 Pro · A9 · 0.435A10GPT-3.5 Turbo · A10 · 0.664GPT-4.1 · A10 · 0.653Gemini 3.1 Flash-Lite · A10 · 0.389Claude Haiku 4.5 · A10 · 0.344Gemini 3 Flash · A10 · 0.622Claude Sonnet 4.6 · A10 · 0.420Claude Opus 4.7 · A10 · 0.573GPT-5.4 · A10 · 0.335GPT-5.5 · A10 · 0.255Gemini 3.1 Pro · A10 · 0.695BnormativeB1GPT-3.5 Turbo · B1 · 0.627GPT-4.1 · B1 · 0.635Gemini 3.1 Flash-Lite · B1 · 0.227Claude Haiku 4.5 · B1 · 0.289Gemini 3 Flash · B1 · 0.382Claude Sonnet 4.6 · B1 · 0.180Claude Opus 4.7 · B1 · 0.313GPT-5.4 · B1 · 0.267GPT-5.5 · B1 · 0.447Gemini 3.1 Pro · B1 · 0.622B2GPT-3.5 Turbo · B2 · 0.622GPT-4.1 · B2 · 0.622Gemini 3.1 Flash-Lite · B2 · 0.051Claude Haiku 4.5 · B2 · 0.149Gemini 3 Flash · B2 · 0.165Claude Sonnet 4.6 · B2 · 0.031Claude Opus 4.7 · B2 · 0.305GPT-5.4 · B2 · 0.064GPT-5.5 · B2 · 0.102Gemini 3.1 Pro · B2 · 0.351B3GPT-3.5 Turbo · B3 · 0.640GPT-4.1 · B3 · 0.645Gemini 3.1 Flash-Lite · B3 · 0.018Claude Haiku 4.5 · B3 · 0.073Gemini 3 Flash · B3 · 0.176Claude Sonnet 4.6 · B3 · 0.031Claude Opus 4.7 · B3 · 0.298GPT-5.4 · B3 · 0.035GPT-5.5 · B3 · 0.089Gemini 3.1 Pro · B3 · 0.269B4GPT-3.5 Turbo · B4 · 0.627GPT-4.1 · B4 · 0.647Gemini 3.1 Flash-Lite · B4 · 0.615Claude Haiku 4.5 · B4 · 0.545Gemini 3 Flash · B4 · 0.549Claude Sonnet 4.6 · B4 · 0.313Claude Opus 4.7 · B4 · 0.313GPT-5.4 · B4 · 0.238GPT-5.5 · B4 · 0.356Gemini 3.1 Pro · B4 · 0.695B5GPT-3.5 Turbo · B5 · 0.647GPT-4.1 · B5 · 0.640Gemini 3.1 Flash-Lite · B5 · 0.089Claude Haiku 4.5 · B5 · 0.084Gemini 3 Flash · B5 · 0.089Claude Sonnet 4.6 · B5 · 0.069Claude Opus 4.7 · B5 · 0.313GPT-5.4 · B5 · 0.038GPT-5.5 · B5 · 0.040Gemini 3.1 Pro · B5 · 0.176B6GPT-3.5 Turbo · B6 · 0.644GPT-4.1 · B6 · 0.620Gemini 3.1 Flash-Lite · B6 · 0.269Claude Haiku 4.5 · B6 · 0.435Gemini 3 Flash · B6 · 0.524Claude Sonnet 4.6 · B6 · 0.109Claude Opus 4.7 · B6 · 0.313GPT-5.4 · B6 · 0.269GPT-5.5 · B6 · 0.564Gemini 3.1 Pro · B6 · 0.584B7GPT-3.5 Turbo · B7 · 0.640GPT-4.1 · B7 · 0.635Gemini 3.1 Flash-Lite · B7 · 0.076Claude Haiku 4.5 · B7 · 0.035Gemini 3 Flash · B7 · 0.093Claude Sonnet 4.6 · B7 · 0.038Claude Opus 4.7 · B7 · 0.313GPT-5.4 · B7 · 0.045GPT-5.5 · B7 · 0.042Gemini 3.1 Pro · B7 · 0.269B8GPT-3.5 Turbo · B8 · 0.645GPT-4.1 · B8 · 0.627Gemini 3.1 Flash-Lite · B8 · 0.489Claude Haiku 4.5 · B8 · 0.536Gemini 3 Flash · B8 · 0.673Claude Sonnet 4.6 · B8 · 0.298Claude Opus 4.7 · B8 · 0.298GPT-5.4 · B8 · 0.545GPT-5.5 · B8 · 0.509Gemini 3.1 Pro · B8 · 0.389CconvergentC1GPT-3.5 Turbo · C1 · 0.004GPT-4.1 · C1 · 0.004Gemini 3.1 Flash-Lite · C1 · 0.111Claude Haiku 4.5 · C1 · 0.004Gemini 3 Flash · C1 · 0.255Claude Sonnet 4.6 · C1 · 0.011Claude Opus 4.7 · C1 · 0.004GPT-5.4 · C1 · 0.002GPT-5.5 · C1 · 0.160Gemini 3.1 Pro · C1 · 0.231C2GPT-3.5 Turbo · C2 · 0.004Gemini 2.5 Pro · C2 · 0.004GPT-4.1 · C2 · 0.004Gemini 3.1 Flash-Lite · C2 · 0.004Claude Haiku 4.5 · C2 · 0.004Gemini 3 Flash · C2 · 0.004Claude Sonnet 4.6 · C2 · 0.004Claude Opus 4.7 · C2 · 0.004GPT-5.4 · C2 · 0.004GPT-5.5 · C2 · 0.165Gemini 3.1 Pro · C2 · 0.004C3GPT-3.5 Turbo · C3 · 0.002GPT-4.1 · C3 · 0.004Gemini 3.1 Flash-Lite · C3 · 0.027Claude Haiku 4.5 · C3 · 0.082Gemini 3 Flash · C3 · 0.384Claude Sonnet 4.6 · C3 · 0.067Claude Opus 4.7 · C3 · 0.004GPT-5.4 · C3 · 0.015GPT-5.5 · C3 · 0.004Gemini 3.1 Pro · C3 · 0.245C4GPT-3.5 Turbo · C4 · 0.004GPT-4.1 · C4 · 0.004Gemini 3.1 Flash-Lite · C4 · 0.004Claude Haiku 4.5 · C4 · 0.042Gemini 3 Flash · C4 · 0.031Claude Sonnet 4.6 · C4 · 0.004Claude Opus 4.7 · C4 · 0.004GPT-5.4 · C4 · 0.215GPT-5.5 · C4 · 0.031Gemini 3.1 Pro · C4 · 0.005C5GPT-3.5 Turbo · C5 · 0.004GPT-4.1 · C5 · 0.004Gemini 3.1 Flash-Lite · C5 · 0.024Claude Haiku 4.5 · C5 · 0.027Gemini 3 Flash · C5 · 0.091Claude Sonnet 4.6 · C5 · 0.024Claude Opus 4.7 · C5 · 0.004GPT-5.4 · C5 · 0.005GPT-5.5 · C5 · 0.002Gemini 3.1 Pro · C5 · 0.178C6GPT-3.5 Turbo · C6 · 0.004GPT-4.1 · C6 · 0.004Gemini 3.1 Flash-Lite · C6 · 0.004Claude Haiku 4.5 · C6 · 0.004Gemini 3 Flash · C6 · 0.178Claude Sonnet 4.6 · C6 · 0.005Claude Opus 4.7 · C6 · 0.004GPT-5.4 · C6 · 0.011GPT-5.5 · C6 · 0.004Gemini 3.1 Pro · C6 · 0.176C7GPT-3.5 Turbo · C7 · 0.004GPT-4.1 · C7 · 0.004Gemini 3.1 Flash-Lite · C7 · 0.016Claude Haiku 4.5 · C7 · 0.118Gemini 3 Flash · C7 · 0.076Claude Sonnet 4.6 · C7 · 0.085Claude Opus 4.7 · C7 · 0.004GPT-5.4 · C7 · 0.004GPT-5.5 · C7 · 0.004Gemini 3.1 Pro · C7 · 0.0040.0 · near zero0.7 · high diversity
Semantic diversity by model and question. Bright is high diversity, dark is near zero. Questions are labelled by code; hover a cell for its value.

Four plausible causes. Possibly all four at once.

  • Reward shaping rewards consistencyHuman raters upvote confident, fluent, repeatable answers. Diversity gets penalised as inconsistency. Across many training rounds, the model learns to converge.
  • Safety guardrails compress around a safe centreIf a creative response has any probability of being problematic, the model learns to avoid the whole region of response space. Safety and creativity trade off.
  • IP protection suppresses the long tailNewer models are trained not to reproduce outputs traceable to specific sources, which inadvertently kills the long tail of creative outputs that would make a model’s fingerprint identifiable.
  • Benchmark overfitting trains away the varianceBenchmarks reward single correct answers. A model scoring 92% on MMLU has no incentive to maintain diversity on questions with “good” answers. The signal optimises away anything benchmarks don’t reward.
Diversity matters · pick GPT-4.1 or Claude OpusTasks that need the distribution.
  • Brainstorming and ideation
  • Marketing copy variants for A/B tests
  • Customer persona generation
  • Exploring solution spaces in product design
  • Any “give me 10 different approaches” workflow
Convergence is fine · GPT-5.x or Claude SonnetTasks that just need the mode.
  • Summarisation
  • Classification and routing
  • Q&A over your own documents
  • Code generation (one right answer)
  • Extraction tasks

If you are building a product on top of a model that is collapsing, your product is silently homogenising too, and you will not see it in your evals.

02 · The problem with prompting

Teaching AI to reason about people, not stereotypes.

A large language model is a remarkably good average human. The trouble starts the moment you ask it to be somebody in particular.

Out of the box, the AI is the average person.

Ask a frontier model the standard personality-test questions, the IPIP items psychologists ask people, with no instruction to play a role. The typical answer it gives is the typical answer people give. Its answers land on the human average. We lead with what it gets right, because that is the trust the rest of this argument spends.

Ask for a kind of person, and it caricatures.

Ask it to answer as a woman rather than a man and it stops being accurate and starts exaggerating. Women really do score a little higher on agreeableness, about 0.6 standard deviations, a mild tilt with heavy overlap: the average woman scores above roughly 3 in 4 men.

The model gets the direction right but inflates the size to about 2 standard deviations, two almost-separate species, 98 in 100. Given only “a woman,” the one prior it has to reach for is the stereotype, so it leans on it hard.

Right direction. Three times too loud.

The solution to bias is more bias.

Counter-intuitive, but true: every extra detail is its own stereotype, but they pull in different directions. Pile on enough and they cancel each other out, and the category resolves into one real person.

“Woman, 40s” is just a category, and the model asks itself whether she is more agreeable. Add that she studied law, had a rural childhood, has two kids and runs marathons, and the pulls start competing. “Woman, 40, lawyer” is a person: the stereotypes are still true, they just stop explaining anything. A stereotype is a fine place to start and a terrible place to stop.

Feed it the whole person, and the caricature collapses.

Cambium does this deliberately: we hand the model the equivalent of a decade of acquaintance at once, hundreds of co-occurring details it has to reconcile into one coherent person. The swings are the competing pulls overshooting and correcting, and they shrink as more traits pin the answer down.

It settles on the real 0.6, not zero, because the difference is real, only small. The remaining gap is not “no difference”; it is the real difference. The leftover gap is signal, not error.

But only if the person could exist.

Real people are not drawn at random. Income tracks education, which tracks occupation, age and place. Feed the model an impossible combination and there is no real person behind it, only the stereotypes it can reach for.

So the lever cuts both ways. A realistic bundle, 41, married, ICU nurse, two kids, mid income, suburban, lets the competing pulls offset each other. A random one, 18, retired neurosurgeon, earns $12k, five kids, no schooling, lets them reinforce, and the bias compounds.

Realistic combinations cancel. Random ones compound.

Cambium’s job: build real people, at scale.

Public data almost always arrives as summaries, “this tract is 52% female, median income $54k”, which throw away the linkages between traits. Cambium’s synthesis engine works the other way, reconstructing individuals whose distributions add up across 200-plus datasets at once.

The result matches reality down to 10 to 20 households, with thousands of variables per person, built only from public aggregates, so no private record is ever touched. An LLM goes from believable but wrong about any group to a faithful stand-in for a real population.

200+ public datasetsCensus, BLS, IRS and state datasets. Summaries, with the links thrown away.
DisaggregationThe synthesis engine: aggregates to individuals. 200-plus datasets reconciled at once, rebuilding the linkages behind them.
Real individualsThousands of attributes each. 10 to 20 household resolution. Public aggregates only.
03 · The rebuild

Rebuilding a realistic population from summary numbers.

Averages describe one column at a time. The method below puts the people back.

Averages describe one column at a time.

A summary like “average income £39,400, average commute 39 minutes” tells you about each thing on its own, but nothing about how they connect. Five numbers, sitting alone.

Behind those five numbers were ten people, each with a sex, an age, years of education, an income and a commute. What leaves the room is the summary: ten numbers, instead of ten people.

CharacteristicAverageSpread
Share female50%n/a
Age40.7± 12.5
Education14.5 yrs± 2.6
Income£39,400± £12,700
Commute39 min± 17.6

“Spread” just means how much people differ from the average.

The characteristics were connected.

More education went with higher income. Longer commutes went with higher pay, too. The averages threw those stories away.

Rebuild the people.

Take the summary, the averages and spread, add a small sample showing how things connect, and reconstruct a full, believable table of individuals.

Nudge the table until the totals match.

Begin with a rough draft table. Match the row totals. Then match the column totals. Repeat until the totals match the real-world summary. Relationships between columns are held fixed, never distorted.

Statisticians call this iterative proportional fitting. We call it the balancing method.

Averages and spreadIdentical to the original summary numbers.
Relationships between columnsPreserved. Education still tracks income.

It connects to everything.

Join on shared characteristics, sex, age, income, to enrich each person: spending and media habits attach to the individuals they belong to. Averages sit alone. Synthetic people can be joined, enriched, and put to work, safely.

From believable but wrong to a faithful stand-in for a real population. Polling, market research, policy and patient populations, on demand, with no privacy risk.

Request a demo →