Skip to main content

Random Baseline

4 runs · 4 datasets · 1 model

slug: random-baseline

0.774
Best SPS · globalopinionqa

Disaggregated subgroup scorecard. Each card below is one published run for this vendor; expand the question-type and demographic-subgroup sections to see the matrix beneath the headline SPS. Where coverage permits, 95% CI bands accompany the point estimate.

What each cell measures

Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup — published survey ground truth, not a model's guess about the subgroup. Higher = closer to how that real subgroup actually answered.

Demographic conditioning here is not stereotyping: every conditioned score is checked against what real subgroup members said, not against assumptions about them — and where the vendor's output diverges from the real subgroup, the score drops.

  • globalopinionqa — ground truth: Durmus et al. 2023, Anthropic — llm_global_opinions
  • gss — ground truth: NORC at the University of Chicago — General Social Survey
  • opinionsqa — ground truth: Santurkar et al., ICML 2023 — Whose Opinions Do LLMs Reflect? (derived from Pew American Trends Panel)
  • subpop — ground truth: Suh et al., ACL 2025 — SubPOP: Subpopulation-Level Opinion Prediction

Low-n cells: cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are reported for transparency, not as findings.

CIs: "no CI — single run" marks point estimates from a single run with no confidence interval yet; treat the uncertainty as unknown, never zero.

Columns: Score = distributional parity (p_dist) restricted to the subgroup · p_cond = conditioning strength vs the unconditioned baseline · N = questions answered under that conditioning · Cov. = coverage dot derived from N (green = high, ≥100 · amber = medium, 50–99 · red = low, <50). Topic tables: SPS = Survey Parity Score for the topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 · p_refuse = 1 − mean(|R_model − R_human|).

Multiple comparisons: with this many subgroup cells, a few extreme cells are expected by chance alone — read patterns across a dimension, not single cells.

globalopinionqa baseline

baseline--random-baseline--tdefault--tplcurrent--8855add7

0.774 ± 0.023
SPS · 95% CI [0.751, 0.797] · n = 100
Question-type breakdown (6 topics)
Topic SPS p_dist p_rank p_refuse N
General Attitudes low n — suggestive only 0.695 0.947 0.443 1.000 4
Politics & Governance 0.677 0.893 0.461 0.963 30
International Relations & Security 0.674 0.884 0.464 0.987 60
Trust & Wellbeing low n — suggestive only 0.662 0.876 0.449 1.000 2
Economy & Work low n — suggestive only 0.591 0.904 0.278 0.949 3
Health & Science low n — suggestive only 0.477 0.955 0.000 1.000 1

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

gss baseline

baseline--random-baseline--tdefault--tplcurrent--2ad2dd26

0.735 ± 0.020
SPS · 95% CI [0.716, 0.756] · n = 75
Question-type breakdown (7 topics)
Topic SPS p_dist p_rank p_refuse N
Health & Science low n — suggestive only 0.692 0.801 0.583 0.967 2
International Relations & Security 0.637 0.829 0.446 0.974 33
Trust & Wellbeing low n — suggestive only 0.631 0.823 0.439 0.951 3
Economy & Work low n — suggestive only 0.627 0.804 0.449 0.981 4
Politics & Governance low n — suggestive only 0.599 0.798 0.399 0.975 9
General Attitudes 0.590 0.828 0.352 0.960 10
Social Values & Religion 0.584 0.797 0.372 0.977 14

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on gss. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

opinionsqa baseline

baseline--random-baseline--tdefault--tplcurrent--971f8ad1

0.763 ± 0.008
SPS · 95% CI [0.755, 0.771] · n = 684
Question-type breakdown (10 topics)
Topic SPS p_dist p_rank p_refuse N
Media & Information 0.672 0.800 0.544 0.993 63
Health & Science 0.668 0.804 0.532 0.987 47
Technology & Digital Life 0.660 0.785 0.535 0.993 26
International Relations & Security 0.660 0.817 0.503 0.988 149
Politics & Governance 0.659 0.814 0.503 0.989 40
Social Values & Religion 0.650 0.795 0.505 0.990 37
Economy & Work 0.646 0.796 0.496 0.991 68
General Attitudes 0.635 0.810 0.460 0.989 190
Trust & Wellbeing 0.628 0.779 0.476 0.995 25
Identity & Demographics 0.626 0.807 0.445 0.985 39

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on opinionsqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

subpop baseline

baseline--random-baseline--tdefault--tplcurrent--c53167f6

0.757 ± 0.014
SPS · 95% CI [0.744, 0.771] · n = 200
Question-type breakdown (9 topics)
Topic SPS p_dist p_rank p_refuse N
Identity & Demographics low n — suggestive only 0.826 0.798 0.854 0.995 1
Trust & Wellbeing low n — suggestive only 0.693 0.886 0.500 0.988 2
Economy & Work 0.689 0.852 0.526 0.991 17
Health & Science low n — suggestive only 0.672 0.807 0.537 0.994 5
Technology & Digital Life 0.654 0.814 0.494 0.993 47
General Attitudes 0.644 0.803 0.485 0.913 37
Social Values & Religion 0.640 0.828 0.452 0.987 36
International Relations & Security 0.639 0.798 0.481 0.991 33
Politics & Governance 0.623 0.821 0.426 0.983 22

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on subpop. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

No demographic conditioning data has been published for this vendor yet. The question-type matrix above shows topic-level parity; subgroup rows fill in once Althing-style conditioned runs land.

← Back to leaderboard