Skip to main content

Gemini Flash Lite

3 runs · 3 datasets · 1 model

slug: gemini-flash-lite

0.776
Best SPS · subpop

Disaggregated subgroup scorecard. Each card below is one published run for this vendor; expand the question-type and demographic-subgroup sections to see the matrix beneath the headline SPS. Where coverage permits, 95% CI bands accompany the point estimate.

What each cell measures

Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup — published survey ground truth, not a model's guess about the subgroup. Higher = closer to how that real subgroup actually answered.

Demographic conditioning here is not stereotyping: every conditioned score is checked against what real subgroup members said, not against assumptions about them — and where the vendor's output diverges from the real subgroup, the score drops.

  • globalopinionqa — ground truth: Durmus et al. 2023, Anthropic — llm_global_opinions
  • opinionsqa — ground truth: Santurkar et al., ICML 2023 — Whose Opinions Do LLMs Reflect? (derived from Pew American Trends Panel)
  • subpop — ground truth: Suh et al., ACL 2025 — SubPOP: Subpopulation-Level Opinion Prediction

Low-n cells: cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are reported for transparency, not as findings.

CIs: "no CI — single run" marks point estimates from a single run with no confidence interval yet; treat the uncertainty as unknown, never zero.

Columns: Score = distributional parity (p_dist) restricted to the subgroup · p_cond = conditioning strength vs the unconditioned baseline · N = questions answered under that conditioning · Cov. = coverage dot derived from N (green = high, ≥100 · amber = medium, 50–99 · red = low, <50). Topic tables: SPS = Survey Parity Score for the topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 · p_refuse = 1 − mean(|R_model − R_human|).

Multiple comparisons: with this many subgroup cells, a few extreme cells are expected by chance alone — read patterns across a dimension, not single cells.

globalopinionqa raw

raw--gemini-2.5-flash-lite--tdefault--tplcurrent--ddafa278

0.761 ± 0.032
SPS · 95% CI [0.729, 0.792] · n = 100
Question-type breakdown (6 topics)
Topic SPS p_dist p_rank p_refuse N
Health & Science low n — suggestive only 0.887 0.775 1.000 1.000 1
Economy & Work low n — suggestive only 0.730 0.793 0.667 0.949 3
International Relations & Security 0.665 0.689 0.641 0.986 60
Politics & Governance 0.631 0.659 0.602 0.962 30
General Attitudes low n — suggestive only 0.594 0.789 0.398 1.000 4
Trust & Wellbeing low n — suggestive only 0.490 0.457 0.523 1.000 2

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

opinionsqa raw

raw--gemini-2.5-flash-lite--tdefault--tplcurrent--bdb40b06

0.738 ± 0.032
SPS · 95% CI [0.704, 0.769] · n = 100
Question-type breakdown (8 topics)
Topic SPS p_dist p_rank p_refuse N
Social Values & Religion 0.799 0.794 0.804 15
Health & Science low n — suggestive only 0.768 0.763 0.774 1
Economy & Work 0.765 0.750 0.781 20
Technology & Digital Life low n — suggestive only 0.764 0.736 0.792 4
Politics & Governance low n — suggestive only 0.763 0.728 0.799 1
General Attitudes 0.736 0.739 0.733 38
Media & Information low n — suggestive only 0.701 0.744 0.658 2
International Relations & Security 0.660 0.661 0.659 19

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on opinionsqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

subpop raw

raw--gemini-2.5-flash-lite--tdefault--tplcurrent--46e46626

0.776 ± 0.029
SPS · 95% CI [0.746, 0.803] · n = 100
Question-type breakdown (9 topics)
Topic SPS p_dist p_rank p_refuse N
Identity & Demographics low n — suggestive only 0.846 0.838 0.854 0.995 1
Technology & Digital Life 0.735 0.700 0.771 0.993 16
International Relations & Security 0.730 0.712 0.747 0.991 15
Health & Science low n — suggestive only 0.693 0.641 0.744 0.994 6
General Attitudes 0.682 0.693 0.671 0.932 17
Economy & Work 0.678 0.629 0.727 0.992 13
Social Values & Religion 0.605 0.587 0.624 0.985 27
Media & Information low n — suggestive only 0.604 0.579 0.629 0.978 1
Politics & Governance low n — suggestive only 0.602 0.585 0.620 0.981 4

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on subpop. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

No demographic conditioning data has been published for this vendor yet. The question-type matrix above shows topic-level parity; subgroup rows fill in once Althing-style conditioned runs land.

← Back to leaderboard