AgentFairBench Leaderboard
Demographic disparity in the actions of LLM agents across hiring, lending,
and medical triage. Lower CFR / MASD / rate-disparity means fairer; higher AFB-Score means fairer -
but read the verification legend before trusting any row.
Read this first
Two things are true here at once, and the point of the benchmark is that they disagree.
Across 7909 decisions on four models spanning two vendors (three hosted tiers plus the
open-weights llama31-8b run locally at temperature 0), the magnitude statistic
and the exchangeability test flag different cells. Read either one alone and you will draw
the wrong conclusion.
On magnitude: comparing the six-group score spread (MASD) against a two-replicate pairwise
noise floor overstates disparity by up to 2.25x through statistic arity alone. Against
an arity-matched floor, over the 23 stable-floor cells of the reported family the ratio runs 0.68 to 2.40 (median 0.88). 2 cells have an interval entirely above the floor, both on the open-weights model and in different domains, and the randomization test is null in both: wide score spreads with no consistent group ordering. 6 near-deterministic cells have a degenerate (near-zero) floor and are held out.
On exchangeability: a within-set randomization test on the range of per-group means finds
4 of 34 cells significant after Benjamini-Hochberg (smallest adjusted p = 0.0034). All
four are hiring, none is on the primary tier, one is the open-weights cross-vendor model, and
all appear only at the expanded 24 matched sets. The effect is small: the group effect's
standard deviation is 0.16 to 0.31 of call-to-call noise. In every surviving cell the
lowest-scoring group is white-male-coded names, the opposite of the direction most
readers will assume.
So one cell has a wide spread with no ordering, and four have a consistent ordering on
spreads that never clear the floor. Do not read this page as "no bias", and do not read it as
"bias found". Four secondary-model cells were excluded as incomplete because collection was
cut short by a rate limit, not selected after the fact.
34 / 4Cells reported / excluded
0.68-1.89Ratio range, stable-floor cells
4 / 24BH-significant after correction
7909Decisions, 4 models
Verification levels
Every row below carries exactly one of three verification levels. They are not interchangeable,
and the difference matters for how much a row can be trusted against gaming.
✓verified
Maintainers re-ran the model on the held-out private split themselves.
This is the only level with held-out gaming resistance: the submitter never saw the eval items.
≈ trace-only
Reproduced from submitted traces - the maintainers replayed and
re-scored transcripts the submitter provided, but ran no held-out evaluation of their own.
Trace-only entries carry no held-out gaming resistance: a submitter who tuned to the
released items, or hand-picked favorable traces, would look identical to one who did not.
○ self-reported
Author-run pilot - run by the AgentFairBench maintainers themselves
on the pilot/dev split, not independently re-run on the held-out private split. All four rows
below are self-reported for that reason: they are the authors' own instrument check, not a
verified submission.
Results
Four-fifths screening flag, sonnet / hiring / C0, at 24 matched sets. Impact ratio
0.774, below the 0.80 conventional threshold. Unlike every other flag in this benchmark,
this one coincides with a randomization test that survives correction (adjusted p = 0.0034), so
the two channels agree.
The direction needs stating plainly, because the number alone invites the wrong reading. The
selection rate is lowest for white-male-coded names (0.500) and highest for
black-female-coded names (0.646); the 0.774 ratio is white-male against that highest group. The
score channel points the same way. This is a disparity, and it runs opposite to the pattern the
four-fifths rule was written to detect.
Two limits on how far to take it. The magnitude is small, with a group variance component
under 0.12 percent of total variance, and the arity-matched ratio for this cell remains below
the noise floor. And a four-fifths ratio computed on 24 synthetic profiles is a screening
statistic on a purposive sample, not an adverse-impact finding about any deployed system.
Excluded cells
fable / lending / C0, fable / lending / C3, sonnet /
lending / C0, and sonnet / lending / C3 are excluded from every table on this page.
Collection of these four cells was cut short by a rate limit; they are omitted rather than reported
on whichever profiles happened to finish, because scoring a partial, non-random subset would bias
exactly the disparity quantity being measured.
Submit a model
External models enter this leaderboard only through the PR-based submission protocol, never
estimated or fabricated. A submission with traces the maintainers can replay lands as
≈ trace-only; a submission the maintainers re-run
themselves on the private split lands as ✓verified.
See the submission protocol for the full process.