AgentFairBench Leaderboard

Demographic disparity in the actions of LLM agents across hiring, lending, and medical triage. Lower CFR / MASD / rate-disparity means fairer; higher AFB-Score means fairer - but read the verification legend before trusting any row.

34 / 4Cells reported / excluded
0.68-1.89Ratio range, stable-floor cells
4 / 24BH-significant after correction
7909Decisions, 4 models

Verification levels

Every row below carries exactly one of three verification levels. They are not interchangeable, and the difference matters for how much a row can be trusted against gaming.

✓verified Maintainers re-ran the model on the held-out private split themselves. This is the only level with held-out gaming resistance: the submitter never saw the eval items.
≈ trace-only Reproduced from submitted traces - the maintainers replayed and re-scored transcripts the submitter provided, but ran no held-out evaluation of their own. Trace-only entries carry no held-out gaming resistance: a submitter who tuned to the released items, or hand-picked favorable traces, would look identical to one who did not.
○ self-reported Author-run pilot - run by the AgentFairBench maintainers themselves on the pilot/dev split, not independently re-run on the held-out private split. All four rows below are self-reported for that reason: they are the authors' own instrument check, not a verified submission.

Results

One row per model tier. Ratio range and BH-significant count are computed over that tier's reported cells in results/v11/analysis.json. Full per-cell detail, confidence intervals, and the excluded-cell log live in results/v11/tables.md.
ModelVerification DecisionsCells (reported/excl.) Replicate depth kArity-matched ratio BH-significantFlags
fable ○ self-reported 671 4 / 2 k=2 0.95-0.99 1 / 4 -
haiku ○ self-reported 3744 20 / 0 k=2-3 0.68-1.02 0 / 20 -
llama31-8b ○ self-reported 578 4 / 0 k=2 0.79-0.79 1 / 4 -
sonnet ○ self-reported 614 4 / 2 k=2 0.90-1.06 2 / 4 4/5ths
Four-fifths screening flag, sonnet / hiring / C0, at 24 matched sets. Impact ratio 0.774, below the 0.80 conventional threshold. Unlike every other flag in this benchmark, this one coincides with a randomization test that survives correction (adjusted p = 0.0034), so the two channels agree.

The direction needs stating plainly, because the number alone invites the wrong reading. The selection rate is lowest for white-male-coded names (0.500) and highest for black-female-coded names (0.646); the 0.774 ratio is white-male against that highest group. The score channel points the same way. This is a disparity, and it runs opposite to the pattern the four-fifths rule was written to detect.

Two limits on how far to take it. The magnitude is small, with a group variance component under 0.12 percent of total variance, and the arity-matched ratio for this cell remains below the noise floor. And a four-fifths ratio computed on 24 synthetic profiles is a screening statistic on a purposive sample, not an adverse-impact finding about any deployed system.

Excluded cells

fable / lending / C0, fable / lending / C3, sonnet / lending / C0, and sonnet / lending / C3 are excluded from every table on this page. Collection of these four cells was cut short by a rate limit; they are omitted rather than reported on whichever profiles happened to finish, because scoring a partial, non-random subset would bias exactly the disparity quantity being measured.

Submit a model

External models enter this leaderboard only through the PR-based submission protocol, never estimated or fabricated. A submission with traces the maintainers can replay lands as ≈ trace-only; a submission the maintainers re-run themselves on the private split lands as ✓verified. See the submission protocol for the full process.

Each row pins the model tier, harness version, and run date in the repository. See results/v11/analysis.json for the source data behind this page and leaderboard/results.json for the machine-readable leaderboard this page is generated alongside.