Human Model Leaderboard
Benchmarks from across the industry for simulation, HCI and models of human behaviour broadly, with who has been scored on each.
68 evals ยท updated 2026-09-14
Overview
Three of the leading benchmarks people have put forward for how well a model can simulate a human being. tau-USI scores how closely a model matches what one real person did in a conversation. SimBench scores how closely it matches how a whole population answered a survey. SOUL-Index runs models across 23 tasks and is the most comprehensive to date.
Talks like a person
tau-USI/
| System | Source | USI |
|---|---|---|
| Human (inter-annotator) | published table | 92.7 |
| DeepSeek-V3.1best general model | published table | 76.0 |
| OSim-8BOdysSim | published table | 75.4 |
| OSim-4B-MidOdysSim | published table | 72.6 |
| OSim-Inst-8BOdysSim | published table | 71.4 |
| CoSER-8B | published table | 67.2 |
| OSim-8B-MidOdysSim | published table | 67.1 |
| OSim-4BOdysSim | published table | 66.8 |
| OSim-Inst-4BOdysSim | published table | 64.1 |
| UserLM-8B | published table | 62.0 |
| HumanLike-7B | published table | 59.8 |
| HumanLM-opinion | published table | 46.9 |
Answers like a population
SimBench16 systems
| System | Source | SimBench score |
|---|---|---|
| Claude 3.7 Sonnet | published table | 40.8 |
| Claude 3.7 Sonnet (4000) | published table | 39.5 |
| GPT-4.1 | published table | 34.5 |
| DeepSeek-R1 | published table | 34.5 |
| o4-mini-high | published table | 29.0 |
| Llama 3.1 405B Instruct | published table | 28.4 |
| Qwen2.5-72B-Instruct | published table | 27.6 |
| Qwen2.5-32B-Instruct | published table | 23.8 |
| OLMo-2-32B-DPO | published table | 19.8 |
| OLMo-2-32B | published table | 15.9 |
| OLMo-2-13B | published table | 13.8 |
| Qwen2.5-72B | published table | 13.3 |
Behaves like a person across 23 tasks
SOUL-Index (OdysSim)/
| System | Source | Average |
|---|---|---|
| Claude Opus 4.7best general model | published table | 65.5 |
| Ditto-v2-8B | published table | 66.5 |
| OSim-8B-Inst-Post | published table | 66.0 |
| OSim-8B-Inst | published table | 65.7 |
| OSim-8B-Inst-Post (no distillation) | published table | 65.3 |
| OSim-8B | published table | 64.6 |
| OSim-8B-Post | published table | 63.8 |
| OSim-4B | published table | 62.6 |
| OSim-4B-Post | published table | 60.5 |
| HumanLM | published table | 48.7 |
| OSim-8B-Inst-Mid | published table | 43.1 |
| OSim-8B-Mid | published table | 41.1 |
| SotopiaRL-7B | published table | 39.7 |