Edison Advances
← Benchmarks

BixBench3

Benchmarking AI agents on research-study-scale computational biology tasks

BixBench3 evaluates the ability of AI agents to carry out analyses on the scale of complete research studies. Each of the 20 tasks mirrors how scientists often use agents today: the scientist sets the research objective and a general methodological plan, then delegates implementation to the agent. Each task is based on a published computational biology paper. The agent must build a working analysis pipeline that analyzes the paper's raw data and produce structured artifacts supporting its findings. Those 138 artifacts – such as protein abundance matrices or tables of differentially expressed genes – are graded programmatically against reference artifacts from the published study, measuring how well the agent executed the required analyses.

Released 2026-08-26 Last updated 2026-09-30
Paper
arXiv↗
Code
GitHub↗
Dataset
Hugging Face↗

Results

Leaderboard

# Model Average score Average cost (USD) Average output tokens Published
1 claude-fable-5.1 (with fallback) 0.515 $65.07 163.9k 2026-09-21
2 claude-opus-5.5 0.502 $42.33 224.6k New 2026-09-30
3 gpt-6-astra 0.492 $44.23 207.5k 2026-09-21
4 grok-4.6 0.488 $64.94 101.6k 2026-09-21
5 gpt-5.6-sol 0.480 $33.88 141.6k 2026-08-26
6 kimi-k3 0.470 $15.34 195.7k 2026-08-26
7 gpt-6-sol 0.470 $13.46 217.6k New 2026-09-30
8 glm-5.2 0.463 $27.01 234.5k 2026-08-26
9 claude-opus-4.8 0.460 $52.52 286.9k 2026-08-26
10 gpt-5.5 0.431 $25.17 73.8k 2026-08-26
11 gemini-3.5-flash 0.421 $43.83 358.8k 2026-08-26
12 claude-opus-5 0.406 $41.17 206.9k 2026-08-26
13 gpt-6-luna 0.369 $1.18 413.2k New 2026-09-30
14 gpt-5.4-mini 0.364 $8.47 421.8k 2026-08-26
15 gpt-5.4-nano 0.275 $3.39 552.7k 2026-08-26
16 gemini-3.1-pro-preview 0.248 $129.14 87.5k 2026-08-26
17 claude-sonnet-4.6 0.153 $88.26 131.2k 2026-08-26
18 claude-haiku-4.5 0.015 $2.20 69.9k 2026-08-26
19 gemini-3.1-flash-lite 0.000 $0.35 17.5k 2026-08-26

Sub-benchmarks

Click any row for the per-sub-benchmark leaderboard.
Overall (20-task cohort)
overall
Leader claude-fable-5.1 (with fallback) 0.515