Edison Advances
← Benchmarks

BixBench3

Benchmarking AI agents on research-study-scale computational biology tasks

BixBench3 evaluates the ability of AI agents to carry out analyses on the scale of complete research studies. Each task mirrors how scientists often use agents today: the scientist sets the research objective and a general methodological plan, then delegates implementation to the agent. Each task is based on a published computational biology paper. The agent must build a working analysis pipeline that analyzes the paper's raw data and produce structured artifacts supporting its findings. Those artifacts – such as protein abundance matrices or tables of differentially expressed genes – are graded programmatically against reference artifacts from the published study, measuring how well the agent executed the required analyses. We evaluated 13 frontier language models across 20 BixBench3 tasks encompassing the generation of 138 unique artifacts. Their average scores ranged from 0.00 for Gemini 3.1 Flash Lite to 0.48 for GPT 5.6 Sol. These were long and computationally intensive runs: an average attempt took 6.8 hours, processed 102 million tokens, and cost $43 while the largest consumed 1.07 billion tokens, ran for 24 hours, and cost $525.

Paper
PDF
Code
GitHub
Dataset
Hugging Face

Top performers

Bars colored by model provider. Repeat runs are averaged with a standard-error whisker. Hover for details.

Aggregate leaderboard

Mean bixbench3 score across all sub-benchmarks. Only models with coverage ≥ 1/1 are listed — comparisons against partial-coverage runs would be misleading.
# Model Mean bixbench3 score Coverage
1 gpt-5.6-sol 0.480 1/1
2 kimi-k3 0.470 1/1
3 glm-5.2 0.463 1/1
4 claude-opus-4.8 0.460 1/1
5 gpt-5.5 0.431 1/1
6 gemini-3.5-flash 0.421 1/1
7 claude-opus-5 0.406 1/1
8 gpt-5.4-mini 0.364 1/1
9 gpt-5.4-nano 0.275 1/1
10 gemini-3.1-pro-preview 0.248 1/1
11 claude-sonnet-4.6 0.153 1/1
12 claude-haiku-4.5 0.015 1/1
13 gemini-3.1-flash-lite 0.000 1/1

Sub-benchmarks

Click any row for the per-sub-benchmark leaderboard.
Overall (20-task cohort)
overall
Leader gpt-5.6-sol 0.480