Skip to main content
Task and benchmark claim graph

Separate the claim from the model.

A benchmark statement is meaningful only with its model version, task, dataset, split, metric, reporting source and caveats. BioAtlas groups claims only when those fields align instead of manufacturing a universal leaderboard.

Canonical claim records

Every curated claim has an inspectable page.

26 claims
26normalized benchmark claims
Protein structure prediction

AlphaFold 2 / 3

AlphaFold 3 · Google DeepMind
peer-reviewed
DatasetCASP14 / complex evaluations
SplitCASP14 blinded targets
Metric or taskStructure accuracy and confidence

Top-performing CASP14 system; consult the primary paper for target-level metrics.

Reported by: Independent community assessment and model developersReplication: multiple independent uses
Known caveats
  • CASP performance does not establish equal accuracy for every target class or drug-relevant complex.
  • AlphaFold 2 and AlphaFold 3 require separate evaluation contexts.
protein-structure|casp14-complex-evaluations|structure-accuracy-and-confidence|casp14-blinded-targetsPrimary source ↗
Protein structure prediction

RoseTTAFold / All-Atom

RoseTTAFold All-Atom · Institute for Protein Design, UW
peer-reviewed
DatasetCASP14 / complex modelling
SplitCASP14 and held-out structure targets
Metric or taskStructure prediction

Competitive three-track structure prediction reported in the primary publication.

Reported by: Model developers with community benchmark contextReplication: independent use documented
Known caveats
  • Results across RoseTTAFold generations are not interchangeable.
  • All-atom complex evaluation requires a task-specific benchmark.
protein-structure|casp14-complex-modelling|structure-prediction|casp14-and-held-out-structure-targetsPrimary source ↗
Biomolecular complex prediction

Chai-1 / Chai-2

Chai-2 · Chai Discovery
developer-reported
DatasetComplex and antibody-design evaluations
SplitPoseBusters and complex-evaluation sets
Metric or taskStructure / design performance

Developer-reported AF3-class complex-prediction performance.

Reported by: Model developersReplication: partial
Known caveats
  • Preprint and developer-reported comparisons require independent reproduction.
  • Benchmark protocol and entity coverage determine comparability.
complex-structure|complex-and-antibody-design-evaluations|structure-design-performance|posebusters-and-complex-evaluation-setsPrimary source ↗
Biomolecular complex prediction

Boltz-1 / Boltz-2

Boltz-2 · MIT (Barzilay & Jaakkola labs)
developer-reported
DatasetPoseBusters and affinity benchmarks
SplitPoseBusters and affinity evaluation sets
Metric or taskStructure and affinity

Open evaluation reports structure prediction and later affinity capabilities.

Reported by: Model developers and open community evaluationsReplication: partial
Known caveats
  • Boltz-1 structure claims and Boltz-2 affinity claims should be separated by version.
  • Affinity performance depends strongly on target family and split design.
complex-structure|posebusters-and-affinity-benchmarks|structure-and-affinity|posebusters-and-affinity-evaluation-setsPrimary source ↗
De novo protein or binder design

RFdiffusion

RFdiffusion2 · Institute for Protein Design, UW
experimental
DatasetExperimental binder validation
SplitProspective designed proteins and binders
Metric or taskDe novo protein design

The primary publication includes experimental validation of generated designs.

Reported by: Model developers with wet-lab validationReplication: experimental primary
Known caveats
  • Experimental success rates depend on design objective and filtering pipeline.
  • A generated backbone still requires sequence design and laboratory confirmation.
protein-design|experimental-binder-validation|de-novo-protein-design|prospective-designed-proteins-and-bindersPrimary source ↗
De novo protein or binder design

ESM3

ESM3 · EvolutionaryScale (now CZ Biohub)
peer-reviewed
DatasetGenerative protein evaluations
SplitGenerative protein evaluations and experimental fluorescent-protein example
Metric or taskSequence, structure and function

Reported multimodal generation includes an experimentally characterized designed protein.

Reported by: Model developersReplication: experimental primary
Known caveats
  • A demonstration protein is not a universal measure of design performance.
  • Hosted and open variants may differ in capability and access.
protein-design|generative-protein-evaluations|sequence-structure-and-function|generative-protein-evaluations-and-experimental-fluorescent-protein-examplePrimary source ↗
Protein–ligand pose prediction

DiffDock

Version history not yet curated · MIT (Barzilay & Jaakkola labs)
peer-reviewed
DatasetPDBBind
SplitPDBBind held-out complexes
Metric or taskTop-ranked docking pose

Peer-reviewed top-ranked docking-pose evaluation.

Reported by: Model developersReplication: independent benchmarks exist
Known caveats
  • Random and temporal splits can produce materially different estimates.
  • Pose accuracy is not binding-affinity accuracy.
binding-pose|pdbbind|top-ranked-docking-pose|pdbbind-held-out-complexesPrimary source ↗
Genomic sequence modelling

Evo / Evo 2

Evo 2 · Arc Institute + Stanford + NVIDIA
peer-reviewed
DatasetGenomic sequence evaluations
SplitHeld-out genomic sequence evaluations
Metric or taskDNA/RNA/protein generation

Peer-reviewed sequence-generation and prediction evaluations across biological scales.

Reported by: Model developersReplication: emerging
Known caveats
  • Generation quality does not establish biological function or safety.
  • Evo and Evo 2 require version-specific evaluation.
genomic-generation|genomic-sequence-evaluations|dna-rna-protein-generation|held-out-genomic-sequence-evaluationsPrimary source ↗
Cell perturbation prediction

State / Stack

Stack · Arc Institute
developer-reported
DatasetPerturbation prediction datasets
SplitPerturbation-prediction datasets
Metric or taskCell-state prediction

Open evaluation of cell-state prediction under perturbation.

Reported by: Model developersReplication: emerging
Known caveats
  • Cell line, tissue, dose and timepoint shifts can dominate performance.
  • In-silico perturbations require experimental validation.
cell-perturbation|perturbation-prediction-datasets|cell-state-prediction|perturbation-prediction-datasetsPrimary source ↗
Single-cell representation learning

scGPT

Version history not yet curated · University of Toronto (Bo Wang Lab)
peer-reviewed
DatasetSingle-cell downstream tasks
SplitSingle-cell downstream-task benchmarks
Metric or taskCell representation

Peer-reviewed evaluation across representation and downstream single-cell tasks.

Reported by: Model developersReplication: multiple independent uses
Known caveats
  • Performance varies by preprocessing, batch correction and downstream task.
  • Representation quality is not equivalent to causal perturbation prediction.
single-cell-representation|single-cell-downstream-tasks|cell-representation|single-cell-downstream-task-benchmarksPrimary source ↗
Antibody sequence modelling

IgLM / AntiBERTy

Version history not yet curated · Johns Hopkins (Gray Lab)
peer-reviewed
DatasetAntibody sequence evaluations
SplitAntibody sequence evaluation sets
Metric or taskAntibody language modelling

Peer-reviewed antibody language-modelling evaluation.

Reported by: Model developersReplication: emerging
Known caveats
  • Sequence plausibility does not guarantee affinity, specificity or developability.
  • Training-set lineage and germline distribution affect generalization.
antibody-language|antibody-sequence-evaluations|antibody-language-modelling|antibody-sequence-evaluation-setsPrimary source ↗
De novo protein or binder design

BoltzGen

BoltzGen · MIT / Boltz team
experimental
DatasetBinder-design evaluations
SplitSplit details not yet normalized
Metric or taskDesign success and structural quality

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
protein-design|binder-design-evaluations|design-success-and-structural-quality|unknown-splitPrimary source ↗
Antibody sequence modelling

RFantibody

RFantibody · Institute for Protein Design, UW
experimental
DatasetExperimental antibody design
SplitSplit details not yet normalized
Metric or taskEpitope-specific de novo design

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
antibody-language|experimental-antibody-design|epitope-specific-de-novo-design|unknown-splitPrimary source ↗
Integrated discovery platform

ESM-2

Version history not yet curated · Meta AI (FAIR)
peer-reviewed
DatasetProtein representation and structure evaluations
SplitSplit details not yet normalized
Metric or taskTransfer / language-model representation

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
platform-discovery|protein-representation-and-structure-evaluations|transfer-language-model-representation|unknown-splitPrimary source ↗
Integrated discovery platform

SaProt

Version history not yet curated · Westlake / Zhejiang collaborators
peer-reviewed
DatasetProtein understanding benchmarks
SplitSplit details not yet normalized
Metric or taskStructure-aware representation

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
platform-discovery|protein-understanding-benchmarks|structure-aware-representation|unknown-splitPrimary source ↗
Protein–ligand pose prediction

Uni-Mol / Uni-Mol2

Version history not yet curated · DP Technology / DeepModeling
developer-reported
DatasetMolecular representation/property benchmarks
SplitSplit details not yet normalized
Metric or task3D molecular representation

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
binding-pose|molecular-representation-property-benchmarks|3d-molecular-representation|unknown-splitPrimary source ↗
Genomic sequence modelling

DNABERT-2

Version history not yet curated · Multi-institution research team
developer-reported
DatasetGenomic downstream tasks
SplitSplit details not yet normalized
Metric or taskGeneral genomic representation

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
genomic-generation|genomic-downstream-tasks|general-genomic-representation|unknown-splitPrimary source ↗
Genomic sequence modelling

Caduceus

Version history not yet curated · Cornell / Tri Dao collaborators
developer-reported
DatasetLong-range genomic benchmarks
SplitSplit details not yet normalized
Metric or taskVariant / sequence prediction

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
genomic-generation|long-range-genomic-benchmarks|variant-sequence-prediction|unknown-splitPrimary source ↗
Genomic sequence modelling

Borzoi

Version history not yet curated · Calico / academic collaborators
peer-reviewed
DatasetFunctional-genomics and RNA-seq evaluations
SplitSplit details not yet normalized
Metric or taskSequence-to-RNA prediction

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
genomic-generation|functional-genomics-and-rna-seq-evaluations|sequence-to-rna-prediction|unknown-splitPrimary source ↗
Integrated discovery platform

RiNALMo

Version history not yet curated · University of Zagreb / A*STAR collaborators
peer-reviewed
DatasetRNA downstream and structure tasks
SplitSplit details not yet normalized
Metric or taskRNA representation / structure prediction

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
platform-discovery|rna-downstream-and-structure-tasks|rna-representation-structure-prediction|unknown-splitPrimary source ↗
Genomic sequence modelling

LucaOne

Version history not yet curated · BioMap / collaborators
peer-reviewed
DatasetDNA/RNA/protein downstream tasks
SplitSplit details not yet normalized
Metric or taskCross-domain biological representation

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
genomic-generation|dna-rna-protein-downstream-tasks|cross-domain-biological-representation|unknown-splitPrimary source ↗
Cell perturbation prediction

Universal Cell Embeddings (UCE)

Version history not yet curated · Stanford / collaborators
peer-reviewed
DatasetCross-dataset single-cell evaluations
SplitSplit details not yet normalized
Metric or taskUniversal cell representation

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
cell-perturbation|cross-dataset-single-cell-evaluations|universal-cell-representation|unknown-splitPrimary source ↗
Cell perturbation prediction

AIDO Cell 1.0

Version history not yet curated · GenBio AI
developer-reported
DatasetVirtual Cell Benchmark 1.0
SplitSplit details not yet normalized
Metric or task31 metrics across five task families

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
cell-perturbation|virtual-cell-benchmark-1-0|31-metrics-across-five-task-families|unknown-splitPrimary source ↗
Integrated discovery platform

AIDO.Tissue

Version history not yet curated · GenBio AI
developer-reported
DatasetSpatial transcriptomics downstream evaluations
SplitSplit details not yet normalized
Metric or taskCell/tissue representation and prediction

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
platform-discovery|spatial-transcriptomics-downstream-evaluations|cell-tissue-representation-and-prediction|unknown-splitPrimary source ↗
Integrated discovery platform

GenBio-PathFM

Version history not yet curated · GenBio AI
developer-reported
DatasetTHUNDER / HEST / PathoROB
SplitSplit details not yet normalized
Metric or taskHistopathology representation and downstream performance

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
platform-discovery|thunder-hest-pathorob|histopathology-representation-and-downstream-performance|unknown-splitPrimary source ↗
Cell perturbation prediction

Tahoe-x1

Version history not yet curated · Tahoe Therapeutics
developer-reported
DatasetFour disease-relevant single-cell evaluation groups
SplitSplit details not yet normalized
Metric or taskEssentiality, cancer hallmarks, cell type and perturbation response

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Reported by: Source authorsReplication: unknown
Known caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.
cell-perturbation|four-disease-relevant-single-cell-evaluation-groups|essentiality-cancer-hallmarks-cell-type-and-perturbation-response|unknown-splitPrimary source ↗