What it is
ProGen showed that an autoregressive language model, prompted with a functional 'tag', can generate artificial proteins (e.g. novel lysozymes) that are actually active in the lab. The line matured at Profluent, which used protein LMs to design OpenCRISPR-1, an open-sourced, AI-generated gene editor.
Evidence trail
BioAtlas keeps the path from source to decision visible. A connection records provenance; it does not imply that evidence is sufficient for every context.
Model passport
How ProGen / ProGen2 / ProGen3 represents biology
Category is navigation. These fields describe the model-specific computational transformation and deliberately override broad category defaults.
Biological scale
Modalities & tasks
Registry, claims and frontier intelligence
Version history not yet curated
1 version record · release year not yet normalized. Model-family identity remains separate from capability and access changes.
Explore version lineage →0 normalized claims
No task, dataset, split and metric claim has been normalized for this record yet.
Open claim intelligence →3 connected frontiers
Genome design · Peer-reviewed capability
Inspect research horizon →Connected research frontiers
These records describe active research directions, not guaranteed capabilities of this model. Evidence stages and unresolved questions are preserved separately.
Genome-scale generative biology
Arc Institute · Stanford · NVIDIA · 2026-03-01Can a foundation model read, predict and design biological sequence continuously from single nucleotides to megabase-scale genomes?
Evidence boundary and unresolved questions
Generative plausibility is not equivalent to biological viability, function or safety. Long generated sequences require extensive synthesis, containment and functional review.
- What biological constraints are learned versus memorized?
- How should whole-genome designs be evaluated before synthesis?
- Can mechanistic interpretability keep pace with model scale?
genome foundation model · long context · sequence design · biosafetyOpen frontier record →Multimodal protein programming
EvolutionaryScale · 2025-01-16Can one generative model reason jointly over protein sequence, structure and function and create functional proteins from mixed prompts?
Evidence boundary and unresolved questions
One striking protein demonstration does not establish general success across enzymes, therapeutics or complex multi-objective design tasks.
- How frequently do generated functions survive experimental testing?
- Can the model optimize potency, stability and safety together?
- How should synthetic training labels affect confidence?
multimodal · protein language model · function generation · synthetic biologyOpen frontier record →Bridge-RNA programmable DNA recombination
Arc Institute · UC Berkeley · Stanford · 2024-06-26Can RNA programmably specify both target and donor DNA to insert, excise or invert large sequences without relying on conventional CRISPR cutting and repair?
Evidence boundary and unresolved questions
The original 2024 work was early-stage and bacterial. Efficiency, specificity, delivery and control in mammalian cells require separate validation.
- Can the system work efficiently and specifically in human cells?
- How are off-target recombination and repeated sequences controlled?
- Can delivery support therapeutically relevant tissues and cargo sizes?
genome editing · bridge RNA · recombinase · large DNA editsOpen frontier record →Inputs and outputs
Inputs
Target structure or design objectiveOutputs
Designed sequencesCandidate structuresScientific and technical profile
Scientific principles
Technology
Scientific lineage
These are transparent concept matches—not claims that one scientist alone caused this model. Each connection is based on the model’s recorded domain, scientific principles, technical terms or an explicit lineage link.
Atomic structures of biologically important molecules by X-ray crystallography
Dorothy Crowfoot HodgkinStructure-based drug design depends on the experimental structural tradition she helped establish.
Information, entropy and communication
Claude E. ShannonSequence modelling, cross-entropy training, language models, mutual information and representation learning all use Shannon’s framework.
The central dogma and directional information transfer
Francis CrickMulti-omic models and sequence foundation models connect genotype, transcript and protein through this information-flow framework.
Anfinsen’s dogma—the thermodynamic hypothesis
Christian B. AnfinsenProtein structure prediction, inverse folding and generative protein design all assume that sequence strongly constrains structure and function.
Transformer self-attention
Ashish Vaswani and colleaguesProtein, genome, molecule and single-cell foundation models use attention to learn dependencies across biological sequences and multimodal inputs.
The alpha helix, beta sheet and hydrogen-bonded protein structure
Linus Pauling, Robert Corey & Herman BransonProtein representation, fold recognition, structural priors and generative protein design all encode these recurring geometric motifs.
Evaluation evidence
BioAtlas has not yet extracted a structured benchmark claim for this record.
Known limitations
- Performance depends on the evaluation dataset and operating conditions.
- A structured benchmark claim has not yet been extracted for this record.
- Outputs require task-specific scientific and experimental validation.
Milestones
ProGen lysozymes were experimentally active.
Profluent's OpenCRISPR-1 was AI-designed and open-sourced.