Foundation models (FMs) are redefining biological design by moving beyond static prediction toward intelligent creation. The latest systems no longer just infer protein structures—they integrate text, sequence, and experimental context to reason across biological scales. By embedding design intent directly into generative logic, they merge understanding and invention into a single computational framework. In part 2, we examine how process-aware models are embedding manufacturability, stability, and scalability into molecular design—and how emerging benchmarks are redefining what it means to validate AI in biomanufacturing.
The New Generation: Multimodal and Controllable Design Foundation Models
The evolution of biological foundation models (FMs) is entering a new phase defined not only by greater scale and multimodality but by intentional control. The earliest models passively captured correlations between sequence and structure; their successors now act as interactive systems capable of reasoning across biological contexts and generating molecules that satisfy precise design goals. In these multimodal and controllable architectures, inference and creation are no longer distinct operations but parts of the same process: a continuous loop in which models can ask and answer biologically meaningful questions, predict outcomes, and design molecules accordingly.
Multimodal Integration
Traditional protein FMs treated sequences or structures as isolated data types. The new generation integrates multiple modalities — sequence, structure, chemical, and textual — into a single representational space. This multimodal learning allows models to connect biological form and function directly, linking what a molecule looks like to what it does and how it behaves experimentally.
ProCyon exemplifies this leap. It is an 11-billion-parameter multimodal foundation model trained across text, protein sequences, 3D structures, and small molecule data to predict phenotypes and perform biological question answering.1 By combining unstructured textual information with molecular and structural embeddings, ProCyon can reason across contexts, retrieving mechanistic explanations or predicting the effects of mutations based on both data and learned biological language. This approach transforms biological modeling from static prediction to active reasoning: users can now pose functional prompts such as “What mutation might improve solubility without destabilizing the active site?” and receive plausible, testable suggestions grounded in multimodal correlations.
A similar model fuses embeddings from sequence, structure, and scientific literature.2 This integration allows knowledge transfer between data-rich and data-poor protein families, substantially improving performance in low-resource domains. By embedding textual descriptions of molecular function and experimental observations alongside structural and sequence data, the model learns to generalize biochemical relationships that are rarely captured in raw sequence alone. Such fusion of empirical and contextual data marks the beginning of what might be called knowledge-grounded biomodeling, in which machine learning draws on both data-driven statistics and semantic understanding of biology.
The value of context is further emphasized by work demonstrating that including metadata — experimental conditions, cellular systems, assay types — during training significantly improves model interpretability and performance.3 Contextual embeddings allow FMs to distinguish between intrinsic molecular properties and condition-dependent effects, enabling more accurate functional predictions. This approach aligns with how experimentalists interpret biological data: meaning is inseparable from context. Embedding this awareness directly into model architecture represents a critical step toward deployable artificial intelligence (AI) systems that can make biologically coherent, process-relevant recommendations.
Together, these developments redefine the role of FMs. They are no longer static pattern recognizers but interactive reasoning systems. In a multimodal framework, a single model can link natural-language hypotheses to sequence edits, structural predictions, or property forecasts. Instead of sequentially querying separate models for structure, function, and manufacturability, researchers can operate within a unified interface — an AI collaborator that “understands” the biochemical and linguistic dimensions of its own data.
Controllable and Generative Models
While multimodal integration allows richer inference, the next frontier is control: generating new molecules that satisfy predefined biological or process constraints. Early generative models relied on brute-force sampling of sequence space, producing millions of candidates that required extensive filtering. The current generation embeds controllability directly into the generation process, steering design toward specific goals such as binding affinity, stability, or ease of manufacture.
PoET-2 exemplifies this approach.4 It is a retrieval-augmented foundation model that supports in-context controllable generation, allowing users to specify desired properties or motifs during sequence synthesis. By referencing examples from relevant protein families stored in memory, PoET-2 can design proteins that retain family-level characteristics while introducing targeted innovations — new binding pockets, altered domains, or structural motifs. This retrieval-based conditioning greatly improves design interpretability: users can trace generated features back to known exemplars, aligning generative creativity with biochemical plausibility.
AlphaDesign introduced another paradigm by coupling generative modeling with structure prediction.5 Rather than retraining a model for each task, AlphaDesign uses AlphaFold itself as an evaluation oracle, iteratively generating sequences, predicting their folds, and refining the most stable outputs. This closed-loop hallucination strategy enables the creation of novel, fold-stable proteins with high confidence in their physical feasibility. It marked one of the first demonstrations that foundation models could autonomously invent functional molecular architectures, leveraging predictive models as evaluators rather than static tools.
Beyond structure, some models now incorporate physical properties directly into their learning objectives. ForceGen was the first generative framework to optimize for nonlinear mechanical and stability properties.6 By embedding elasticity, rigidity, and thermodynamic stability into its loss function, ForceGen moves AI-driven design closer to real-world manufacturability. Instead of producing merely fold-stable proteins, it generates molecules likely to withstand the thermal and mechanical stresses of industrial processing. In doing so, it introduces the concept of design-for-manufacture into the AI workflow, a crucial bridge between molecular discovery and scalable production.
Complementing these top-down generative methods, Local Atomistic Environment models take a bottom-up approach.7 Rather than treating residues as tokens in a global sequence, they encode the fine-grained physical neighborhoods around each atom or residue. This focus on local context captures the microphysics that govern folding pathways and local stability, providing a bridge between atomistic simulation and large-scale learning. By integrating these detailed local representations into broader FMs, researchers can combine the interpretability and mechanistic accuracy of physics-based modeling with the scalability and generalization power of deep learning.
The unifying trend among these advances is controllability. Modern FMs are no longer opaque generators of arbitrary molecules; they are guided systems that can be steered through prompts, conditioning, or multi-objective optimization to produce therapeutically relevant outputs. This shift replaces the brute-force sampling of traditional generative models with guided creativity, where each design is traceable to specific goals or constraints, such as fold type, binding site geometry, process stability, or manufacturability metrics. In the context of drug development, this means AI can now explore chemical and biological space with purpose rather than chance, proposing molecules that are both novel and practical.
Taken together, multimodal and controllable foundation models mark the point where AI truly begins to merge with experimental biology. They enable both understanding and invention: models that interpret data, generate hypotheses, and design molecules aligned with defined functional or process requirements. This evolution represents a profound transformation of the discovery pipeline from static prediction toward interactive, context-aware design. In the next phase, these capabilities will extend beyond the molecular level, integrating process data and manufacturing constraints directly into generative reasoning, an advance that could ultimately yield therapeutic candidates designed not only for efficacy but for seamless translation to production.
Table 1. Comparison of Major Foundation Models for Biological Design and Manufacturability
Designing for Manufacturability
The earliest AI-driven design tools in structural biology focused on getting the fold right — predicting the native three-dimensional conformation of a sequence with high accuracy. For research, that was revolutionary. For industry, it was only the beginning. In biomanufacturing, a protein that folds correctly is not necessarily one that expresses efficiently, purifies cleanly, or remains stable through formulation and storage. True innovation requires designing molecules not just for foldability, but for process readiness. The next generation of FMs is beginning to meet this demand, embedding manufacturability considerations — solubility, aggregation resistance, expression yield, and stability — into their design objectives.
From Structure Correctness to Process Readiness
Traditional computational design prioritized thermodynamic stability and structural accuracy, metrics that correlate only loosely with manufacturability. A protein that folds perfectly in silico might still aggregate during expression or precipitate in bioreactor media. Similarly, an mRNA optimized for translation could degrade under standard storage or fill–finish conditions. As therapeutic candidates progress toward production, the bottleneck shifts from molecular discovery to process viability.
Modern FMs extend their learning objectives beyond structural correctness to encompass developability. By integrating experimental and simulated data on solubility, viscosity, and expression levels, these models begin to capture the complex interplay between molecular properties and process performance. This shift mirrors a broader trend in the life sciences: the recognition that therapeutic design and manufacturing are not sequential stages but coupled systems. Incorporating manufacturability at the design stage reduces late-stage attrition, lowers costs, and accelerates the transition from bench to bioreactor.8
Manufacturability as a Multi-Objective Optimization
Manufacturability can be viewed as a multi-objective optimization problem spanning the entire development continuum. The same model must balance conflicting constraints: stability versus flexibility, solubility versus binding affinity, expression yield versus post-translational complexity. Modern FMs address this challenge through multitask learning, training simultaneously on data sets that reflect each dimension of developability.
Predictive submodules within these systems can estimate protein expression efficiency in microbial or mammalian hosts, forecast solubility under varying pH or ionic conditions, and model aggregation or viscosity relevant to fill/finish operations. By learning from both successful and failed production campaigns, these models infer patterns that human intuition alone might overlook, such as subtle sequence motifs correlated with high-yield expression or low viscosity formulations.
A critical advantage of FMs is their ability to simulate stress conditions in latent space. By perturbing learned representations to mimic process stresses such as shear, heat, or oxidative environments, they can evaluate the robustness of candidate molecules without the need for exhaustive experimental screening.6 This allows iterative optimization directly within the model: sequences that exhibit favorable predicted performance across these simulated stressors can be prioritized for laboratory validation.
This integration of process awareness into model architecture transforms molecular design into a form of virtual process development. Instead of waiting for empirical feedback from pilot-scale production, manufacturability metrics are embedded from the outset, ensuring that only process-viable candidates advance to experimental testing.
Learning from Mechanical and Physical Properties
The transition from “AI for structure” to “AI for process engineering” is perhaps most visible in generative systems that incorporate mechanical and physical properties into their loss functions. ForceGen is the first generative framework designed to model nonlinear mechanical behaviors, such as elasticity, shear response, and thermal resilience.6 By encoding physical constraints into the design process, ForceGen can generate proteins predicted to resist denaturation and mechanical degradation, critical for large-scale mixing, filtration, and formulation steps. This marks a fundamental expansion of AI’s design remit: from reproducing nature’s structures to inventing molecules engineered for industrial realities.
Another model introduced by Rettie and colleagues takes a related approach in the domain of macrocycles, developing a generative model that co-optimizes binding affinity with physical synthesis feasibility.9 The model identifies scaffolds likely to fold and cyclize efficiently during chemical synthesis, a critical determinant of yield and cost. In doing so, it demonstrates that generative AI can integrate synthetic accessibility — a key manufacturability parameter — into molecular design. Both ForceGen and the macrocycle model exemplify a shift toward designing for production rather than designing in isolation from it.
Further refinement is coming from models that link these macroscopic properties back to atomistic representations. The Local Atomistic Environment framework encodes residue-level physical neighborhoods, capturing the micro-interactions that drive folding kinetics and local stability.7 By blending fine-grained atomistic insight with scalable deep learning, these models help identify structural motifs that improve both functional performance and manufacturability, such as disulfide arrangements that enhance mechanical robustness or surface chemistries that prevent aggregation.
Industrial Implications
For industry, these computational breakthroughs translate directly into efficiency and risk reduction. Embedding manufacturability into design accelerates the lead-to-production transition by ensuring that only candidates with favorable process profiles enter development. This can reduce the number of experimental iterations required to reach a viable formulation and drastically cut the cost of late-stage failure.
Improved predictive power also increases the design hit rate. In traditional workflows, only a small fraction of AI-generated sequences prove both functional and manufacturable. By including process-aware objectives, foundation models elevate that fraction, producing candidates more likely to succeed in experimental validation. This creates a more efficient feedback loop between digital and physical design, with each iteration reinforcing the model’s understanding of what makes a molecule not just active but producible.
Perhaps most transformative is the emergence of manufacturing-aware design pipelines, where digital design, process simulation, and predictive modeling are integrated within a unified framework. Here, manufacturability metrics — solubility, viscosity, yield, formulation stability — are not evaluated post hoc but serve as embedded constraints in the model’s generative logic. This approach could enable in silico design of biotherapeutics tailored to specific production platforms, whether microbial, mammalian, or cell-free.
The downstream implications extend to regulatory readiness as well. By quantifying and documenting manufacturability attributes during early design, developers generate a transparent digital trail of decisions and parameters that can inform comparability and quality assessments later in development.8 In effect, AI-driven design becomes part of the quality-by-design (QbD) framework, linking molecular innovation directly to validated manufacturing outcomes.
The convergence of generative AI and process engineering thus represents more than an incremental advance; it redefines what “design” means in biopharmaceutical development. Foundation models trained for manufacturability are beginning to collapse the traditional separation between R&D and production, turning molecular design into a continuous, data-driven optimization of both biology and process. In the years ahead, as models gain access to richer datasets from bioprocess monitoring and digital twins, their predictive scope will only expand, bringing the vision of fully integrated, intelligent biomanufacturing closer to reality.
Evaluating and Benchmarking the Field
As FMs for biology continue to grow in scale and ambition, the field faces a paradox: predictive power is advancing faster than the ability to measure it. While the past few years have produced an explosion of architectures and pretraining strategies, there remains no consistent framework for evaluating performance across tasks or for validating predictions against real-world outcomes. For industrial and regulatory stakeholders, this is a critical gap. Without reproducible, standardized benchmarks linking computational outputs to wet-lab performance, even the most sophisticated models remain difficult to trust in a manufacturing or clinical context.
Current Evaluation Issues
The most immediate challenge lies in the lack of uniform evaluation metrics. Different groups assess their models using incompatible criteria — perplexity for sequence prediction, TM-score for structural alignment, or stability score for thermodynamic predictions — each capturing a narrow dimension of performance.10 As a result, cross-model comparisons are often unreliable, with apparent gains reflecting data set idiosyncrasies rather than genuine architectural improvements.
This fragmentation extends to data practices as well. Many models are pretrained and fine-tuned on overlapping data sets, leading to hidden data leakage that inflates benchmark scores. Because biological data are highly redundant — homologous sequences recur across multiple databases — standard train-test splits fail to ensure independence, meaning that models can “memorize” evolutionary families instead of generalizing from first principles. The outcome is a proliferation of models that appear powerful on paper but deliver inconsistent results when applied to new proteins or real experimental systems.
Perhaps most consequentially, current metrics seldom connect computational performance with wet-lab relevance. A model that predicts folding with high accuracy may still generate sequences that fail to express or remain soluble in a bioreactor. Bridging this gap requires benchmarks that not only assess accuracy in silico but also track correspondence between predicted and empirical outcomes.
Standardization Efforts
Recognizing these limitations, several initiatives have begun to establish common standards for biological FMs. Among them, ProteinBench stands out as a significant step forward.11 Designed as a community-driven benchmarking platform, it defines unified evaluation tasks across key categories — folding, mutation effect prediction, thermostability, and design diversity — each measured with consistent and interpretable metrics. Importantly, ProteinBench includes standardized data splits that minimize redundancy and enforce strict separation between training and evaluation sets, enabling more meaningful comparisons across models.
Beyond metrics, ProteinBench also promotes transparency in model reporting. Participants are encouraged to disclose data sources, fine-tuning procedures, and any evidence of data leakage or overlap with benchmark sets. This level of openness is essential for ensuring that reported gains reflect genuine improvements in modeling biological complexity rather than artifacts of training protocol. In time, these community standards may form the foundation for regulatory evaluation frameworks, paralleling how validation benchmarks emerged in other AI-heavy domains, such as medical imaging and natural language processing.
Intrinsic-Dimension Studies
A complementary line of work seeks to understand what, exactly, these large biological FMs have learned. Studies have explored the intrinsic dimensionality of embedding spaces within models like ESM and ProtT5, revealing that despite their massive parameter counts, much of the learned information collapses into low-dimensional manifolds encoding a kind of “structural grammar.”12 In practical terms, this means the latent spaces of different models trained independently on dissimilar data sets often converge toward similar organizational principles, implicitly representing evolutionary and structural relationships among proteins.
This discovery suggests that FMs may already approximate an internalized language of biology, with embedding dimensions corresponding to interpretable features such as hydrophobicity, secondary structure, or catalytic motifs. Understanding this latent grammar could yield multiple benefits: model compression (reducing redundancy without performance loss), improved interpretability (linking learned features to physical phenomena), and enhanced transfer learning (reusing compact representations across tasks). For regulators and developers alike, these insights represent progress toward explainable AI, essential for justifying algorithmic design decisions that impact therapeutic outcomes.
The Need for Manufacturability Benchmarks
While structure- and function-focused benchmarks have matured, a crucial next step is expanding evaluation frameworks to include manufacturability metrics. For industrial adoption, models must be validated not only on their ability to predict folding or binding but also on their predictive power for process-relevant parameters such as solubility, expression yield, and formulation stability. As AI-driven design becomes integral to drug development, these process-linked benchmarks will define whether computational models can be trusted to replace parts of the experimental pipeline.
Such benchmarking requires integrating process data sets — thermal shift assays, solubility screens, viscosity data, and bioreactor yield measurements — into open-access platforms. Doing so would create a feedback loop between computational biology and bioprocess engineering: models trained on these datasets could predict manufacturability traits, while experimental results continually refine the benchmarks. This evolution would also align model validation with regulatory expectations, where reproducibility and traceability are prerequisites for approval.
Ultimately, the field must move beyond evaluating how accurately models imitate nature toward assessing how effectively they engineer it. A future iteration of ProteinBench 2.0 might include standardized tasks for predicting stability under stress, aggregation propensity during expression, or long-term storage behavior under controlled conditions. These benchmarks would transform manufacturability from a downstream hurdle into a measurable design target.
By coupling structural, functional, and process-level metrics within unified evaluation frameworks, the next generation of benchmarks will do more than rank models—they will define readiness for translation. In this way, reproducibility, transparency, and benchmarking are not bureaucratic burdens but the infrastructure that will determine whether foundation models can truly reshape how biologics are discovered, designed, and manufactured.10–12
References
1. Queen, Owen, Robert Calef, and Marinka Zitnik. “ProCyon: A Multimodal Foundation Model for Protein Phenotypes.” Kempner Institute. 19 Dec. 2024.
2. Zhang, Li, et al. “ProteinAligner: A Multi-modal Pretraining Framework for Protein Foundation Models.” bioRxiv. 6 Oct. 2024.
3. Li, Michelle M and Marrinka Zitnik. “Context Matters for Foundation Models in Biology.” Kempner Institute. 16 Aug. 2024.
4. Truong, Timothy Fei Jr. and Tristan Bepler. “A multimodal foundation model for controllable protein generation and representation learning.” OpenProtein AI. 11 Feb, 2025.
5. Jendrusch, Michael A, et al. “AlphaDesign: a de novo protein design framework based on AlphaFold.” Mol. Syst. Biol. 21: 1166–1189 (2025).
6. Ni, Bo, David L Kaplan, and Markus J Buehler. “ForceGen: End-to-end de novo protein generation based on nonlinear mechanical unfolding responses using a language diffusion model.” Science Advances. 7 Feb. 2024.
7. Bojan, Meital and Sanketh Vedula. “Representing Local Protein Environments With Atomistic Foundation Models.” Rowan. 20 Jun. 2025.
8. Guo, Fei, et al. “Foundation models in bioinformatics.” National Science Review. 12: nwaf028 (2025).
9. Rettie, Stephen A, et al. “Accurate de novo design of high-affinity protein-binding macrocycles using deep learning.” Nature Chemical Biology. 20 Jun. 2025.
10. Bjerregaard, Andreas, et al. “Foundation models of protein sequences: A brief overview.” Current Opinion in Structural Biology. 91: 103004 (2025).
11. Ye, Fei, et al. “ProteinBench: A Holistic Evaluation of Protein Foundation Models.” Bytedance Research. Accessed 16 Oct. 2025.
12. Yang, Soojung, et al. “Probing the Embedding Space of Protein Foundation Models through Intrinsic Dimension Analysis.” 38th Conference on Neural Information Processing Systems. 2024.












