European Commission’s Joint Research Centre (JRC) examines the rapidly developing field of AI models trained to work with biological information such as DNA, RNA and proteins. Based on a dataset of 480 biological AI models, the research finds that progress is strongest in data-rich areas, particularly protein structure prediction, protein-function annotation and molecular design. These advances are already relevant to drug discovery and other biotechnology applications.
The report highlights a major difference between biological fields: data availability and standardisation strongly influence AI progress. Protein AI has benefited from decades of research and curated repositories such as the Protein Data Bank and UniProt. By contrast, areas such as single-cell biology remain less mature because their datasets are more fragmented and inconsistently structured. The researchers also note that protein models are increasingly using synthetic data generated from AI predictions, creating a growing dependence on predicted rather than exclusively experimental data.
One of the report's most important findings is what it calls a “maturity paradox.” Some models, including AlphaFold and ESM3, are highly mature from a scientific perspective but remain at low-to-mid technology-readiness levels because they have not been comprehensively validated for clinical or industrial deployment. In other words, performing exceptionally well on a scientific benchmark does not automatically mean an AI model is ready for real-world use. The gap creates challenges around validation, governance, regulation and potentially biosecurity, particularly when publicly available models could be misused for activities such as pathogen or toxin engineering.
For Europe, the JRC recommends better biological data infrastructure, stronger coordination among EU researchers, support for European biological foundation models as public goods, and clearer frameworks that evaluate both scientific maturity and technology readiness. It also recommends expanding support for emerging areas such as single-cell biology and multimodal models. The broader message is that the next generation of biological AI will depend not only on bigger models and more computing power, but on high-quality data, collaboration, validation and responsible governance that can turn impressive research systems into reliable real-world technologies.