Engineering PapersSearch

DOE OSTI · 3374411

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Abel, Jr., Gary R. [Fourth Eon Bio, San Diego, CA (United States); Johns Hopkins Univ., Baltimore, MD (United States)]·Alexanian, Tessa [International Biosecurity and Biosafety Initiative for Science (IBBIS), Geneva (Switzerland)]·Bartling, Craig [Battelle Memorial Institute, Columbus, OH (United States)]·Beal, Jacob [RTX BBN Technologies, Cambridge, MA (United States)]·Curtis, Samuel [National Institute of Standards and Technology (NIST), Washington, DC (United States). Center for AI Standards and Innovation (CAISI)]·Flyangolts, Kevin [Aclid, New York, NY (United States)]·Foner, Leonard [SecureDNA, Basel (Switzerland)]·Forry, Samuel P. [National Institute of Standards and Technology (NIST), Gaithersburg, MD (United States)]·Godbold, Gene D. [Signature Science LLC, Charlottesville, VA (United States)]·Horvitz, Eric [Microsoft Corporation, Redmond, WA (United States)]·Hu, Bin [Los Alamos National Laboratory (LANL), Los Alamos, NM (United States)] (ORCID:0000000202788466)·Hudson, Corey M. [The Align Foundation, Covina, CA (United States)]·Jagla, Caitlin [RTX BBN Technologies, Cambridge, MA (United States)]·Lababidi, Rassin [International Biosecurity and Biosafety Initiative for Science (IBBIS), Geneva (Switzerland)]·Lin-Gibson, Sheng [National Institute of Standards and Technology (NIST), Gaithersburg, MD (United States)]·Magalis, Brittany Rife [Univ. of Louisville, KY (United States)]·Pannu, Jaspreet [Johns Hopkins Univ., Baltimore, MD (United States)]·Rivera, Sebastian [Engineering Biology Research Consortium, Emeryville, CA (United States)]·Ross, David [National Institute of Standards and Technology (NIST), Gaithersburg, MD (United States)]·Wittmann, Bruce J. [Microsoft Corporation, Redmond, WA (United States)]·Diggans, James [Twist Bioscience, San Francisco, CA (United States); International Gene Synthesis Consortium, Emeryville, CA (United States)]

Abstract

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Abel, Jr., Gary R. [Fourth Eon Bio, San Diego, CA (United States); Johns Hopkins Univ., Baltimore, MD (United States)], Alexanian, Tessa [International Biosecurity and Biosafety Initiative for Science (IBBIS), Geneva (Switzerland)], Bartling, Craig [Battelle Memorial Institute, Columbus, OH (United States)], Beal, Jacob [RTX BBN Technologies, Cambridge, MA (United States)], Curtis, Samuel [National Institute of Standards and Technology (NIST), Washington, DC (United States). Center for AI Standards and Innovation (CAISI)], Flyangolts, Kevin [Aclid, New York, NY (United States)], Foner, Leonard [SecureDNA, Basel (Switzerland)], Forry, Samuel P. [National Institute of Standards and Technology (NIST), Gaithersburg, MD (United States)], Godbold, Gene D. [Signature Science LLC, Charlottesville, VA (United States)], Horvitz, Eric [Microsoft Corporation, Redmond, WA (United States)], Hu, Bin [Los Alamos National Laboratory (LANL), Los Alamos, NM (United States)] (ORCID:0000000202788466), Hudson, Corey M. [The Align Foundation, Covina, CA (United States)], Jagla, Caitlin [RTX BBN Technologies, Cambridge, MA (United States)], Lababidi, Rassin [International Biosecurity and Biosafety Initiative for Science (IBBIS), Geneva (Switzerland)], Lin-Gibson, Sheng [National Institute of Standards and Technology (NIST), Gaithersburg, MD (United States)], Magalis, Brittany Rife [Univ. of Louisville, KY (United States)], Pannu, Jaspreet [Johns Hopkins Univ., Baltimore, MD (United States)], Rivera, Sebastian [Engineering Biology Research Consortium, Emeryville, CA (United States)], Ross, David [National Institute of Standards and Technology (NIST), Gaithersburg, MD (United States)], Wittmann, Bruce J. [Microsoft Corporation, Redmond, WA (United States)], Diggans, James [Twist Bioscience, San Francisco, CA (United States); International Gene Synthesis Consortium, Emeryville, CA (United States)]. 2026-05-14. Beyond sequence similarity: toward function-based screening of nucleic acid synthesis. https://doi.org/10.3389/fbioe.2026.1832724

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES

Genome-resolved insights into microbial diversity and elemental cycling in Winogradsky columns

We retained 18 MAGs with ≥50% completion and <10% contamination (i.e., at least medium quality). Of these, 10 had >90% completion and <5% contamination; however, only one (Paceibacteria Bin.003_MG) can be described as high-quality, as the others lacked a full suite of 5S, 16S, and 23S rRNA genes. To maximize the diversity of our recovered MAGs, we also retained one MAG (Chromatiaceae Bin.008_AM) with >40% (but less than 50%) completion and <5% contamination, as well as one (Rhodopseudomonas Bin.015_MK) with >90% completion and <20% (but>10%) contamination. Interestingly, significant chimerism was not detected in this MAG (40) , suggesting that the elevated contamination (20%) may instead reflect two closely related strains collapsing into a single bin. Consistent with this, contig coverage was bimodal, with roughly 17% of the assembly at ~115x and the remaining 83% at ~282x, while GC content remained uniform across both groups (~64%), arguing against contamination from a taxonomically distinct source.

59 BASIC BIOLOGICAL SCIENCES