Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Biological Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Biotechnological solutions for critical mineral recovery from unconventional feedstocks

Secure and sustainable metal recovery from unconventional feedstocks is needed to meet the mineral demands of energy, defense, and electronic technologies. Here, we highlight the potential to leverage nature’s ability to extract and differentiate metal ions in biotechnologies that could become the next generation of mining and refining. We describe bulk and trace processes and then discuss the advances and opportunities of two key bioprocesses: microbially mediated solubilization of metal ions from solid matrices (termed ‘bioleaching’) and bio-based separation of solubilized ions via selective adsorption to proteins. Both biotechnologies have advantages such as reduced energy input for leaching low-grade feedstocks and reduced organic solvent demand for separating ions with similar physiochemical properties but require more development for industrial scale recovery from unconventional feedstocks. Innovation in biological science and engineering may bring timely solutions to key challenges toward recovering critical minerals from unconventional feedstocks.

organic

Gas-Grain Simulation Facility (GGSF)

The goal of the Gas-Grain Simulation Facility project is to provide a microgravity laboratory to facilitate research relevant to exobiology (the study of the origin and evolution of life in the universe). Such a facility will also be useful in other areas of study important to NASA including planetary science, biology, atmospheric science, astrophysics, chemistry, and physics. To achieve this goal, the project will develop and support the GGSF, a modular facility-class payload planned for inclusion on Space Station Freedom. The GGSF will consist of the following: an experiment chamber(s) supported by subsystems that provide chamber environment regulation and monitoring capabilities; sample generation, injection, positioning, and retrieval capabilities; and computer control, data acquisition, and housekeeping capabilities. The facility will also provide analytical tools such as light-scattering measurement systems, aerosol size-spectrum measurement devices, and optical imaging systems.

Greenwald, Ken

Light Microscopy Module: On-Orbit Microscope Planned for the Fluids Integrated Rack on the International Space Station

The Light Microscopy Module (LMM) is planned as a remotely controllable, automated, on-orbit facility, allowing flexible scheduling and control of physical science and biological science experiments within the Fluids Integrated Rack (FIR) on the International Space Station. Initially four fluid physics experiments in the FIR will use the LMM the Constrained Vapor Bubble, the Physics of Hard Spheres Experiment-2, Physics of Colloids in Space-2, and Low Volume Fraction Entropically Driven Colloidal Assembly. The first experiment will investigate heat conductance in microgravity as a function of liquid volume and heat flow rate to determine, in detail, the transport process characteristics in a curved liquid film. The other three experiments will investigate various complementary aspects of the nucleation, growth, structure, and properties of colloidal crystals in microgravity and the effects of micromanipulation upon their properties.

Motil, Susan M.

Light Microscopy Module Fan Disturbance Characterized Through Microgravity Emissions Laboratory Testing

A Light Microscopy Module (LMM) is being engineered, designed, and developed at the NASA Glenn Research Center. The LMM is planned as a remotely controllable on-orbit microscope subrack facility, allowing flexible scheduling and control of physical science and biological science experiments within Glenn s Fluids Integrated Rack on the International Space Station. The LMM concept is a modified commercial research imaging light microscope with powerful laser-diagnostic hardware and interfaces, creating a one-of-a-kind, state-of-the-art microscopic research facility. The microscope will house several different objectives, corresponding to magnifications of 10, 40, 50, 63, and 100. Features of the LMM include high-resolution color video microscopy, brightfield, darkfield, phase contrast, differential interference contrast, spectrophotometry, and confocal microscopy combined in a single configuration. Also, laser tweezers are integrated with the diagnostics as a sample manipulation technique. As part of the development phase of the LMM, it was necessary to quantify the microgravity disturbances generated by the control box fan. Isolating the fan was deemed necessary to reduce the fan speed harmonic amplitudes and to eliminate any broadband disturbances across the 60- to 70-Hz and 160- to 170-Hz frequency ranges. The accelerations generated by a control box fan component of the LMM were measured in the Microgravity Emissions Laboratory (MEL). The MEL is a low-frequency measurement system developed to simulate and verify the on-orbit International Space Station (ISS) microgravity environment. The accelerations generated by various operating components of the ISS, if too large, could hinder the science performed onboard by disturbing the microgravity environment. The MEL facility gives customers a test-verified way of measuring their compliance with ISS limitations on vibratory disturbance levels. The facility is unique in that inertial forces in 6 degrees of freedom can be characterized simultaneously for an operating test article. Vibratory disturbance levels are measured for engineering or flight-level hardware following development from component to subassembly through the rack-level configuration. The MEL can measure accelerations as small as 10-7g, the accuracy needed to confirm compliance with ISS requirements.

McNelis, Anne M.

An MLCommons Scientific Benchmarks Ontology

Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper introduces an ontology for scientific benchmarking developed through a unified, community-driven effort that extends the MLCommons ecosystem to cover physics, chemistry, materials science, biology, climate science, and more. Building on prior initiatives such as XAI-BENCH, FastML Science Benchmarks, PDEBench, and the SciMLBench framework, our effort consolidates a large set of disparate benchmarks and frameworks into a single taxonomy of scientific, application, and system-level benchmarks. New benchmarks can be added through an open submission workflow coordinated by the MLCommons Science Working Group and evaluated against a six-category rating rubric that promotes and identifies high-quality benchmarks, enabling stakeholders to select benchmarks that meet their specific needs. The architecture is extensible, supporting future scientific and AI/ML motifs, and we discuss methods for identifying emerging computing patterns for unique scientific workloads. The MLCommons Science Benchmarks Ontology provides a standardized, scalable foundation for reproducible, cross-domain benchmarking in scientific machine learning. A companion webpage for this work has also been developed as the effort evolves: https://mlcommons-science.github.io/benchmark/

Hawks, Ben [Fermilab] (ORCID:0000000157000288)

Geological applications and training in remote sensing

Some of the experiences, methods, and opinions developed during 15 years of teaching an introductory course in remote sensing at several universities in the Southern California area are related. Although the course is offered in Geology departments, every class includes significant numbers of students from other disciplines including geography, computer science, biology, and environmental science. The instructor or teaching assistant provides a few hours of tutorial lectures (outside of regular class time) on basic geology for these nongeologists. This approach is successful because the grade distribution for nongeologists is similar to that for geologists. The schedule for a typical one-semester course is given.

Sabins, F. F., Jr.

NASA Thesaurus Data File

The NASA Thesaurus contains the authorized NASA subject terms used to index and retrieve materials in the NASA Aeronautics and Space Database (NA&SD) and NASA Technical Reports Server (NTRS). The scope of this controlled vocabulary includes not only aerospace engineering, but all supporting areas of engineering and physics, the natural space sciences (astronomy, astrophysics, planetary science), Earth sciences, and the biological sciences. The NASA Thesaurus Data File contains all valid terms and hierarchical relationships, USE references, and related terms in machine-readable form. The Data File is available in the following formats: RDF/SKOS, RDF/OWL, ZThes-1.0, and CSV/TXT.

Source record

GeneLab for High Schools – Bioinformatic Training For Students And Educators

Modern biological sciences are increasingly based on high-throughput molecular techniques, including genomics, transcriptomics, and proteomics. NASA’s GeneLab program has collected extensive data from ‘omics’ studies, curated them into an accessible platform and provided data analysis/visualization tools to facilitate the generation of new hypotheses and research directions. GeneLab for High Schools (GL4HS), launched in 2017, has endeavored to utilize this database and provide tools for students to understand and analyze omics datasets whilst also learning about spaceflight research. The GL4HS program ran in person at Ames from 2017-2019 and has run virtually since 2020. Each year fifteen high school students are trained to analyze and interpret GeneLab transcriptomic data. Additionally, in the last several years we have expanded our “teacher training program” to include 10 teachers total in an effort to enable this program to be utilized in classrooms across the USA. Teachers also join the NASA GeneLab Education Working Group (EWG) enabling support as they implement custom GL4HS modules into their classrooms. The GL4HS program consists of three main components – (1) core learning modules, (2) networking and teamwork, and (3) an independent learning project. Students are also taught critical networking and science communication skills facilitating their ability to ‘sell their science’ in innovative and creative ways. This program has enabled students to learn about biology in space and to have a glimpse into the world of research for the first time. Many of the students in this program shared that the course was transformative to their perception about biological sciences and how it linked to other areas of STEM. The ultimate and long-term goal of GL4HS is to expand the program to multiple locations thereby facilitating the reach of NASA Space Biology beyond NASA-centric regions.

GeneLab

Current and future directions in network biology

Network biology is an interdisciplinary field bridging computational and biological sciences that has proved pivotal in advancing the understanding of cellular functions and diseases across biological systems and scales. Although the field has been around for two decades, it remains nascent. It has witnessed rapid evolution, accompanied by emerging challenges. These stem from various factors, notably the growing complexity and volume of data together with the increased diversity of data types describing different tiers of biological organization. We discuss prevailing research directions in network biology, focusing on molecular/cellular networks but also on other biological network types such as biomedical knowledge graphs, patient similarity networks, brain networks, and social/contact networks relevant to disease spread. In more detail, we highlight areas of inference and comparison of biological networks, multimodal data integration and heterogeneous networks, higher-order network analysis, machine learning on networks, and network-based personalized medicine. Following the overview of recent breakthroughs across these five areas, we offer a perspective on future directions of network biology. Additionally, we discuss scientific communities, educational initiatives, and the importance of fostering diversity within the field. This article establishes a roadmap for an immediate and long-term vision for network biology.

59 BASIC BIOLOGICAL SCIENCES

Bio-Nanotechnology: Challenges for Trainees in a Multidisciplinary Research Program

The recent developments in the field of nanotechnology have provided scientists with a new set of nanoscale materials, tools and devices in which to investigate the biological science thus creating the mulitdisciplinary field of bio-nanotechnology. Bio-nanotechnology merges the biological sciences with other scientific disciplines ranging from chemistry to engineering. Todays students must have a working knowledge of a variety of scientific disciplines in order to be successful in this new field of study. This talk will provide insight into the issue of multidisciplinary education from the perspective of a graduate student working in the field of bio-nanotechnology. From the classes we take to the research we perform, how does the modern graduate student attain the training required to succeed in this field?

Koehne, Jessica Erin

National Aeronautics and Space Administration Biological and Physical Research Enterprise Strategy

As the 21st century begins, NASA's new Vision and Mission focuses the Agency's Enterprises toward exploration and discovery.The Biological and Physical Research Enterprise has a unique and enabling role in support of the Agency's Vision and Mission. Our strategic research seeks innovations and solutions to enable the extension of life into deep space safely and productively. Our fundamental research, as well as our research partnerships with industry and other agencies, allow new knowledge and tech- nologies to bring improvements to life on Earth. Our interdisciplinary research in the unique laboratory of microgravity addresses opportunities and challenges on our home planet as well as in space environments. The Enterprise maintains a key role in encouraging and engaging the next generation of explorers from primary school through the grad- uate level via our direct student participation in space research.The Biological and Physical Research Enterprise encompasses three themes. The biological sciences research theme investigates ways to support a safe human presence in space. This theme addresses the definition and control of physiological and psychological risks from the space environment, including radiation,reduced gravity, and isolation. The biological sciences research theme is also responsible for the develop- ment of human support systems technology as well as fundamental biological research spanning topics from genomics to ecologies. The physical sciences research theme supports research that takes advantage of the space environment to expand our understanding of the fundamental laws of nature. This theme also supports applied physical sciences research to improve safety and performance of humans in space. The research partnerships and flight support theme establishes policies and allocates space resources to encourage and develop entrepreneurial partners access to space research.Working together across research disciplines, the Biological and Physical Research Enterprise is performing vital research and technology development to extend the reach of human space flight.

Source record

Space Station Biological Research Project

To meet NASA's objective of using the unique aspects of the space environment to expand fundamental knowledge in the biological sciences, the Space Station Biological Research Project at Ames Research Center is developing, or providing oversight, for two major suites of hardware which will be installed on the International Space Station (ISS). The first, the Gravitational Biology Facility, consists of Habitats to support plants, rodents, cells, aquatic specimens, avian and reptilian eggs, and insects and the Habitat Holding Rack in which to house them at microgravity; the second, the Centrifuge Facility, consists of a 2.5 m diameter centrifuge that will provide acceleration levels between 0.01 g and 2.0 g and a Life Sciences Glovebox. These two facilities will support the conduct of experiments to: 1) investigate the effect of microgravity on living systems; 2) what level of gravity is required to maintain normal form and function, and 3) study the use of artificial gravity as a countermeasure to the deleterious effects of microgravity observed in the crew. Upon completion, the ISS will have three complementary laboratory modules provided by NASA, the European Space Agency and the Japanese space agency, NASDA. Use of all facilities in each of the modules will be available to investigators from participating space agencies. With the advent of the ISS, space-based gravitational biology research will transition from 10-16 day short-duration Space Shuttle flights to 90-day-or-longer ISS increments.

NASA Discipline General Space Life Sciences

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2019 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 436 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Location Identifiers, Metadata, and Map for Field Measurements at the East-Taylor Watershed Community Observatory, Colorado, USA (Version 3.3)

This dataset contains identifiers, metadata, and a map of the locations where field measurements have been conducted at the East-Taylor Watershed Community Observatory located in the Upper Colorado River Basin, United States. This is version 3.3 of the dataset and replaces the prior version 3.2 (see below for details on changes between the versions). Dataset description: The East River-Taylor Watershed is the primary field site of the Watershed Function Scientific Focus Area (WFSFA) and the Rocky Mountain Biological Laboratory. Researchers from several institutions generate highly diverse hydrological, biogeochemical, climate, vegetation, geological, remote sensing, and model data at the East-Taylor Watershed in collaboration with the WFSFA. Thus, the purpose of this dataset is to maintain an inventory of the field locations and instrumentation to provide information on the field activities in the East-Taylor Watershed and coordinate data collected across different locations, researchers, and institutions. The dataset contains (1) a README file with information on the various files, (2) three csv files describing the metadata collected for each surface point location, plot and region registered with the WFSFA, (3) csv files with metadata and contact information for each surface point location registered with the WFSFA, (4) a csv file with with metadata and contact information for plots, (5) a csv file with metadata for geographic regions and sub-regions within the watershed, (6) a compiled xlsx file with all the data and metadata which can be opened in Microsoft Excel, (7) a kml map of the locations plotted in the watershed which can be opened in Google Earth, (8) a jpg image of the kml map which can be viewed in any photo viewer, and (9) a zipped file with the registration templates used by the SFA team to collect location metadata. The zipped template file contains two csv files with the blank templates (point and plot), two csv files with instructions for filling out the location templates, and one compiled xlsx file with the instructions and blank templates together. Additionally, the templates in the xlsx include drop down validation for any controlled metadata fields. Persistent location identifiers (Location_ID) are determined by the WFSFA data management team and are used to track data and samples across locations. Dataset uses: This location metadata is used to update the Watershed SFA’s publicly accessible Field Information Portal (an interactive field sampling metadata exploration tool; https://wfsfa-data.lbl.gov/watershed/), the kml map file included in this dataset, and other data management tools internal to the Watershed SFA team. Version Information: The latest version of this dataset publication is version 3.3. This version contains 167 new point locations, 1 new plot, and 2 new geographic regions. Overall, there are a total of 1439 point locations, 75 plots, and 54 geographic regions. Additionally, the kml map of locations and image now includes two boundaries (Upper Ohio Creek (UO) and Carbon Creek (CA)) outside of the East River watershed (USGS HUC-10) and accompanying stream network that represents areas of focus. Refer to methods for further details on the version history. This dataset will be updated on a periodic basis with new measurement location information. Researchers interested in having their East-Taylor Watershed measurement locations added to this list should reach out to the WFSFA data management team at wfsfa-data@googlegroups.com. Acknowledgments: Please cite this dataset if using any of the location metadata in other publications or derived products. If using the location metadata for the 2018 NEON hyperspectral campaign, additionally cite Chadwick et al. (2020). doi:10.15485/1618130. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

2018 NEON and 2025 CHESS Campaigns

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence