Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

CAML: Commutative Algebra Machine Learning─A Case Study on Protein–Ligand Binding Affinity Prediction

Recently, Suwayyid and Wei introduced commutative algebra as an emerging paradigm for machine learning and data science. In this work, we propose commutative algebra machine learning (CAML) for the prediction of protein−ligand binding affinities. Specifically, we apply persistent Stanley−Reisner theory, a key concept in combinatorial commutative algebra, to the affinity predictions of protein−ligand binding and metalloprotein−ligand binding. We present three new algorithms, i.e., element-specific commutative algebra, category-specific commutative algebra, and commutative algebra on bipartite complexes, to tackle the complexity of data involved in (metallo) protein−ligand complexes. We show that the proposed CAML outperforms other state-of-theart methods in (metallo) protein−ligand binding affinity predictions, indicating the great potential of commutative algebra learning.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recent progress on the mesoscale modeling of architected thin-films via phase-field formulations of physical vapor deposition

Thin-film coatings can be found everywhere in modern technological applications due to desirable electrical, mechanical, chemical, and optical properties. These properties directly depend upon the thin-film’s microstructural features, which are themselves influenced by the materials and vapor-deposition processing conditions used for fabrication. As such, understanding processing-microstructure relationships is essential to designing thin-films with optimized properties, and discovering new processing conditions that allow for novel thin-films with multifunctional microstructures. Here, a short review is presented on recent developments that utilize the phase-field method to simultaneously model the vapor-deposition process and corresponding microstructure formation at the mesoscale. Also phase-field-based vapor-deposition models that simulate thin-film growth of immiscible alloy and polycrystalline systems are highlighted in addition to machine-learning-based surrogate models that can facilitate accelerated high-fidelity simulations along with materials design and exploration studies.

36 MATERIALS SCIENCE↗

Selecting Post-Processing Schemes for Accurate Detection of Small Objects in Low-Resolution Wide-Area Aerial Imagery

In low-resolution wide-area aerial imagery, object detection algorithms are categorized as feature extraction and machine learning approaches, where the former often requires a post-processing scheme to reduce false detections and the latter demands multi-stage learning followed by post-processing. In this paper, we present an approach on how to select post-processing schemes for aerial object detection. We evaluated combinations of each of ten vehicle detection algorithms with any of seven post-processing schemes, where the best three schemes for each algorithm were determined using average F-score metric. The performance improvement is quantified using basic information retrieval metrics as well as the classification of events, activities and relationships (CLEAR) metrics. We also implemented a two-stage learning algorithm using a hundred-layer densely connected convolutional neural network for small object detection and evaluated its degree of improvement when combined with the various post-processing schemes. The highest average F-scores after post-processing are 0.902, 0.704 and 0.891 for the Tucson, Phoenix and online VEDAI datasets, respectively. The combined results prove that our enhanced three-stage post-processing scheme achieves a mean average precision (mAP) of 63.9% for feature extraction methods and 82.8% for the machine learning approach.

54 ENVIRONMENTAL SCIENCES↗

Elucidating Abnormal Grain Growth in Thermomagnetic Processed Materials with Transfer Learning and Reinforcement Learning

The goal of this research program is to establish the mechanism governing local grain boundary motion, which is needed to design and process desirable microstructures for better performance, by identifying the relative contributions of grain boundary (GB) energy and mobility to grain growth. Classical models for grain growth assume that the primary mechanism for reducing the total interfacial energy is area reduction and that GB restructuring is not significant. This assumption implies that grain growth is locally driven by curvature. However, recent experimental observations using new non-destructive 3D x-ray diffraction microscopy techniques (3D-XRM) reveal that classic descriptors (i.e., curvature, number of neighbors, grain size) do not predict real grain growth. Instead, local GB motion appears to be governed by its energy relative to its neighbors such that low-energy boundaries replace those of higher energy. However, simulations that incorporate GB energy anisotropy still fail to reproduce these observations. These discrepancies suggest that the common assumption for grain growth theory must be re-examined to predict and, thus, control microstructure evolution in real polycrystals. A significant challenge to testing this assumption is due to anisotropic GB mobility. Mobility may cause abnormal grain growth or affect the final grain shapes or growth rate but its true contributions are unknown because it is difficult to measure. For example, observations in Fe have found that grains associated with high energy and high mobility boundaries tend to experience abnormal grain growth, whereas abnormal grain growth is associated with low energy and high mobility boundaries in alumina. As mobility and energy both control GB motion, it is challenging to isolate the local driving forces necessary to test the common assumption that the primary mechanism is area reduction. The novelty of this work is the use of machine learning tools to capture GB mobility and energy from 3D-XRM measurements in polycrystals to test the common assumption used in grain growth models. Machine learning can capture high-order correlations in dynamic systems like those found in the evolving GB topology. The PIs have developed a physics-regularized interpretable machine learning microstructure evolution (PRIMME) model that accurately replicates the grain growth behavior of its trained data set.

36 MATERIALS SCIENCE↗

Materials representation and transfer learning for multi-property prediction

The adoption of machine learning in materials science has rapidly transformed materials property prediction. Hurdles limiting full capitalization of recent advancements in machine learning include the limited development of methods to learn the underlying interactions of multiple elements as well as the relationships among multiple properties to facilitate property prediction in new composition spaces. To address these issues, we introduce the Hierarchical Correlation Learning for Multi-property Prediction (H-CLMP) framework that seamlessly integrates: (i) prediction using only a material's composition, (ii) learning and exploitation of correlations among target properties in multi-target regression, and (iii) leveraging training data from tangential domains via generative transfer learning. The model is demonstrated for prediction of spectral optical absorption of complex metal oxides spanning 69 three-cation metal oxide composition spaces. H-CLMP accurately predicts non-linear composition-property relationships in composition spaces for which no training data are available, which broadens the purview of machine learning to the discovery of materials with exceptional properties. This achievement results from the principled integration of latent embedding learning, property correlation learning, generative transfer learning, and attention models. The best performance is obtained using H-CLMP with transfer learning [H-CLMP(T)] wherein a generative adversarial network is trained on computational density of states data and deployed in the target domain to augment prediction of optical absorption from composition. H-CLMP(T) aggregates multiple knowledge sources with a framework that is well suited for multi-target regression across the physical sciences.

36 MATERIALS SCIENCE↗

Reformulation of the No-Free-Lunch Theorem for Entangled Datasets

The No-Free-Lunch (NFL) theorem is a celebrated result in learning theory that limits one’s ability to learn a function with a training data set. With the recent rise of quantum machine learning, it is natural to ask whether there is a quantum analog of the NFL theorem, which would restrict a quantum computer’s ability to learn a unitary process with quantum training data. However, in the quantum setting, the training data can possess entanglement, a strong correlation with no classical analog. In this work, we show that entangled data sets lead to an apparent violation of the (classical) NFL theorem. This motivates a reformulation that accounts for the degree of entanglement in the training set. As our main result, we prove a quantum NFL theorem whereby the fundamental limit on the learnability of a unitary is reduced by entanglement. We employ Rigetti's quantum computer to test both the classical and quantum NFL theorems. In conclusion, our work establishes that entanglement is a commodity in quantum machine learning.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine learning sheds light on microbial dark proteins

In this article, metagenomics projects have revealed more than 8 billion non-redundant microbial protein sequences from across the Earth’s biosphere. Of these, 1.17 billion proteins do not have recognizable homologues in any of the more than 100,000 reference genomes available1. Understanding the function of these microbial proteins is a daunting task. Fortunately, machine learning has recently achieved unprecedented accuracy in modelling complex biological data and making predictions. At the forefront of these advancements are machine learning-based approaches that can confidently predict atomic-level protein structures for many (but not all) amino acid sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Computational Framework for Machine-Learning-Enabled 13 C Fluxomics

13 C metabolic flux analysis (MFA) has emerged as a powerful tool for synthetic biology. This optimization-based approach suffers long computation time and unstable solutions depending on the initial guess. Here, we develop a machine-learning-based framework for 13 C fluxomics. Specifically, training and test data sets are generated by metabolic network decomposition and flux sampling, in which flux ratios at metabolic nodes and simulated labeling patterns of metabolites are used as training targets and features, respectively. To improve prediction accuracy and simplify the model, automated processes are developed for flux ratio selection based on solvability and feature screening based on importance. We found that predictive performance can be significantly improved using both amino acids and central carbon metabolites in comparison with amino acids alone. Together with measured external fluxes, the predicted flux ratios determine the mass balance system, yielding global flux distributions. This approach is validated by flux estimation using both simulated and experimental data in comparison with canonical 13 C MFA. The approach represents a reliable fluxomics method readily applicable to high-throughput metabolic phenotyping, which highlights the advances of intelligent learning algorithms in synthetic biology, specifically in the Test and Learn stage of the Design-Build-Test-Learn cycle.

13C metabolic flux analysis↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Benchmarking a Tunable Quantum Neural Network on Trapped-Ion and Superconducting Hardware

We implement a quantum generalization of a neural network on trapped-ion and IBM superconducting quantum computers to classify MNIST images, a common benchmark in computer vision. The network feedforward involves qubit rotations whose angles depend on the results of measurements in the previous layer. The network is trained via simulation, but inference is performed experimentally on quantum hardware. The classical-to-quantum correspondence is controlled by an interpolation parameter, $a$, which is zero in the classical limit. Increasing $a$ introduces quantum uncertainty into the measurements, which is shown to improve network performance at moderate values of the interpolation parameter. We then focus on particular images that fail to be classified by a classical neural network but are detected correctly in the quantum network. For such borderline cases, we observe strong deviations from the simulated behavior. We attribute this to physical noise, which causes the output to fluctuate between nearby minima of the classification energy landscape. Such strong sensitivity to physical noise is absent for clear images. We further benchmark physical noise by inserting additional single-qubit and two-qubit gate pairs into the neural network circuits. Our work provides a springboard toward more complex quantum neural networks on current devices: while the approach is rooted in standard classical machine learning, scaling up such networks may prove classically non-simulable and could offer a route to near-term quantum advantage.

FOS: Physical sciences↗

Good practices for documenting AI-based studies on energy and buildings

Artificial intelligence has transformed building science research over the past decade, with applications spanning energy modeling, energy prediction, HVAC optimization and controls, fault detection, and occupancy modeling. However, many studies lack adequate documentation of datasets, algorithms, training procedures, and validation methods. Building science research faces additional challenges including inconsistent evaluation metrics, limited generalizability across building types, climates, and significant gaps between experimental studies and deployed systems. This communication provides practical guidance for good practices in documenting and publishing AI-based research following established standards from the computer science and machine learning communities. By adopting frameworks such as Datasheets for Datasets, Model Cards, and standardized reproducibility checklists, researchers can ensure their work meets the rigorous documentation standards necessary for reproducible, comparable, and impactful building science research.

Hong, Tianzhen [Lawrence Berkeley National Laborat↗

FracML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage

Poster on “FRACML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage” for the CCUS 2025 conference held in Houston, Texas March 3-5, 2025. The accurate characterization of subsurface fracture networks is essential for the secure operation of carbon capture, utilization, and storage (CCUS) projects. A thorough understanding of the spatial distribution of subsurface faults and fractures is crucial for predicting CO2 plume evolution and minimizing risks such as potential leakage into overlying formations or induced seismicity. In this context, robust fracture network quantification plays a pivotal role in reservoir management, providing the data necessary to fine-tune operational parameters, and ensure the environmental and economic viability of CCUS projects. As part of the U.S. Department of Energy’s SMART (Science-informed Machine Learning for Accelerating Real-time Decisions in Subsurface Applications) initiative, we focused on the development and application of a machine learning-based tool (FRACML) designed to quantify and map fracture networks using real-world (non-synthetic) data from an active CO2 injection site. Our objective is to demonstrate the utility of this tool in improving operational efficiency and safety across CCUS sites.

artifical intelligence / machine learning (AI/ML)↗

Protein sequence design with a learned potential

The task of protein sequence design is central to nearly all rational protein engineering problems, and enormous effort has gone into the development of energy functions to guide design. Here, we investigate the capability of a deep neural network model to automate design of sequences onto protein backbones, having learned directly from crystal structure data and without any human-specified priors. The model generalizes to native topologies not seen during training, producing experimentally stable designs. We evaluate the generalizability of our method to a de novo TIM-barrel scaffold. The model produces novel sequences, and high-resolution crystal structures of two designs show excellent agreement with in silico models. Our findings demonstrate the tractability of an entirely learned method for protein sequence design.

59 BASIC BIOLOGICAL SCIENCES↗

Supervised Machine Learning Approach for Classifying Earth Science Publications

The data collections archived and distributed by the GES DISC NASA data center are widely utilized for various Earth Science studies. As these collections are created, many research works are published regarding these collections' algorithms, their validation, and their applications. As NASA data centers collect these publications for public use, it is helpful to categorize them based on how they relate to their associated datasets. Specifically, whether the publication linked to the GES DISC dataset is using it for applicational research, describing the algorithm used for the dataset creation, validating the dataset, or providing a general overview of the data collection. Currently, this process requires simple manual labeling, and as such, it may be possible to solve via automation. To approach this problem, machine learning classifiers were developed to predict a publication's category. Manually labeled publications were used as the training data for the supervised machine learning algorithms, specifically Random Forest and Multinomial Naïve Bayes. After balancing the dataset and implementing the Multinomial Naïve Bayes algorithm, the classification accuracy achieved was substantially higher than the baseline accuracy, thus significantly improving the efficiency of publication labeling.

Rohan Dayal↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Enhanced Surgical Decision-Making Tools in Breast Cancer: Predicting 2-Year Postoperative Physical, Sexual, and Psychosocial Well-Being following Mastectomy and Breast Reconstruction (INSPiRED 004)

Background: We sought to predict clinically meaningful changes in physical, sexual, and psychosocial well-being for women undergoing cancer-related mastectomy and breast reconstruction 2 years after surgery using machine learning (ML) algorithms trained on clinical and patient-reported outcomes data. Patients and Methods: We used data from women undergoing mastectomy and reconstruction at 11 study sites in North America to develop three distinct ML models. We used data of ten sites to predict clinically meaningful improvement or worsening by comparing pre-surgical scores with 2 year follow-up data measured by validated Breast-Q domains. We employed ten-fold cross-validation to train and test the algorithms, and then externally validated them using the 11th site’s data. We considered area-under-the-receiver-operating-characteristics-curve (AUC) as the primary metric to evaluate performance. Results: Overall, between 1454 and 1538 patients completed 2 year follow-up with data for physical, sexual, and psychosocial well-being. In the hold-out validation set, our ML algorithms were able to predict clinically significant changes in physical well-being (chest and upper body) (worsened: AUC range 0.69–0.70; improved: AUC range 0.81–0.82), sexual well-being (worsened: AUC range 0.76–0.77; improved: AUC range 0.74–0.76), and psychosocial well-being (worsened: AUC range 0.64–0.66; improved: AUC range 0.66–0.66). Baseline patient-reported outcome (PRO) variables showed the largest influence on model predictions. Conclusions: Machine learning can predict long-term individual PROs of patients undergoing postmastectomy breast reconstruction with acceptable accuracy. This may better help patients and clinicians make informed decisions regarding expected long-term effect of treatment, facilitate patient-centered care, and ultimately improve postoperative health-related quality of life.

60 APPLIED LIFE SCIENCES↗

Modeling Spatial Distribution of Snow Water Equivalent by Combining Meteorological and Satellite Data with Lidar Maps

Abstract An accurate characterization of the water content of snowpack, or snow water equivalent (SWE), is necessary to quantify water availability and constrain hydrologic and land surface models. Recently, airborne observations (e.g., lidar) have emerged as a promising method to accurately quantify SWE at high resolutions (scales of ∼100 m and finer). However, the frequency of these observations is very low, typically once or twice per season in the Rocky Mountains of Colorado. Here, we present a machine learning framework that is based on random forests to model temporally sparse lidar-derived SWE, enabling estimation of SWE at unmapped time points. We approximated the physical processes governing snow accumulation and melt as well as snow characteristics by obtaining 15 different variables from gridded estimates of precipitation, temperature, surface reflectance, elevation, and canopy. Results showed that, in the Rocky Mountains of Colorado, our framework is capable of modeling SWE with a higher accuracy when compared with estimates generated by the Snow Data Assimilation System (SNODAS). The mean value of the coefficient of determination R 2 using our approach was 0.57, and the root-mean-square error (RMSE) was 13 cm, which was a significant improvement over SNODAS (mean R 2 = 0.13; RMSE = 20 cm). We explored the relative importance of the input variables and observed that, at the spatial resolution of 800 m, meteorological variables are more important drivers of predictive accuracy than surface variables that characterize the properties of snow on the ground. This research provides a framework to expand the applicability of lidar-derived SWE to unmapped time points. Significance Statement Snowpack is the main source of freshwater for close to 2 billion people globally and needs to be estimated accurately. Mountainous snowpack is highly variable and is challenging to quantify. Recently, lidar technology has been employed to observe snow in great detail, but it is costly and can only be used sparingly. To counter that, we use machine learning to estimate snowpack when lidar data are not available. We approximate the processes that govern snowpack by incorporating meteorological and satellite data. We found that variables associated with precipitation and temperature have more predictive power than variables that characterize snowpack properties. Our work helps to improve snowpack estimation, which is critical for sustainable management of water resources.

54 ENVIRONMENTAL SCIENCES↗