Engineering PapersSearch

SEARCH · Engineering Papers

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Discovering type I cis-AT polyketides through computational mass spectrometry and genome mining with Seq2PKS

Type 1 polyketides are a major class of natural products used as antiviral, antibiotic, antifungal, antiparasitic, immunosuppressive, and antitumor drugs. Analysis of public microbial genomes leads to the discovery of over sixty thousand type 1 polyketide gene clusters. However, the molecular products of only about a hundred of these clusters are characterized, leaving most metabolites unknown. Characterizing polyketides relies on bioactivity-guided purification, which is expensive and time-consuming. To address this, we present Seq2PKS, a machine learning algorithm that predicts chemical structures derived from Type 1 polyketide synthases. Seq2PKS predicts numerous putative structures for each gene cluster to enhance accuracy. The correct structure is identified using a variable mass spectral database search. Benchmarks show that Seq2PKS outperforms existing methods. Applying Seq2PKS to Actinobacteria datasets, we discover biosynthetic gene clusters for monazomycin, oasomycin A, and 2-aminobenzamide-actiphenol.

60 APPLIED LIFE SCIENCES

AI‐Driven Defect Engineering for Advanced Thermoelectric Materials

Thermoelectric materials offer a promising pathway to directly convert waste heat to electricity. However, achieving high performance remains challenging due to intrinsic trade-offs between electrical conductivity, the Seebeck coefficient, and thermal conductivity, which are further complicated by the presence of defects. This review explores how artificial intelligence (AI) and machine learning (ML) are transforming thermoelectric materials design. Advanced ML approaches including deep neural networks, graph-based models, and transformer architectures, integrated with high-throughput simulations and growing databases, effectively capture structure-property relationships in a complex multiscale defect space and overcome the “curse of dimensionality”. This review discusses AI-enhanced defect engineering strategies such as composition optimization, entropy and dislocation engineering, and grain boundary design, along with emerging inverse design techniques for generating materials with targeted properties. Finally, it outlines future opportunities in novel physics mechanisms and sustainability, highlighting the critical role of AI in accelerating the discovery of thermoelectric materials.

36 MATERIALS SCIENCE

Application of machine learning to discover new intermetallic catalysts for the hydrogen evolution and the oxygen reduction reactions

The adsorption energies for hydrogen, oxygen, and hydroxyl were calculated by means of density functional theory on the lowest energy surface of 24 pure metals and 332 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B taking into account the effect of biaxial elastic strains. This information was used to train two random forest regression models, one for the hydrogen adsorption and another for the oxygen and hydroxyl adsorption, based on 9 descriptors that characterized the geometrical and chemical features of the adsorption site as well as the applied strain. All the descriptors for each compound in the models could be obtained from physico-chemical databases. The random forest models were used to predict the adsorption energy for hydrogen, oxygen, and hydroxyl of ≈2700 binary intermetallic compounds with stoichiometries AB, A 2 B, and A 3 B made of metallic elements, excluding those that were environmentally hazardous, radioactive, or toxic. This information was used to search for potential good catalysts for the HER and ORR from the criteria that their adsorption energy for H and O/OH, respectively, should be close to that of Pt. Further, this investigation shows that the suitably trained machine learning models can predict adsorption energies with an accuracy not far away from density functional theory calculations with minimum computational cost from descriptors that are readily available in physico-chemical databases for any compound. Moreover, the strategy presented in this paper can be easily extended to other compounds and catalytic reactions, and is expected to foster the use of ML methods in catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Locating Undocumented Wells Using Historical Oil and Gas Exploration Maps: A Case Study in Osage County, Oklahoma

Undocumented oil and gas wells lack reliable information about their locations and characteristics, making them difficult to identify. These wells can result in unanticipated delays and costs in the development of nearby surface and subsurface resources, and, if improperly plugged, can cause contamination. This study leverages historical petroleum exploration maps to locate such wells, focusing on Osage County, Oklahoma. Two sets of early 20th century oil and gas exploration maps by the United States Geological Survey were georeferenced and analyzed using a computer vision model to detect well symbols. The locations of detected wells were compared to the location of known wells in the database from the Bureau of Indian Affairs Osage Agency to identify potential undocumented wells. The analysis yielded over 500 potential undocumented wells, with dry holes constituting the largest fraction. Field verification confirmed the presence of some undocumented wells. Comparison with prior work revealed limited overlap, underscoring the complementary value of historical oil and gas maps for locating undocumented wells. This approach demonstrates the utility of integrating historical cartographic resources with modern geospatial and machine learning techniques to improve the identification and management of undocumented wells.

Energy - Petroleum

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics

Insights into coordination and ligand trends of lanthanide complexes from the Cambridge Structural Database

Abstract Understanding lanthanide coordination chemistry can help develop new ligands for more efficient separation of lanthanides for critical materials needs. The Cambridge Structural Database (CSD) contains tens of thousands of single crystal structures of lanthanide complexes that can serve as a training ground for both fundamental chemical insights and future machine learning and generative artificial intelligence models. This work aims to understand the currently available structures of lanthanide complexes in CSD by analyzing the coordination shell, donor types, and ligand types, from the perspective of rare-earth element (REE) separations. We obtain four sets of lanthanide complexes from CSD: Subset 1, all Ln-containing complexes (49472 structures); Subset 2, mononuclear Ln complexes (27858 structures); Subset 3, mononuclear Ln complexes without cyclopentadienyl ligands (Cp) (26156 structures); Subset 4, Ln complexes with at least one 1,10-phenanthroline (phen) or its derivative as a coordinating ligand (2226 structures). The subsequent analysis of lanthanide complexes in these subsets examines the trends in coordination numbers and first shell distances as well as identifies and characterizes the ligands and donor groups. In addition, examples of Ln-complexes with commercially available complexants and phen-based ligands are interrogated in detail. This systematic investigation lays the groundwork for future data-driven ligand designs for REE separations based on the structural insights into the lanthanide coordination chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

INSPIRED: Inelastic neutron scattering prediction for instantaneous results and experimental design

Inelastic neutron scattering (INS) has unique advantages in probing how atoms vibrate and how the vibrations propagate and interact. Such dynamic information is crucial in understanding various material properties, from heat capacity, thermal conductivity, phase transitions, and chemical reactions to more exotic quantum behavior. The analysis and interpretation of the INS spectra often start from a model structure of the sample, followed by a series of calculations to obtain the simulated spectra to compare with experiments. The conventional way to perform such calculations usually requires significant time, computing resources, and specialized expertise. Here, we present a new program named INSPIRED (Inelastic Neutron Scattering Prediction for Instantaneous Results and Experimental Design), which enables users to perform rapid INS simulations in several different ways on their personal computers in just a few clicks, with the crystal structure as the only input file. Specifically, the users can choose a pre-trained symmetry-aware neural network (coupled with an autoencoder) to predict the phonon density of states (DOS), 1D S(E) and 2D S(|Q|,E) spectra for any given structure. One can also choose an existing density functional theory (DFT) calculation from a database (containing over 12,000 crystals), and quickly obtain the simulated INS spectra for single crystals and powders. It is also possible to use pre-trained universal machine learning force fields to relax a given crystal structure, calculate the phonon dispersion and DOS, and, subsequently, the INS spectra. All these functions are implemented with a PyQt graphic user interface. Finally, we expect these new tools will benefit broad user communities and significantly improve the efficiency of experiment design, execution, and data analysis for INS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Augmenting machine learning of Grad–Shafranov equilibrium reconstruction with Green's functions

This work presents a method for predicting plasma equilibria in tokamak fusion experiments and reactors. The approach involves representing the plasma current as a linear combination of basis functions using principal component analysis of plasma toroidal current densities (J t ) from the EFIT-AI equilibrium database. Then utilizing EFIT's Green's function tables, basis functions are created for the poloidal flux (ψ) and diagnostics generated from the toroidal current (J t ). Similar to the idea of a physics-informed neural network (NN), this physically enforces consistency between ψ, J t , and the synthetic diagnostics. First, the predictive capability of a least squares technique to minimize the error on the synthetic diagnostics is employed. The results show that the method achieves high accuracy in predicting ψ and moderate accuracy in predicting J t with median R 2 = 0.9993 and R 2 = 0.978, respectively. A comprehensive NN using a network architecture search is also employed to predict the coefficients of the basis functions. The NN demonstrates significantly better performance compared to the least squares method with median R 2 = 0.9997 and 0.9916 for J t and ψ, respectively. The robustness of the method is evaluated by handling missing or incorrect data through the least squares filling of missing data, which shows that the NN prediction remains strong even with a reduced number of diagnostics. Additionally, the method is tested on plasmas outside of the training range showing reasonable results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Neural Network Analysis of Nuclear Magnetic Resonance and Infrared Spectra

Nuclear magnetic resonance (NMR) spectroscopy and infrared (IR) spectroscopy are powerful chemical characterization techniques with broad general usage. However, the manual evaluation of the resulting spectra is time-consuming and requires significant expertise, preventing insights from being used in real-time applications. With recent advances in computation and artificial intelligence (AI), new tools are available for automating spectral interpretation. In this work, machine learning (ML) algorithms using 1-dimensional convolutional neural networks (CNNs) were applied to identify common functional groups from spectral information. Raw spectra were collected virtually from the Human Metabolome Database (HMDB) and National Institute of Standards and Technology (NIST) Chemistry WebBook and processed into a suitable standard. Algorithm design was tailored to best fit the nature of the problem, with built-in flexibility to accommodate relevant parameters beyond the raw spectral input, specifically solvent identity and magnetic frequency for NMR. The predictive capability of the algorithm in identifying functional groups is displayed in several examples. This methodology has been compiled into a code repository and could easily be modified to adapt alternative data sources, including other spectrum types. To mitigate overfitting, a common problem in mathematical modeling where overfamiliarity with training data produces trends that are not representative of the general data, a novel metric was developed, referred to as Accufit. Accufit includes a parameter that penalizes substantial differences in the training accuracy and the accuracy of an independent validation set. Examples are presented showing the effectiveness of Accufit in maintaining the model’s predictive capability while controlling the overfitting when used as a custom metric for hyperparameter tuning.

Sturgill, James

Constituent Data Replacement Tool

The purpose of this tool is to estimate key parameters that may be missing in public wastewater composition datasets. The tool can be applied to develop complete treatment and critical mineral extraction profiles for leachate, produced water and other aqueous waste streams. The tool applies machine learning algorithms to replace missing data in a user’s water data set that are adjusted based on user preferences for options including algorithm type, number of features, and classification variables. The tool can use the user’s data alone or combine user data with the NEWTS USGS Produced Water Database for more robust training. This research was funded by the U.S. Department of Energy’s Office Fossil Energy and Carbon Management (FECM) through National Energy Technology Laboratory’s ongoing research under the Water Management for Power System Field Work Proposal, DE-FECM 1022428 and Critical Minerals Field Work Proposal, DE-FECM 1022420.

Aqueous Chemistry

Machine-Learning-Driven Discovery of Water Splitting BaFe 2 O 4 and Human-in-the-Loop Improvement via Al-Substitution for Increased Thermal Stability

Thermochemical hydrogen (TCH) production offers a promising method for converting thermal energy into hydrogen fuel through heat-driven redox cycles of metal oxides. Here, in this work a defect graph neural network (dGNN) was used to predict oxygen vacancy formation energies ΔH V O combined with Materials Project predictions of oxygen chemical potential stability to screen candidate oxides via high-throughput database analysis. BaFe 2 O 4 was identified as a promising material for experimental validation based on its predicted ΔH V O , oxygen chemical potential stability range, and potential for tunable substitutions to improve thermal properties. Experimental validation using thermogravimetric analysis (TGA), stagnation flow reactor (SFR), X-ray diffraction (XRD), and electron microscopy confirmed positive water-splitting behavior but also revealed limitations in thermal stability under aggressive reduction conditions. To address this, a human-in-the-loop modification strategy was employed introducing Al substitution in BaFe 2–x Al x O 4 ; this modification improves thermal stability, alters the crystal structure and enhances overall performance. These results demonstrate a combined computational and experimental workflow in which machine learning accelerates identification of promising candidates, while targeted experimental design enables optimization of functional performance. This approach advances the development of robust, cost-effective TCH materials and highlights the importance of integrating data-driven discovery with human-guided materials design in paving the way for scalable hydrogen production technologies.

organic

Machine learning-assisted design of metal–organic frameworks for hydrogen storage: A high-throughput screening and experimental approach

Various theoretical approaches, including big data and high-throughput screening techniques, have been explored in developing new materials due to their significant potential time-saving advantages. However, it remains a significant challenge to experimentally realize new materials that are predicted. In this study, we propose a novel materials design strategy that utilizes machine-learning (ML) techniques to predict new porous materials that show promise for hydrogen storage and are likely to be feasible to synthesize. By leveraging ML techniques and metal–organic framework (MOF) databases, we are able to predict the synthesizability of MOF structures. This is evidenced by the successful synthesis of a new vanadium-based MOF that exhibits excellent performance for cryogenic H 2 storage. Notably, the total gravimetric and volumetric H 2 uptakes are as high as 9.0 wt% and 50.0 g/L at 77 K and 150 bar. This ML-assisted materials design offers an efficient and promising approach for developing hydrogen storage materials.

08 HYDROGEN

Benchmarking universal machine learning interatomic potentials for rapid analysis of inelastic neutron scattering data

The accurate calculation of phonons and vibrational spectra remains a significant challenge, requiring highly precise evaluations of interatomic forces. Traditional methods based on the quantum description of the electronic structure, while widely used, are computationally expensive and demand substantial expertise. Emerging universal machine learning interatomic potentials (uMLIPs) offer a transformative alternative by employing pre-trained neural network surrogates to predict interatomic forces directly from atomic coordinates. This approach dramatically reduces computation time and minimizes the need for technical knowledge. In this paper, we produce a phonon database comprising nearly 5000 inorganic crystals to benchmark the performance of several leading uMLIPs. We further assess these models in real-world applications by using them to analyze experimental inelastic neutron scattering data collected on a variety of materials. Through detailed comparisons, we identify the strengths and limitations of these uMLIPs, providing insights into their accuracy and suitability for fast calculations of phonons and related properties, as well as the potential for real-time interpretation of neutron scattering spectra. Our findings highlight how the rapid advancement of AI in science is revolutionizing experimental research and data analysis.

inelastic neutron scattering

Non-Electricity Based Renewable Fuels: Theory and Computation for Solar Thermochemical Hydrogen

Dominated by photovoltaics and wind, current renewable energy sources generate mostly electricity, but 80% of the global final energy consumption occurs in form of fuels. Therefore, direct solar fuel generation would be a major breakthrough for the energy transition. Solar thermochemical hydrogen (STCH) is one of the very few potential routes towards scalable renewable fuels, but currently suffers from lack of an oxide working material that could optimally perform energy conversion within the thermodynamic boundary conditions. Theory and computation can contribute in two distinct ways, through materials search and discovery, but also by providing detailed mechanistic models for specific systems so to advance our understanding of possible design strategies. To enable high-throughput materials screening, we developed a defect graph neural network (dGNN) machine learning approach,[1] which accelerates the prediction of defect formation energies by replacing the tedious density functional theory (DFT) supercell calculations for all possible defect sites. This approach enables high-throughput database screening of oxides, which was integrated with thermodynamic modeling to extract the reduction entropies as additional selection criterion for STCH. Once potential candidate materials are identified, detailed models can guide materials design by predicting performance characteristics. One challenge is to quantitatively predict thermochemical equilibria at high concentrations when the redox active defects start to interact with each other, thereby impeding the formation of additional defects. Introducing a model for the free energy of defect interaction, parametrized on the basis of DFT data, we simulated the complete STCH redox cycle for (Sr,Ce)MnO3 alloys, achieving near-quantitative agreement with experimental data.[2] The analysis of these simulations reveals how defect interactions diminish the reduction entropy and H2 yield, suggesting to include these interactions in design considerations. Finally, we revisit the popular van't Hoff method for analyzing reduction enthalpies and entropies. This method is not ideal, as it involves a temperature-dependent convolution of gas-phase and solid-state entropies, causing uncertainties in the same order of magnitude as the physical quantities of interest. To avoid this problem, we suggest a simple alternative approach which can be applied to experimental and simulated data alike.

first-principles calculations

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES

Artificial Intelligence and Machine Learning Applications in Modern Power Systems

Machine learning (ML) and artificial intelligence (AI) algorithms offer valuable tools for the analysis and interpretation of large datasets. These tools have the capability to uncover insights that may not be readily apparent within these datasets. In recent years, the integration of ML and AI has become increasingly prevalent in various applications within the power system domain. One of the earliest instances of machine learning in power systems can be traced back to demand forecasting, where artificial neural networks were employed for short-term load forecasting. In contemporary power systems, an abundance of high-resolution geospatial and temporal data is generated at various time intervals, ranging from sub-seconds (Phasor Measurement Units or PMUs) to seconds (Supervisory Control and Data Acquisition or SCADA), minutes (Process Information or PI), and extending to days, months, and years. These datasets contain valuable information concerning system reliability and performance. This information holds the potential to offer critical insights into system operations, as well as solutions for predicting and mitigating contingencies to prevent cascading outages. Despite the immense power of machine learning tools, system operators, planners, and utilities often exhibit hesitancy in fully embracing AI-enabled system operations and planning. This cautious approach persists, even as numerous diverse applications of machine learning continue to emerge in the realm of power systems. In this chapter, our focus will delve deep into ML and AI applications tailored for power systems. These applications aim to furnish system operators with enhanced situational awareness and augment their decision-making capabilities, especially during challenging operating conditions. Specific areas of interest encompass root cause analyses of electricity market datasets and the strategic selection of representative samples from vast power system databases for training ML/AI models. Finally, the chapter will conclude with a short discussion on the future of ML/AI in power systems and possible directions that the industry is moving towards.

power system applications, machine learning (ML),

Temperature‐Dependent Crystallization in Two‐Step Perovskite Deposition Revealed by In Situ GIWAXS and Machine Learning‐Guided Analysis

The performance and stability of perovskite solar cells are strongly governed by the crystallization behavior of their active layer. In two-step sequential deposition, early-stage film formation plays a decisive role in determining final phase purity and device quality. Guided by a data-driven analysis of nearly 39 000 devices in the FAIR perovskite database, we identified solvent-mediated quenching and thermal processing as key variables affecting power conversion efficiency (PCE), particularly in two-step fabrication. Here, to investigate these effects in real time, we designed and implemented a custom-built, temperature-controlled spin-coating system, enabling precise thermal modulation during precursor deposition. Using this platform, we performed in situ GIWAXS measurements to study the crystallization dynamics of FA 0.5 MA 0.5 PbI 3 films over a temperature range of 30°C–90°C. Our results reveal a non-monotonic relationship between spin-coating temperature and α-phase formation, governed by the interplay between precursor interdiffusion, PbI 2 crystallinity, and δ-phase suppression. The custom thermal control enabled us to isolate and quantify these competing effects during the earliest stages of film formation, providing mechanistic insight into how spin-coating temperature governs both phase purity and kinetic pathways in two-step perovskite systems. Temperature-dependent SEM and photovoltaic device measurements further demonstrate that early-stage crystallization pathways directly translate into differences in morphology, charge-transport continuity, and device performance. These findings inform targeted strategies for optimizing deposition protocols to balance rapid nucleation, phase stability, and device performance.

Saadawy, Ahmed [King Fahd University of Petroleum

A machine learning estimator trained on synthetic data for real-time earthquake ground-shaking predictions in Southern California

Abstract After large-magnitude earthquakes, a crucial task for impact assessment is to rapidly and accurately estimate the ground shaking in the affected region. To satisfy real-time constraints, intensity measures are traditionally evaluated with empirical Ground Motion Models that can drastically limit the accuracy of the estimated values. As an alternative, here we present Machine Learning strategies trained on physics-based simulations that require similar evaluation times. We trained and validated the proposed Machine Learning-based Estimator for ground shaking maps with one of the largest existing datasets (<100M simulated seismograms) from CyberShake developed by the Southern California Earthquake Center covering the Los Angeles basin. For a well-tailored synthetic database, our predictions outperform empirical Ground Motion Models provided that the events considered are compatible with the training data. Using the proposed strategy we show significant error reductions not only for synthetic, but also for five real historical earthquakes, relative to empirical Ground Motion Models.

Environmental Sciences & Ecology