Engineering PapersSearch

SEARCH · Engineering Papers

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Use of Frit‐Disc Crucible Sets to Make Solution Growth More Quantitative and Versatile

The recent availability of step‐edge, frit‐disc crucible sets (generally sold as Canfield Crucible Sets or CCS) has led to multiple innovations associated with the group's use of solution growth. The use of CCS allows for the clean separation of liquid from solid phases during the growth process. This clean separation enables the reuse of the decanted liquid, either allowing for simple, economic, savings associated with recycling expensive precursor elements or allowing for the fractionation of a growth into multiple, small steps, revealing the progression of multiple solidifications. Clean separation of liquid from solid phases also allows for the determination of the liquidus line (or surface) and the creation, or correction, of composition–temperature phase diagrams. The reuse of clean decanted liquid has also allowed to prepare liquids ideally suited for the growth of large single crystals of specific phases by tuning the composition of the melt to the optimal composition for growth of the desired phase, often with reduced nucleation sites. Finally, it is discussed how solution growth and CCS use can be harnessed to provide a plethora of composition–temperature data points defining liquidus lines or surfaces with differing degrees of precision to either test or anchor artificial intelligence and/or machine‐learning‐based attempts to augment and extend the limited experimentally determined database.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES

Virtual node graph neural network for full phonon prediction

Understanding the structure-property relationship is crucial for designing materials with desired properties. The past few years have witnessed remarkable progress in machine-learning methods for this connection. However, substantial challenges remain, including the generalizability of models and prediction of properties with materials-dependent output dimensions. Here we present the virtual node graph neural network to address the challenges. By developing three virtual node approaches, we achieve Γ-phonon spectra and full phonon dispersion prediction from atomic coordinates. We show that, compared with the machine-learning interatomic potentials, our approach achieves orders-of-magnitude-higher efficiency with comparable to better accuracy. This allows us to generate databases for Γ-phonon containing over 146,000 materials and phonon band structures of zeolites. Additionally, our work provides an avenue for rapid and high-quality prediction of phonon band structures enabling materials design with desired phonon properties. The virtual node method also provides a generic method for machine-learning design with a high level of flexibility. In this study, the authors present a virtual node graph neural network to enable the prediction of material properties with variable output dimensions. This method offers fast and accurate predictions of phonon band structures in complex solids.

36 MATERIALS SCIENCE

Machine learning surrogates for ion energy–angle distributions in thermal and RF plasma sheaths

Ion energy–angle distributions (IEADs) at material surfaces are a critical input for plasma–material interaction (PMI) studies in fusion devices, yet they are computationally expensive to obtain using particle-in-cell (PIC) simulations. In this work, we develop a machine learning surrogate based on a deep deconvolutional neural network (DDeCNN) trained on large databases generated with the hPIC2 code. The surrogate is capable of reconstructing IEADs from sheath parameters for both thermal and radio-frequency (RF) plasmas, including cases with multiple ion species. Across thousands of test cases, the model achieves high accuracy, with over 97 % of predictions classified as good or average based on standard error metrics (MAE, MSE, L2). Even in the more challenging RF and multi-species regimes, the surrogate reliably captures the multi-peak structure of PIC results. Once trained, the surrogate produces IEADs in milliseconds on a common workstation, yielding speedups of six to seven orders of magnitude compared with running a full PIC simulation. This computational gain enables dense parameter scans and direct coupling of IEAD predictions with PMI and erosion models on whole-device scales in fusion-relevant conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

hashin_shtrikman_mp: a package for the optimal design and discovery of multi-phase composite materials

hashin_shtrikman_mp is a tool for composites designers who have desired composite properties in mind, but who do not yet have an underlying formulation. The library utilizes the tightest theoretical bounds on the effective properties of composite materials with unspecified microstructure – the Hashin-Shtrikman bounds – to identify candidate theoretical materials, find real materials that are close to the candidates, and determine the optimal volume fractions for each of the constituents in the resulting composite. Its features include (i) leveraging of materials in the Materials Project database, (ii) integration with the Materials Project API, (iii) use of genetic machine-learning, (iv) agnosticism to underlying microstructure, and (v) ultimate engineering application, make it a tool with much broader applications than its predecessors.

97 MATHEMATICS AND COMPUTING

G2Aero Database of Airfoils - Curated Airfoils

This dataset contains a curated set of 19,164 airfoil shapes from various applications and the data-driven design space of separable shape tensors (PGA space), which can be used as a parameter space for machine-learning applications focused on airfoil shapes. We constructed the airfoil dataset in two main stages. First, we identified 13 baseline airfoils from the NREL 5MW and IEA 15MW reference wind turbines. We reparameterized these shapes using least-squares fits of 8-order CST parametrizations, which involve 18 coefficients. By uniformly perturbing all 18 CST coefficients by +/-20% around each baseline airfoil, we generated 1,000 unique airfoils. Each airfoil was sampled with 1,001 shape landmarks whose x-coordinates followed a cosine distribution along the chord. This process resulted in a total of 13,000 airfoil shapes, each with 1,001 landmarks. In the second phase, we gathered additional airfoils from the extensive BigFoil database, which consolidates data from sources such as the University of Illinois Urbana-Champaign (UIUC) airfoil database, the JavaFoil database, the NACA-TR-824 database, and others. We undertook a thorough pre-processing step to filter out shapes with sparse, noisy, or incomplete data. We also removed airfoils with sharp leading edge and those exceeding our threshold for trailing edge thickness. Additionally, we thinned out the collection of NACA airfoils-- parametric sweeps of NACA airfoils with increasing thickness and camber present in BigFoil database-- by selecting every fourth step in the parameter sweeps. Finally, we regularized the airfoils by reparametrizing them with an 8-order CST parametrization (with 1,001 shape landmarks with x coordinated following cosine distribution along the chord) and removing airfoils with high reconstruction errors. This data pre-processing resulted in a set of 6,164 airfoils. In total, our curated airfoil dataset comprises 19,164 airfoils, each with 1,001 landmarks, and is stored in the curated_airfoils.npz file. Using this curated airfoil dataset, we utilized the separable shape tensors framework to develop a data-driven parameterization of airfoils based on principal geodesic analysis (PGA) of separable shape tensors. This PGA space is provided in PGAspace.npz file.

airfoils

Quantitative Analysis and Prediction of Thermal Runaway Metrics of High-Nickel Oxide Cathodes by Machine Learning Models

The pursuit of higher energy density in lithium-ion batteries has made high-nickel (Ni) layered oxides leading cathode candidates for next-generation electric vehicles. However, their poor thermal stability, particularly at Ni contents ≥ 90%, increases the risk of cathode-initiated thermal runaway. Furthermore, we present a data-driven framework combining linear and nonlinear machine learning models to predict key thermal runaway descriptors from a high-throughput differential scanning calorimetry database. With cathode composition and state of charge (SOC) as input features, the ensemble model accurately predicts peak temperature, heat release, and peak heat flow. SHAP analysis identifies Ni content and SOC as the dominant factors controlling thermal runaway temperature, while SOC primarily governs heat release and peak heat flow. Al, Mg, and Mn improve thermal stability by strengthening metal–oxygen bonding and delaying structural transformation, whereas B mainly reduces heat release through surface passivation. Validation with a new cathode composition confirms accurate prediction of SOC-dependent thermal runaway behavior and critical SOC.

25 ENERGY STORAGE

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science

Zentropy Theory for Transformative Functionalities of Magnetic and Superconducting Materials

The proposed research developed the zentropy theory through applications to complex magnetic materials and superconductors under the hypothesis that the emergent properties of complex magnetic materials and superconductors can be predicted by statistical mechanics of ergodic microstates with their partition functions computed from DFT-predicted free energies. The key objective is to develop approaches to systematically determine the types and number of microstates and the supercell size in DFT-based calculations through convergency of macroscopic functionalities, with the incorporation of our mixed-space approach accounting for the interactions between periodic supercells. In addition to use scientific intuitions to guide the design of important microstates, the key innovation of the proposed research is to integrate the domain knowledge and the material-property-descriptor database (MPDD) with 4 million microstates, which is supported by our deep neural network machine learning models (SIPFENN: structure-informed prediction of formation energy using neural networks) and integrated with our high throughput DFT Tool Kit (DFTTK). For complex magnetic materials, one of the objectives is to develop approaches to calculate short-range ordering from the statistical distribution of each microstate. For superconductors, the divergency of quasiparticle effective mass at a quantum critical point will be investigated, and the superconducting and non-superconducting microstates will be delineated through analysis of electronic band structure, density of states, charge density, and Fermi surface.

36 MATERIALS SCIENCE

Methods in PES-Learn: Direct-Fit Machine Learning of Born–Oppenheimer Potential Energy Surfaces

The release of PES-L EARN version 1.0 as an open-source software package for the automatic construction of machine learning models of semi-global molecular potential energy surfaces (PESs) is presented. Improvements to PES-L EARN ’s interoperability are stressed with new Python API that simplifies workflows for PES construction via interaction with QCSchema input and output infrastructure. In addition, a new machine learning method is introduced to PES-L EARN : kernel ridge regression (KRR). The capabilities of KRR are emphasized with examination of select semi-global PESs. All machine learning methods available in PES-L EARN are benchmarked with benzene and ethanol datasets from the rMD17 database to illustrate PES-L EARN ’s performance ability. Fitting performance and timings are assessed for both systems. Finally, the ability to predict gradients with neural network models is presented and benchmarked with ethanol and benzene. PES-L EARN is an active project and welcomes community suggestions and contributions.

kernel ridge regression

Accelerating catalytic advancements through the precision of high-throughput experiments & calculations

The growing demand for energy-efficient processes to support a sustainable future drives the need for research to rapidly explore chemical and material space through accelerated catalyst discovery initiatives. Recent breakthroughs in high-throughput experimental and computational methods are transforming the catalysis field, surpassing traditional approaches to manipulating variables in catalytic processes. Key advancements in innovation include the integration of machine learning for efficient catalyst screening, high-throughput experimentation, data-driven methodologies employing comprehensive databases, and in situ and in operando techniques for realistic observations. This progress has undoubtedly been intertwined with a collaborative framework across disciplines, reshaping catalyst discovery methods in both industry and academia. This Opinion article presents a multifaceted perspective from coauthors with expertise spanning various stages of the Technology Readiness Level spectrum, highlighting both opportunities and persistent challenges in integrating computational and experimental approaches in catalysis. These challenges span from obtaining high-quality experimental data, scaling simulations to industrially relevant materials and process conditions to navigating the complexity and predictive accuracy of computational models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE

Developing predictive models for µ opioid receptor binding using machine learning and deep learning techniques

Opioids exert their analgesic effect by binding to the µ opioid receptor (MOR), which initiates a downstream signaling pathway, eventually inhibiting pain transmission in the spinal cord. However, current opioids are addictive, often leading to overdose contributing to the opioid crisis in the United States. Therefore, understanding the structure-activity relationship between MOR and its ligands is essential for predicting MOR binding of chemicals, which could assist in the development of non-addictive or less-addictive opioid analgesics. This study aimed to develop machine learning and deep learning models for predicting MOR binding activity of chemicals. Chemicals with MOR binding activity data were first curated from public databases and the literature. Molecular descriptors of the curated chemicals were calculated using software Mold2. The chemicals were then split into training and external validation datasets. Random forest, k-nearest neighbors, support vector machine, multi-layer perceptron, and long short-term memory models were developed and evaluated using 5-fold cross-validations and external validations, resulting in Matthews correlation coefficients of 0.528–0.654 and 0.408, respectively. Furthermore, prediction confidence and applicability domain analyses highlighted their importance to the models’ applicability. Our results suggest that the developed models could be useful for identifying MOR binders, potentially aiding in the development of non-addictive or less-addictive drugs targeting MOR.

Research & Experimental Medicine

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE

The northeast materials database for magnetic materials

The discovery of magnetic materials with high operating temperature ranges and optimized performance is essential for advanced applications. Current data-driven approaches are limited by the lack of accurate, comprehensive, and feature-rich databases. This study aims to address this challenge by using Large Language Models (LLMs) to create a comprehensive, experiment-based, magnetic materials database named the Northeast Materials Database (NEMAD), which consists of 67,573 magnetic materials entries (www.nemad.org). The database incorporates chemical composition, magnetic phase transition temperatures, structural details, and magnetic properties. Enabled by NEMAD, we trained machine learning models to classify materials and predict transition temperatures. Our classification model achieved an accuracy of 90% in categorizing materials as ferromagnetic (FM), antiferromagnetic (AFM), and non-magnetic (NM). The regression models predict Curie (Néel) temperature with a coefficient of determination (R 2 ) of 0.87 (0.83) and a mean absolute error (MAE) of 56K (38K). These models identified 25 (13) FM (AFM) candidates with a predicted Curie (Néel) temperature above 500K (100K) from the Materials Project. This work shows the feasibility of combining LLMs for automated data extraction and machine learning models to accelerate the discovery of magnetic materials.

Ferromagnetism