Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Machine Learning Classification Strategy to Improve Streamflow Estimates in Diverse River Basins in the Colorado River Basin

Streamflow in the Colorado River Basin (CRB) is significantly altered by human activities including land use/cover alterations, reservoir operation, irrigation, and water exports. Climate is also highly varied across the CRB which contains snowpack-dominated watersheds and arid, precipitation-dominated basins. Recently, machine learning methods have improved the generalizability and accuracy of streamflow models. Previous successes with LSTM modeling have primarily focused on unimpacted basins, and few studies have included human impacted systems in either regional or single-basin modeling. We demonstrate that the diverse hydrological behavior of river basins in the CRB are too difficult to model with a single, regional model. We propose a method to delineate catchments into categories based on the level of predictability, hydrological characteristics, and the level of human influence. Lastly, we model streamflow in each category with climate and anthropogenic proxy data sets and use feature importance methods to assess whether model performance improves with additional relevant data. Overall, land use cover data at a low temporal resolution was not sufficient to capture the irregular patterns of reservoir releases, demonstrating the importance of having high-resolution reservoir release data sets at a global scale. On the other hand, the classification approach reduced the complexity of the data and has the potential to improve streamflow forecasts in human-altered regions.

54 ENVIRONMENTAL SCIENCES↗

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗

Machine learning uncovers independently regulated modules in the Bacillus subtilis transcriptome

The transcriptional regulatory network (TRN) of Bacillus subtilis coordinates cellular functions of fundamental interest, including metabolism, biofilm formation, and sporulation. Here, we use unsupervised machine learning to modularize the transcriptome and quantitatively describe regulatory activity under diverse conditions, creating an unbiased summary of gene expression. We obtain 83 independently modulated gene sets that explain most of the variance in expression and demonstrate that 76% of them represent the effects of known regulators. The TRN structure and its condition-dependent activity uncover putative or recently discovered roles for at least five regulons, such as a relationship between histidine utilization and quorum sensing. The TRN also facilitates quantification of population-level sporulation states. As this TRN covers the majority of the transcriptome and concisely characterizes the global expression state, it could inform research on nearly every aspect of transcriptional regulation in B. subtilis.

59 BASIC BIOLOGICAL SCIENCES↗

Quantum annealing-assisted lattice optimization

High Entropy Alloys (HEAs) have drawn great interest due to their exceptional properties compared to conventional materials. The configuration of HEA system is considered a key to their superior properties, but exhausting all possible configurations of atom coordinates and species to find the ground energy state is extremely challenging. In this work, we proposed a quantum annealing-assisted lattice optimization (QALO) algorithm, which is an active learning framework that integrates the Field-aware Factorization Machine (FFM) as the surrogate model for lattice energy prediction, Quantum Annealing (QA) as an optimizer and Machine Learning Potential (MLP) for ground truth energy calculation. By applying our algorithm to the NbMoTaW alloy, we reproduced the Nb depletion and W enrichment observed in bulk HEA. We found our optimized HEAs to have superior mechanical properties compared to the randomly generated alloy configurations. Our algorithm highlights the potential of quantum computing in materials design and discovery, laying a foundation for further exploring and optimizing structure-property relationships.

36 MATERIALS SCIENCE↗

AI for nuclear physics: the EXCLAIM project

An overview of the recent activity of the newly funded EXCLusives with AI and Machine learning (EXCLAIM) collaboration is presented. The main goal of the collaboration is to develop a framework to implement AI and machine learning techniques in problems emerging from the phenomenology of high energy exclusive scattering processes from nucleons and nuclei, maximizing the information that can be extracted from various sets of experimental data, while implementing theoretical constraints from lattice QCD. A specific perspective embraced by EXCLAIM is to use the methods of theoretical physics to understand the working of ML, beyond its standardized applications to physics analyses which most often rely on industrially provided tools, in an automated way.

Analysis and statistical methods↗

DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic Hardware

Spiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches.

Computer Science↗

Hpc Natural Language Understanding (nlu) Dataset

This code provides natural language annotation and labels for multiple activities in high performance. The purpose is to enable machine learning for natural language processing algorithms with high performance computing systems.

Biggs, BrandonS↗

Scale Bridging through Active Learning [PowerPoint]

The slides discuss Active Learning and building dynamic surrogate models with confidence with Machine Learning. A prototype example involving mixing and transport in an inertial confinement and fusion experiment illustrates the concept.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Systems and methods for active learning from sparse training data

A method for active learning using sparse training data can include training a machine learning model using less than ten first training data points to generate a candidate machine learning model. The method can include performing a Monte Carlo process to sample one or more first outputs of the candidate machine learning model. The method can include testing the one or more first outputs to determine if each of the one or more first outputs satisfy a respective convergence condition. The method can include, responsive to at least one first output not satisfying the respective convergence condition, training the candidate machine learning model using at least one second training data point corresponding to the at least one first output. The method can include, responsive to the one or more first outputs each satisfying the respective convergence condition, outputting the candidate machine learning model.

Sankaranarayanan, Subramanian↗

Unraveling Hydrogen Induced Geochemical Reaction Mechanisms through Coupled Geochemical Modeling and Machine Learning

Underground hydrogen storage (UHS) provides a promising large-scale, long-term energy storage solution. A reasonable recovery of stored hydrogen is critical for a successful storage scheme. However, in subsurface reservoirs hydrogen is subject to active geochemical reactions that might result in hydrogen loss. In this study, we implemented a geochemical modeling approach coupled with an unsupervised machine learning technique called non-negative matrix factorization (NMF) to unravel the complex brine-rock-H 2 geochemical processes responsible for hydrogen losses, with particular focus on sulfate reduction reactions. NMF is applied to modeled mineral evolution and fluid component profiles to retrieve profiles that can be interpreted to more easily assess competing processes. NMF decouples simulated competing equilibrium reactions. This facilitates separation of overlapping reaction profiles from redox processes, dissolution fronts, and secondary precipitation while considering the effects of simulation parameters such as salinity, temperature, and total H 2 pressure. NMF successfully discriminates these competing effects in nonlinear ways, allowing robust interpretation. In addition, NMF reveals subtle coupled mineral associations and reaction fronts that are invisible to conventional model analysis. This integrated approach strengthens the conceptual understanding of complex nonlinear hydrogen-brine-rock interactions and advances geochemical research on UHS systems to resolve complexities in modeled geochemical systems without the need for direct experiments or prior knowledge. Furthermore, this study highlights the efficacy of combining geochemical modeling with machine learning techniques to enhance the interpretability of the intricate geochemical simulation output through deciphering the overlapping reaction path that cannot be achieved only using conventional analysis of geochemical models alone.

08 HYDROGEN↗

A Wrapper to Use a Machine-Learning-Based Algorithm for Earthquake Monitoring

Seismology is one of the main sciences used to monitor volcanic activity worldwide. Fast, efficient, and accurate seismicity detectors are crucial to assess the activity level of a volcano in near–real time and to issue timely warnings. Traditional real–time seismic processing software uses phase onset pickers followed by a phase association algorithm to declare an event and estimate its location. The pickers typically do not identify whether the detected phase is a P or S arrival, which can have a negative impact on hypocentral location quality and complicates phase association. We implemented the deep–neural–network–based method PhaseNet to identify in real time P and S seismic waves on data from one– and three–component seismometers. We tuned the Earthworm binder_ew associator module to use the phase identification from PhaseNet to detect and locate the events, which we archive in a SeisComP3 database. We assessed the performance of the algorithm by comparing the results with existing catalogs built to monitor seismic and volcanic activity in Mayotte and the Lesser Antilles region. Our algorithm, which we refer to as PhaseWorm, showed promising results in both contexts and clearly outperformed the previous automatic method implemented in Mayotte. As a result, this innovative real–time processing system is now operational for seismicity monitoring in Mayotte and Martinique.

58 GEOSCIENCES↗

Unveiling the effect of composition on nuclear waste immobilization glasses’ durability by nonparametric machine learning

Abstract Ensuring the long-term chemical durability of glasses is critical for nuclear waste immobilization operations. Durable glasses usually undergo qualification for disposal based on their response to standardized tests such as the product consistency test or the vapor hydration test (VHT). The VHT uses elevated temperature and water vapor to accelerate glass alteration and the formation of secondary phases. Understanding the relationship between glass composition and VHT response is of fundamental and practical interest. However, this relationship is complex, non-linear, and sometimes fairly variable, posing challenges in identifying the distinct effect of individual oxides on VHT response. Here, we leverage a dataset comprising 654 Hanford low-activity waste (LAW) glasses across a wide compositional envelope and employ various machine learning techniques to explore this relationship. We find that Gaussian process regression (GPR), a nonparametric regression method, yields the highest predictive accuracy. By utilizing the trained model, we discern the influence of each oxide on the glasses’ VHT response. Moreover, we discuss the trade-off between underfitting and overfitting for extrapolating the material performance in the context of sparse and heterogeneous datasets.

36 MATERIALS SCIENCE↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

Accelerated screening of functional atomic impurities in halide perovskites using high-throughput computations and machine learning

The pressing need for novel materials that can serve rising demands in solar cell and optoelectronic technologies makes the nexus of halide perovskites, high-throughput computations, and machine learning, very promising. Ever increasing amounts of data on the structure, fundamental properties, and device performance of halide perovskites provide opportunities for learning chemical rules and design principles that make these materials attractive, and applying them across wide chemical spaces. In this work, we show that impurity properties of halide perovskites computed using density functional theory (DFT) can be combined with machine learning (ML) to deliver predictive models and quick identification of optoelectronically active impurity atoms. Our computation lead to the largest reported dataset of the formation energies and charge transition levels of Pb-site impurities in methylammonium lead halide (MAPbX 3 ) perovskites. Descriptors are defined to uniquely represent any impurity atom in any MAPbX 3 compound and mapped to the computed impurity properties using regression techniques such as Gaussian process regression, neural networks, and random forests. We use the best optimized predictive models to make predictions for hundreds of impurities across 9 MAPbX 3 compounds and create lists of dominating impurities, that is, impurities that can shift the equilibrium Fermi level in the perovskite as determined by native point defects. Finally, this accelerated screening powered by computations and machine learning can guide the identification of problematic impurities that may cause undesired recombination of charge carriers, as well as impurities that can be deliberately introduced to tune the perovskite conductivity and resulting photovoltaic absorption.

36 MATERIALS SCIENCE↗

Detecting Satellite Laser Ranging Station Data and Operational Anomalies with Machine Learning Isolation Forests at NASA's CDDIS

The International Laser Ranging Service (ILRS) is currently composed of 45 active satellite laser ranging (SLR) stations with several more set to join the network over the next several years. Station changes and histories are logged to files, but not always in real time. Sometimes these details are not added until long after changes have been made to the station –on occasion, years later. This in addition to unexpected hardware errors and other system issues that are not immediately detected impact the products generated by analysts. The ILRS Central Bureau (CB) and NASA’s Crustal Dynamics Data Information System (CDDIS) have worked to provide tools for station engineers to use. This includes the creation of station plots which contain temperature and pressure information along with LAser GEOdynamic Satellite (LAGEOS) and LAser RElativity Satellite (LARES) tracking information that enable the monitoring of station performance and todetermine whether the station has undergone any changes. As next steps, the CDDIS is working to enhance these station performance monitoring tools through machine learning. Isolation forest is an unsupervised machine learning algorithm commonly applied to anomaly detection. In this poster, the CDDIS details the steps taken to track anomalies within SLR station performance using isolation forest with LAGEOS and LARES satellite data.

Benjamin P Michael↗

Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems

Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.We address such issues by building an end-to-end machine learn- ing (ML) workflow that can generate I/O traces for large-scale HPC applications. We leverage ML based feature selection and gener- ative models for I/O trace generation. The generative models are trained on I/O traces collected by the darshan I/O characterization tool over a period of one year. We present a two-step generation process consisting of two deep-learning models, called the feature generator and the trace generator. The combination of two-step generative models provides robustness by reducing the bias of the model and accounting for the stochastic nature of the I/O traces across different runs of an application. We evaluate the performance of the generative models and show that the two-step model can generate time-series I/O traces with less than 20% root mean square error.

Paul, Arnab↗

Chemically Enabled CO 2 -Enhanced Oil Recovery in Multi-Porosity, Hydrothermally Altered Carbonates in the Southern Michigan Basin - Task 2 Topical Report

This attachment A is a detailed Topical Report for Task 2 (Advanced Field Characterization and Machine Learning Based Data Integration) under the project "Chemically Enabled CO 2 -Enhanced Oil Recovery in Multi-Porosity, Hydrothermally Altered Carbonates in the Southern Michigan Basin." The overall project activity and finding, including the field pilot testing of CO 2 injection are summarized in the companion Final Technical Report The Advanced Field Characterization and Machine Learning Based Data Integration task (Task 2) involved a systematic geologic characterization of the Trenton-Black River (TBR) play in the SMB which included the development of comprehensive datasets, advanced data analytics, risk assessment, and piggyback field characterization. These activities aimed to address the complex carbonate systems by evaluating the extent of fractures, dolomitization, facies, and reservoir properties with the primary objective of informing the static and dynamic modeling, field injection test, and providing input into the development strategy plan. The task was divided into four subtasks: • Subtask 2.1 – Data compilation, review, and analysis • Subtask 2.2 – Risk Assessment • Subtask 2.3 – Advanced Field Characterization • Subtask 2.4 – Integrated Physics-Based Machine Learning and Advanced Data Analytics

02 PETROLEUM↗

CoRE MOF DB: A curated experimental metal-organic framework database with machine-learned properties for integrated material-process screening

Here, we present an updated version of the Computation-Ready, Experimental (CoRE) Metal-Organic Framework (MOF) database, which includes a curated set of computation-ready MOF crystal structures designed for high-throughput computational materials discovery. Data collection and curation procedures were improved from the previous version to enable more frequent updates in the future. Machine-learning-predicted properties, such as stability metrics and heat capacities, are included in the dataset to streamline screening activities. An updated version of MOFid was developed to provide detailed information on metal nodes, organic linkers, and topologies of an MOF structure. DDEC6 partial atomic charges of MOFs were assigned based on a machine-learning model. Gibbs ensemble Monte Carlo simulations were used to classify the hydrophobicity of MOFs. The finalized dataset was subsequently used to perform integrated material-process screening for various carbon-capture conditions using high-fidelity temperature-swing adsorption (TSA) simulations. Our workflow identified multiple MOF candidates that are predicted to outperform CALF-20 for these applications.

CoRE MOF database↗