Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “active machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Solving the structure of “single-atom” catalysts using machine learning – assisted XANES analysis

We show that "single-atom” catalysts (SACs) have demonstrated excellent activity and selectivity in challenging chemical transformations such as photocatalytic CO 2 reduction. For heterogeneous photocatalytic SAC systems, it is essential to obtain sufficient information of their structure at the atomic level in order to understand reaction mechanisms. In this work, a SAC was prepared by grafting a molecular cobalt catalyst on a light-absorbing carbon nitride surface. Due to the sensitivity of the X-ray absorption near edge structure (XANES) spectra to subtle variances in the Co SAC structure in reaction conditions, different machine learning (ML) methods, including principal component analysis, K-means clustering, and neural network (NN), were utilized for in situ Co XANES data analysis. As a result, we obtained quantitative structural information of the SAC nearest atomic environment thereby extending the NN-XANES approach previously demonstrated for nanoparticles and size-selective clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An Intelligent Adaptable Monitoring Package. Final Report

The “Intelligent Adaptable Monitoring Package” project was a four-year effort that demonstrated the feasibility of integrated sensing packages at tidal and wave energy sites. Such integration is generally required by the breadth of sensors required to understand environmental effects at marine energy sites and the operational difficulty of deploying, maintaining, and recovering such sensors. Over the course of the project, the Adaptable Monitoring Package (AMP) was deployed in multiple settings, each corresponding to a project budget period: - Budget Period 1: Demonstration of cabled deployment at Pacific Northwest National Laboratory’s Marine Science Laboratory. The deployment highlighted AMP hardware endurance over a 4-month deployment in a tidally-dominated environment and laid the groundwork for machine learning algorithms to detect and classify targets present in active sonar data. - Budget Period 2: Demonstration of an autonomous deployment at PacWave South off the coast of Newport, Oregon. The deployment highlighted the stability of AMP hardware and software, with the autonomous package collecting data on a duty cycle over a 1.5-month deployment. - Budget Period 3: Demonstration of an autonomous deployment powered by a wave energy converter at the U.S. Navy’s Wave Energy Test Site. The deployment highlighted the potential of wave energy to power ocean observatories and led to the development of machine learning algorithms to detect and classify targets in optical camera data. In aggregate, this project’s greatest success was demonstrating the AMP’s flexibility in a range of deployment scenarios. Each budget period represented a “first of a kind” demonstration of integrated instrumentation – cabled AMP, autonomous AMP, wave-energy powered AMP – and each deployment helped to identify and set goals for the next. Further, despite the exploratory nature of these deployments, each one achieved high system up-time and proved that flexible integration of multiple sensors in a single package represents a viable strategy for marine energy environmental monitoring. The key lessons learned from the project are: - Without continuous power, either from a shore cable or in situ source, many of the benefits of integration are lost (If continuous power is not available, the ability to detect rare events is lost, as is the ability to minimize the risk of behavioral changes through adaptive sensing. However, even on a duty cycle, there is still value in being able to acquire synchronous data from multiple sensors.); and - Observations from a moving platform present substantially greater data processing challenges than those from stationary platforms. Finally, these deployments also demonstrate an important truth: successful integration alone does not guarantee that relevant data are collected. To grow the knowledge base about environmental interactions with marine energy converters, integrated systems, like the AMP, need to include the right sensor mix and connect the data pipelines to effective processing algorithms. These deployments establish a strong foundation for future collaborations with the environmental research community: not only to understand the environmental effects of marine energy, but also to improve our general ability to study life in the sea.

16 TIDAL AND WAVE POWER↗

Geochemistry and Multiomics Data Differentiate Streams in Pennsylvania Based on Unconventional Oil and Gas Activity

Unconventional oil and gas (UOG) extraction is increasing exponentially around the world, as new technological advances have provided cost-effective methods to extract hard-to-reach hydrocarbons. While UOG has increased the energy output of some countries, past research indicates potential impacts in nearby stream ecosystems as measured by geochemical and microbial markers. Here, we utilized a robust data set that combines 16S rRNA gene amplicon sequencing (DNA), metatranscriptomics (RNA), geochemistry, and trace element analyses to establish the impact of UOG activity in 21 sites in northern Pennsylvania. These data were also used to design predictive machine learning models to determine the UOG impact on streams. We identified multiple biomarkers of UOG activity and contributors of antimicrobial resistance within the order Burkholderiales. Furthermore, we identified expressed antimicrobial resistance genes, land coverage, geochemistry, and specific microbes as strong predictors of UOG status. Of the predictive models constructed (n = 30), 15 had accuracies higher than expected by chance and area under the curve values above 0.70. The supervised random forest models with the highest accuracy were constructed with 16S rRNA gene profiles, metatranscriptomics active microbial composition, metatranscriptomics active antimicrobial resistance genes, land coverage, and geochemistry (n = 23). The models identified the most important features within those data sets for classifying UOG status. These findings identified specific shifts in gene presence and expression, as well as geochemical measures, that can be used to build robust models to identify impacts of UOG development.

16S rRNA↗

Opportunities and Challenges for Machine Learning-Assisted Enzyme Engineering

Enzymes can be engineered at the level of their amino acid sequences to optimize key properties such as expression, stability, substrate range, and catalytic efficiency or even to unlock new catalytic activities not found in nature. Because the search space of possible proteins is vast, enzyme engineering usually involves discovering an enzyme starting point that has some level of the desired activity followed by directed evolution to improve its “fitness” for a desired application. Recently, machine learning (ML) has emerged as a powerful tool to complement this empirical process. ML models can contribute to (1) starting point discovery by functional annotation of known protein sequences or generating novel protein sequences with desired functions and (2) navigating protein fitness landscapes for fitness optimization by learning mappings between protein sequences and their associated fitness values. In this Outlook, we explain how ML complements enzyme engineering and discuss its future potential to unlock improved engineering outcomes.

60 APPLIED LIFE SCIENCES↗

MP-ALOE: an r2SCAN dataset for universal machine learning interatomic potentials

We present MP-ALOE, a dataset of nearly 1 million DFT calculations using the accurate r2SCAN meta-generalized gradient approximation. Covering 89 elements, MP-ALOE was created using active learning and primarily consists of off-equilibrium structures. We benchmark a machine learning interatomic potential trained on MP-ALOE, and evaluate its performance on a series of benchmarks, including predicting the thermochemical properties of equilibrium structures; predicting forces of far-from-equilibrium structures; maintaining physical soundness under static extreme deformations; and molecular dynamic stability under extreme temperatures and pressures. MP-ALOE shows strong performance on all of these benchmarks and is made public for the broader community to utilize.

Kuner, Matthew C↗

Biomimetic Control over Bimetallic Nanoparticle Structure and Activity via Peptide Capping Ligand Sequence

Here, the controlled design of bimetallic nanoparticles (BNPs) is a key goal in tailoring their catalytic properties. Recently, biomimetic pathways demonstrated potent control over the distribution of different metals within BNPs, but a direct understanding of the peptide effect on the compositional distribution at the interparticle and intraparticle levels remains lacking. We synthesized two sets of PtAu systems with two peptides and correlated their structure, composition, and distributions with the catalytic activity. Structural and compositional analyses were performed by a combined machine learning-assisted refinement of X-ray absorption spectra and Z-contrast measurements by scanning transmission electron microscopy. The difference in the catalytic activities between nanoparticles synthesized with different peptides was attributed to the details of interparticle distribution of Pt and Au across these markedly heterogeneous systems, comprising Pt-rich, Au-rich, and Au core/Pt shell nanoparticles. The total amount of Pt in the shells of the BNPs was proposed to be the key catalytic activity descriptor. This approach can be extended to other systems of metals and peptides to facilitate the targeted design of catalysts with the desired activity.

36 MATERIALS SCIENCE↗

DETECTING FIRE WITH MACHINE LEARNING-ENABLED VISUAL MONITORING FOR NUCLEAR POWER PLANT ENVIRONMENTS

Nuclear power plants are experiencing significant cost challenges to remain competitive with other energy-generation utilities. Unlike other industries, the cost of operation and maintenance activities is mostly attributed to workforce costs. To mitigate this, nuclear power plant stakeholders are increasingly interested in the development and deployment of machine learning methods to potentially automate or augment manually intensive tasks to reduce costs, especially for monitoring activities. One monitoring function that is visually demanding and that can occur frequently to meet the requirements of a fire protection program is visually monitoring an area for fire occurrence. Currently, fire watch activities consist of a worker physically stationed at a given location with the sole responsibility of observing a given area to ensure a fire is detected and mitigated promptly. This effort focused on the development and evaluation of a suitable deep convolutional neural network to classify individual video frames at a sub-second frequency for the occurrence of “fire” and “no fire” in varying industrial environments similar to nuclear power plants. It is believed that a trained neural network model could be integrated with existing facility video surveillance camera feeds to generate alerts when fire inferences occur in individual frames captured at sub-second temporal resolutions. Extensive effort was dedicated to identifying and curating suitable imagery training data representing varying environments and scene settings with and without flame features to maximize generalization in nuclear power plant environments. The data collection effort resulted in the aggregation of a large, labeled image library exceeding 12,000 images to support model training for diverse industrial environments. A deep neural network model incorporating parallel multi-scale capabilities was developed and trained to support accurate image-based detection of flame incidents of varying sizes and spectral feature properties within heterogeneous scenes. Analysis results show that the trained model can achieve high inference accuracy despite heterogeneous scene environments and components. Testing accuracy exceeded 95.0 percent with very low false positive and false negative inferences.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Energy-Efficient Self-Organization and Swarm Behavior in Active Matter

Living systems have the unique ability to form hierarchical assemblies, in which individual constituents can perform tasks cooperatively and emergently. Harnessing such properties is a long-standing challenge for the rational design of dynamic materials, that can respond to their environment, communicate with one another, and undergo a rapid, reversible, assembly through the transduction of energy. Recent developments in the design of smart and active colloidal building blocks have led to tremendous breakthroughs, with, for instance, the onset of synthetic photoactivated active assemblies. In this project, we develop a combined experimental, computational, theoretical and Machine Learning framework to shed light on the physical underpinnings of such assembly processes and program the assembly of smart active materials.

36 MATERIALS SCIENCE↗

Machine learning-guided design, synthesis, and characterization of atomically dispersed electrocatalysts

The recent integration of machine learning into materials design has revolutionized the understanding of structure–property relationships and optimization of material properties beyond the trial-and-error paradigm. On one hand, machine learning has significantly accelerated the development of atomically dispersed metal-nitrogen-carbon (M-N-C) electrocatalysts, which traditionally heavily relied on heuristic approaches. On the other hand, the primary challenge of leveraging machine learning to expedite M-N-C materials discovery lies in the cost associated with data collection. Here, we review recent machine learning integration strategies for M-N-C catalyst development, including discussions on the typical algorithms such as symbolic regression and convolutional neural networks employed for the theoretical design, synthesis optimization via active learning, and advanced microscopy characterization. Subsequently, we provide our perspective on potential near-future directions for furthering machine learning-assisted development of new M-N-C catalysts and elucidating the complex physicochemical mechanisms governing the selectivity, activity, and durability in this class of materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine‐Learning‐Driven Exploration of Surface Reconstructions of Reduced Rutile TiO 2

Titanium dioxide (TiO 2 ) is widely used as a catalyst support due to its stability, tunable electronic properties, and surface oxygen vacancies, which are crucial for catalytic processes such as the reverse water-gas shift (RWGS) reaction. Reduced TiO 2 surfaces undergo complex surface reconstructions that endow unique properties but are computationally challenging to describe. In this study, we utilize machine-learning interatomic potentials (MLIPs) integrated with an active-learning workflow to efficiently explore reduced rutile TiO 2 surfaces. This approach enabled the prediction of a phase diagram as a function of oxygen chemical potential, revealing a variety of reconstructed phases, including a previously unreported subsurface shear plane structure. We further investigate the electronic properties of these surfaces and validate our results by comparing experimental and theoretical high-resolution transmission electron microscopy (HRTEM). Our findings provide new insights into how extreme surface reductions influence the structural and electronic properties of TiO 2 , with potential implications for catalyst design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Development of neural network force fields for corrosion studies

To fully understand the chemistry and physics of corrosion, novel methods of simulation must be developed. One approach is designing machine learning (ML) algorithms integrated with density functional theory to develop adaptive force fields to gain insight into corrosion behavior namely at the surface of metal oxides. Current methods of modeling corrosion are slow due to the computational cost of resolving both reaction mechanics and mass transport processes. Machine learning methods can be implemented to obtain structure-activity relationships at both the molecular and bulk scale while still retaining the accuracy of density functional theory (DFT) and significantly decreasing the time needed for simulations of complex chemical processes in the various environments of corrosion. Multiscale models are needed for corrosion studies to fully understand its processes not only at the atomic length scale (chemical bonding, energies, and forces), but also at the nano and meso length scales (solid-state physics and material science processes). Current methods of study include DFT, molecular dynamics, and Monte Carlo. The limitation of DFT is that only a small number of atoms or molecules can be simulated at that level of theory. Density functional theory is used to study the electronic structure of atoms and molecules, and calculate the force component of each atom. However, these calculations are limited to about 1000 atoms. Custom periodic boundary conditions (PBC) can be used to describe the various environments and defects that affect the atomic forces to produce a large data set from which a training set can be derived. Machine learning can be utilized to overcome the barrier of modeling macroscopic and multi-scale processes from ab initio calculations through the development of adaptive force fields. Local environments determine the atomic forces of a given system, therefore adaptive force fields must be created to produce reliable quantum mechanical calculations. This can be achieved by developing a learning algorithm that uses the mapped atomic forces or fingerprint as an input to produce energies and magnetic moments as output. A systematic approach was used to begin to build a data set in order to accurately describe the atomic forces in various environments. In Figure 4 below, a simple PBC cell of Fe{sub 2}O{sub 3} was first optimized. A surface optimization was performed next, followed by a hydroxylated surface optimization. Once this calculation has converged, the adsorption of halide species to the hydroxylated surface will be investigated. TensorFlow is an open source platform for machine learning developed by Google. Using a high level application program interface (API) such as Keras allows for building and training ML models easily in a number of different environments and languages. For this project, a neural network was developed within Anaconda in Python. Future Work: Further development of reference data set; Refining neural network and learning algorithm; Fingerprinting atomic environment to enable mapping of atomic force components; Choosing appropriate training set from reference data; Learning from training set and enabling non-linear mapping of training set fingerprints and the atomic forces; Estimation of uncertainty to identify ranges of outside applicability; Testing and analysis of molecular dynamic simulations.

36 MATERIALS SCIENCE↗

Developing predictive models for µ opioid receptor binding using machine learning and deep learning techniques

Opioids exert their analgesic effect by binding to the µ opioid receptor (MOR), which initiates a downstream signaling pathway, eventually inhibiting pain transmission in the spinal cord. However, current opioids are addictive, often leading to overdose contributing to the opioid crisis in the United States. Therefore, understanding the structure-activity relationship between MOR and its ligands is essential for predicting MOR binding of chemicals, which could assist in the development of non-addictive or less-addictive opioid analgesics. This study aimed to develop machine learning and deep learning models for predicting MOR binding activity of chemicals. Chemicals with MOR binding activity data were first curated from public databases and the literature. Molecular descriptors of the curated chemicals were calculated using software Mold2. The chemicals were then split into training and external validation datasets. Random forest, k-nearest neighbors, support vector machine, multi-layer perceptron, and long short-term memory models were developed and evaluated using 5-fold cross-validations and external validations, resulting in Matthews correlation coefficients of 0.528–0.654 and 0.408, respectively. Furthermore, prediction confidence and applicability domain analyses highlighted their importance to the models’ applicability. Our results suggest that the developed models could be useful for identifying MOR binders, potentially aiding in the development of non-addictive or less-addictive drugs targeting MOR.

Research & Experimental Medicine↗

Machine learning for the redox potential prediction of molecules in organic redox flow battery

Here, organic redox flow batteries (ORFB) are recognized as an innovative technology for the large-scale storage of renewable energy. The redox potential of organic redox-active molecules plays a vital role in their performance. Advanced screening techniques like high-throughput experiment and machine learning (ML) have significantly enhanced organic material performance and transformed the field of ORFB. However, the scarcity of experimental data poses a considerable challenge for ML model development in this domain. In our study, we developed lightweight graph-based Gaussian process regression (GPR) models with GPU-accelerated marginalized graph kernel and hybrid kernel to predict the redox potentials of organic redox-active molecules for ORFBs, specifically focusing on small datasets. To evaluate model accuracy, we created a new experimental database of organic redox-active molecules by the data from hundreds of published papers and assembled previous computational datasets. We also considered some key parameters, such as pH conditions and solvent type, to assess their impact on redox potential prediction. Our GPR model predicted redox potentials with high accuracy across all datasets using minimal training data. The study provides powerful tools for molecule screening and design and delivers valuable guidance on designing training datasets for costly experiments.

25 ENERGY STORAGE↗

Computational catalyst discovery: Active classification through myopic multiscale sampling

We report the recent boom in computational chemistry has enabled several projects aimed at discovering useful materials or catalysts. We acknowledge and address two recurring issues in the field of computational catalyst discovery. First, calculating macro-scale catalyst properties is not straightforward when using ensembles of atomic-scale calculations [e.g., density functional theory (DFT)]. We attempt to address this issue by creating a multi-scale model that estimates bulk catalyst activity using adsorption energy predictions from both DFT and machine learning models. The second issue is that many catalyst discovery efforts seek to optimize catalyst properties, but optimization is an inherently exploitative objective that is in tension with the explorative nature of early-stage discovery projects. In other words, why invest so much time finding a “best” catalyst when it is likely to fail for some other, unforeseen problem? We address this issue by relaxing the catalyst discovery goal into a classification problem: “What is the set of catalysts that is worth testing experimentally?” Here, we present a catalyst discovery method called myopic multiscale sampling, which combines multiscale modeling with automated selection of DFT calculations. It is an active classification strategy that seeks to classify catalysts as “worth investigating” or “not worth investigating” experimentally. Our results show an ~7–16 times speedup in catalyst classification relative to random sampling. These results were based on offline simulations of our algorithm on two different datasets: a larger, synthesized dataset and a smaller, real dataset.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multireference Methods are Realistic and Useful Tools for Modeling Catalysis

Abstract Highly correlated systems, in particular those that include transition metals, are ubiquitous in catalysis. The significant static correlation found in such systems is often poorly accounted for using Kohn Sham density functional theory methods, as they are single determinantal in nature. Applications to catalysis of more rigorous and appropriate multiconfigurational methods have been reported in select instances, but their use remains rare. We discuss obstacles that hinder the routine application of multireference (MR) wave function theoretical calculations to catalytic systems and the current state of the art with respect to removing those obstacles.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Insights from machine learning of carbon electrodes for electric double layer capacitors

Recent years have witnessed the broad use of carbon electrodes for electric double layer capacitors (EDLCs) because of large surface area, high porosity and low cost. Whereas experimental investigations are mostly focused on the device performance, computational studies have been rarely concerned with electrochemical properties at conditions remote from equilibrium, limiting their direct applications to materials design. Through a comprehensive analysis of extensive experimental data with various machine-learning methods, we report herein quantitative correlations between the structural features of carbon electrodes and the in-operando behavior of EDLCs including energy and power density. Machine learning allows us to identify important characteristics of activated carbons useful to optimize their efficiency in energy storage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-driven building energy modeling with feature selection and active learning for data predictive control

Three gaps impede the development of cost-effective and accurate data-driven building energy modeling/models (DBEM) for energy forecasting and predictive control strategies. Gap 1: data bias is common in building operation data, but this topic is hardly studied in DBEM; Gap 2: high data dimensionality is common in DBEM, but a systematic and scalable methodology is lacking to solve the problem; Gap 3: the interactions between data bias and high dimensionality have not been systematically studied for DBEM and predictive control in buildings. In this work, to address the three gaps mentioned above, we develop a framework that integrates active learning and feature selection for DBEM used for whole building data predictive control (or DPC, which is a branch of model predictive control). The framework provides a systematic methodology and automatic workflow that starts with raw data from building automation systems to the establishment of data-driven energy models for DPC controllers. The developed strategies and framework are evaluated in a virtual testbed based on EnergyPlus and BCVTB. Improved performance and reduced computational complexity are observed from the DBEM built with the developed framework, as well as the DPC controller based on that DBEM, indicating the effectiveness of the developed framework.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Quantification of neural networks uncertainties with applications to SAFARI-1 axial neutron flux profiles

Deep Neural Networks (DNNs) have been widely used as a data-driven modelling tool in nuclear engineering. However, as a Machine Learning model, Artificial Neural Network (ANN) predictions are subjected to uncertainties originating from the noise in training data, incomplete coverage of the domain, and imperfect neural network architectures. In this work, we target at quantifying the prediction/approximation uncertainties of ANNs using Monte Carlo Dropout (MCD), as well as Bayesian Neural Networks (BNNs) which are solved by variational inference. With a demonstration problem in which neural networks are used to predict the assembly axial neutron flux profiles, the results have shown that the three different neural network models (regular DNNs, DNNs solved with MCD and BNNs) can produce results that agree very well among each other and with the measurement data, on cycles that are not used in the training process. Besides the excellent generalization capability, the uncertainty bands produced by MCD and BNN agree very well, and in general, they can fully envelop the noisy measurement data points. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗