Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

iRF v2.0

A predictive, stable, and interpretable machine learning tool: the iterative random forest algorithm (iRF). iRF discovers high-order interactions among variables with the same order of computational cost as random forests (RF). We have demonstrated the utility of iRF in several applications in the biological and environmental sciences. It is a general purpose machine learning framework for building "explainable" predictive engines.

Brown, JamesB.↗

Automated Machine Learning to Evaluate the Information Content of Tropospheric Trace Gas Columns for Fine Particle Estimates Over India: A Modeling Testbed

India is largely devoid of high-quality and reliable on-the-ground measurements of fine particulate matter (PM 2.5 ). Ground-level PM 2.5 concentrations are estimated from publicly available satellite Aerosol Optical Depth (AOD) products combined with other information. Prior research has largely overlooked the possibility of gaining additional accuracy and insights into the sources of PM using satellite retrievals of tropospheric trace gas columns. We evaluate the information content of tropospheric trace gas columns for PM 2.5 estimates over India within a modeling testbed using an Automated Machine Learning (AutoML) approach, which selects from a menu of different machine learning tools based on the data set. We then quantify the relative information content of tropospheric trace gas columns, AOD, meteorological fields, and emissions for estimating PM 2.5 over four Indian sub-regions on daily and monthly time scales. Our findings suggest that, regardless of the specific machine learning model assumptions, incorporating trace gas modeled columns improves PM 2.5 estimates. We use the ranking scores produced from the AutoML algorithm and Spearman’s rank correlation to infer or link the possible relative importance of primary versus secondary sources of PM 2.5 as a first step toward estimating particle composition. Our comparison of AutoML-derived models to selected baseline machine learning models demonstrates that AutoML is at least as good as user-chosen models. The idealized pseudo-observations (chemical-transport model simulations) used in this work lay the groundwork for applying satellite retrievals of tropospheric trace gases to estimate fine particle concentrations in India and serve to illustrate the promise of AutoML applications in atmospheric and environmental research.

Machine learning↗

Integrated Design of Ultradurable, Low CO 2 Alternative Binder Systems via Machine Learning

This ARPA-E project developed a machine learning tool to use in formulation design of cementitious binders for concrete having 50% less embodied CO 2 and possessing twice the durability compared to concrete based on ordinary portland cement (OPC) binders. The technical focus was on limestone/calcined clay cement (LC3), the leading replacement for OPC. Here, hierarchical machine learning (HML) was used to model the flowability, set time, strength, and durability of LC3 concrete. This methodology identifies latent variables derived from domain knowledge and empirical models that develop an accurate model for a response surface from small datasets. For the flowability metric, particle packing was a dominant factor, while strength and durability were both strongly determined by the fraction of metakaolin and the water:solids ratio. Under constraints of water:binder ratio, material performance metrics, embodied CO 2 , and cost per tonne of OPC, multi-objective optimization was used to design binders parameterized by the mineral composition replacing OPC, particle size distributions, and water:solids ratio. The trained algorithm was able to predict multiple mixes met these performance criteria, and experimental testing validated the predictions. The model demonstrated here is relevant for North America, where pure kaolin deposits are found broadly. The approach is being taken forward into commercial application by Ansatz AI, a materials informatics company founded by PI Washburn and co-PI Poczos. Through collaborations with the cement and concrete industry, and funding from SBIR programs, a commercial software will be developed in future research.

36 MATERIALS SCIENCE↗

Identification of carbohydrate gene clusters obtained from in vitro fermentations as predictive biomarkers of prebiotic responses

Prebiotic fibers are non-digestible substrates that modulate the gut microbiome by promoting expansion of microbes having the genetic and physiological potential to utilize those molecules. Although several prebiotic substrates have been consistently shown to provide health benefits in human clinical trials, responder and non-responder phenotypes are often reported. These observations had led to interest in identifying, a priori, prebiotic responders and non-responders as a basis for personalized nutrition. In this study, we conducted in vitro fecal enrichments and applied shotgun metagenomics and machine learning tools to identify microbial gene signatures from adult subjects that could be used to predict prebiotic responders and non-responders. Using short chain fatty acids as a targeted response, we identified genetic features, consisting of carbohydrate active enzymes, transcription factors and sugar transporters, from metagenomic sequencing of in vitro fermentations for three prebiotic substrates: xylooligosacharides, fructooligosacharides, and inulin. A machine learning approach was then used to select substrate-specific gene signatures as predictive features. These features were found to be predictive for XOS responders with respect to SCFA production in an in vivo trial. Our results confirm the bifidogenic effect of commonly used prebiotic substrates along with inter-individual microbial responses towards these substrates. We successfully trained classifiers for the prediction of prebiotic responders towards XOS and inulin with robust accuracy (≥ AUC 0.9) and demonstrated its utility in a human feeding trial. Overall, the findings from this study highlight the practical implementation of pre-intervention targeted profiling of individual microbiomes to stratify responders and non-responders.

59 BASIC BIOLOGICAL SCIENCES↗

Parameters, Properties, and Process: Conditional Neural Generation of Realistic SEM Imagery Toward ML-Assisted Advanced Manufacturing

Abstract The research and development cycle of advanced manufacturing processes traditionally requires a large investment of time and resources. Experiments can be expensive and are hence conducted on relatively small scales. This poses problems for typically data-hungry machine learning tools which could otherwise expedite the development cycle. We build upon prior work by applying conditional generative adversarial networks (GANs) to scanning electron microscope (SEM) imagery from an emerging advanced manufacturing process, shear-assisted processing and extrusion (ShAPE). We generate realistic images conditioned on temper and either experimental parameters or material properties. In doing so, we are able to integrate machine learning into the development cycle, by allowing a user to immediately visualize the microstructure that would arise from particular process parameters or properties. This work forms a technical backbone for a fundamentally new approach for understanding manufacturing processes in the absence of first-principle models. By characterizing microstructure from a topological perspective, we are able to evaluate our models’ ability to capture the breadth and diversity of experimental scanning electron microscope (SEM) samples. Our method is successful in capturing the visual and general microstructural features arising from the considered process, with analysis highlighting directions to further improve the topological realism of our synthetic imagery.

36 MATERIALS SCIENCE↗

A representation-independent electronic charge density database for crystalline materials

Abstract In addition to being the core quantity in density-functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With the increasing prevalence of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

A representation-independent electronic charge density database for crystalline materials

In addition to being the core quantity in density functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With increasing utilization of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

Manufacturing of Fabric Electrodes using a High-Throughput Screening Platform for Redox Flow Batteries

The objective of this project is to establish a new manufacturing methodology with machine learning- based high-throughput screening for the design and development of hierarchical structured, high-performance fabric electrodes for redox flow batteries (RFBs). The end goal of the project is to design and manufacture fabric electrodes for RFB applications that can provide 250 mA/cm2 current density operation for 100-cycles with 80% average energy efficiency. This was accomplished by first examining the structure-performance-property linkages of the electrodes provided by our partner, AvCarb. The electrodes’ microstructure was characterized by determining their pore size distribution, tortuosity, specific surface area, and porosity. The ohmic, charge transfer and mass transfer resistances were then calculated using electrochemical impedance spectroscopy. Carbon cloth electrodes showed the greatest resistance, which was dominated by charge transfer resistance, which we believe is related to the surface functionalization. Full cell cycling was used in order to determine the area specific resistance and energy efficiency of the cells. All of this experimental data and the results of the mathematical model (to increase the amount of inputs with parametric sweeping) were used to develop a machine learning-based model for the design of high-performance fabric electrodes. Using the results from the machine learning tool, optimized electrodes were fabricated by AvCarb. The ohmic, charge transfer and mass transfer resistances for these new electrodes were measured, and both performed better than any of the initial samples which had been provided by AvCarb.

25 ENERGY STORAGE↗

Use of Machine Learning to Reduce Uncertainties in Particle Number Concentration and Aerosol Indirect Radiative Forcing Predicted by Climate Models

The radiative forcing of anthropogenic aerosols associated with aerosol–cloud interactions (RF(sub aci)) remains the largest source of uncertainty in climate prediction. The calculation of particle number concentration (PNC), one of the critical parameters affecting RF(sub aci), is generally simplified in climate models. Here we employ outputs from long-term (30-years) simulations of a global size-resolved (sectional) aerosol microphysics model and a machine-learning tool to develop a Random Forest Regression Model (RFRM) for PNC. We have implemented the PNC RFRM in GISS-ModelE2.1 with a mass-based One-Moment Aerosol module, which is one of CMIP6 models. Compared to the default setting, the GISS-ModelE2.1 simulation based on RFRM reduces the changes of cloud droplet number concentration associated with anthropogenic emissions, and decreases the RF(sub aci) from −1.46 W⋅m(exp −2) to −1.11 W⋅m(exp −2). This work highlights a promising approach based on machine learning to reduce uncertainties of climate models in predicting PNC and RF(sub aci) without compromising their computing efficiency.

Radiative forcing↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Forensic Analysis of SOHO Router Binaries

Small Office/Home Office (SOHO) routers are used by millions of consumers across the United States, and are commensurately vulnerable. Forensic analysis of SOHO router firmware helps to understand and mitigate those vulnerabilities. This poster focused particularly on analysis of BusyBox executables, a software suite that provides several Unix utilities in a single file. Three main tools were used to analyze the binaries. BinWalk was used to extract the files, but also to build entropy graphs, extract Linux kernel images, and identify CPU architectures; WiiBin processed the binaries to find endianness, architecture, the percent compressed/encrypted, and compiler data; and @DisCo, a machine learning tool used to determine function similarity in disassembled binaries, analyzed similarities and determined versions of extracted BusyBox files from each router. These tools found that venders from all five routers utilized the same version of the BusyBox software across different firmware updates, demonstrating the importance of constant firmware scrutiny to protect against security vulnerabilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Coupling 1D xRAGE simulations with machine learning for graded inner shell design optimization in double shell capsules

Advances in machine learning provide the ability to leverage data from expensive simulations of high-energy-density experiments to significantly cut down on computational time and costs associated with the search for optimal target designs. This study presents an application of cutting-edge Bayesian optimization methods to the one-dimensional (1D) design optimization of double shell graded layer targets for inertial confinement fusion experiments. This investigation attempts to reduce hydrodynamic instabilities while retaining high yields for future NIF experiments. Machine learning methods can use predictive physics simulations to identify graded layer designs from within the vast design space that demonstrate high predicted performance, including novel designs with high uncertainty in performance that may hold unexpected promise. By applying machine learning tools to the simulation design, we map the trade-off between 1D yield and instability, specifically isolating parameter ranges, which maintain high performance while showing significantly improved Rayleigh–Taylor stability over the point design. Furthermore, the groundwork laid in this study will be a useful design tool for future NIF experiments with graded layer targets.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

Autodifferentiable Ensemble Kalman Filters

Data assimilation is concerned with sequentially estimating a temporally evolving state. This task, which arises in a wide range of scientific and engineering applications, is particularly challenging when the state is high-dimensional and the state-space dynamics are unknown. This paper introduces a machine learning framework for learning dynamical systems in data assimilation. Here, our auto-differentiable ensemble Kalman filters (AD-EnKFs) blend ensemble Kalman filters for state recovery with machine learning tools for learning the dynamics. In doing so, AD-EnKFs leverage the ability of ensemble Kalman filters to scale to high-dimensional states and the power of automatic differentiation to train high-dimensional surrogate models for the dynamics. Numerical results using the Lorenz-96 model show that AD-EnKFs outperform existing methods that use expectation-maximization or particle filters to merge data assimilation and machine learning. In addition, AD-EnKFs are easy to implement and require minimal tuning.

autodifferentiation↗

On-the-fly closed-loop materials discovery via Bayesian active learning

Active learning—the field of machine learning (ML) dedicated to optimal experiment design—has played a part in science as far back as the 18th century when Laplace used it to guide his discovery of celestial mechanics. In this work, we focus a closed-loop, active learning-driven autonomous system on another major challenge, the discovery of advanced materials against the exceedingly complex synthesis-processes-structure-property landscape. We demonstrate an autonomous materials discovery methodology for functional inorganic compounds which allow scientists to fail smarter, learn faster, and spend less resources in their studies, while simultaneously improving trust in scientific results and machine learning tools. This robot science enables science-over-the-network, reducing the economic impact of scientists being physically separated from their labs. The real-time closed-loop, autonomous system for materials exploration and optimization (CAMEO) is implemented at the synchrotron beamline to accelerate the interconnected tasks of phase mapping and property optimization, with each cycle taking seconds to minutes. We also demonstrate an embodiment of human-machine interaction, where human-in-the-loop is called to play a contributing role within each cycle. This work has resulted in the discovery of a novel epitaxial nanocomposite phase-change memory material.

36 MATERIALS SCIENCE↗

Data-Driven Strategies for Accelerated Materials Design

The ongoing revolution of the natural sciences by the advent of machine learning and artificial intelligence sparked significant interest in the material science community in recent years. The intrinsically high dimensionality of the space of realizable materials makes traditional approaches ineffective for large-scale explorations. Modern data science and machine learning tools developed for increasingly complicated problems are an attractive alternative. An imminent climate catastrophe calls for a clean energy transformation by overhauling current technologies within only several years of possible action available. Tackling this crisis requires the development of new materials at an unprecedented pace and scale. For example, organic photovoltaics have the potential to replace existing silicon-based materials to a large extent and open up new fields of application. In recent years, organic light-emitting diodes have emerged as state-of-the-art technology for digital screens and portable devices and are enabling new applications with flexible displays. Reticular frameworks allow the atom-precise synthesis of nanomaterials and promise to revolutionize the field by the potential to realize multifunctional nanoparticles with applications from gas storage, gas separation, and electrochemical energy storage to nanomedicine. In the recent decade, significant advances in all these fields have been facilitated by the comprehensive application of simulation and machine learning for property prediction, property optimization, and chemical space exploration enabled by considerable advances in computing power and algorithmic efficiency. In this Account, we review the most recent contributions of our group in this thriving field of machine learning for material science. We start with a summary of the most important material classes our group has been involved in, focusing on small molecules as organic electronic materials and crystalline materials. Specifically, we highlight the data-driven approaches we employed to speed up discovery and derive material design strategies. Subsequently, our focus lies on the data-driven methodologies our group has developed and employed, elaborating on high-throughput virtual screening, inverse molecular design, Bayesian optimization, and supervised learning. We discuss the general ideas, their working principles, and their use cases with examples of successful implementations in data-driven material discovery and design efforts. Furthermore, we elaborate on potential pitfalls and remaining challenges of these methods. Finally, we provide a brief outlook for the field as we foresee increasing adaptation and implementation of large scale data-driven approaches in material discovery and design campaigns.

36 MATERIALS SCIENCE↗

Using Satellite Soil Moisture and Rainfall in the Landslide Hazard Assessment for Situational Awareness System

The Landslide Hazard Assessment for Situational Awareness system(LHASA)gives a global view of landslide hazard in nearly real time. Currently, it is being upgraded from version 1 to version 2, which entails improvements along several dimensions. These include the incorporation of new predictors, machine learning, and new event-based landslide inventories. As a result, LHASA version 2 substantially improves on the prior performanceand introduces a probabilistic element to the global landslide nowcast. Data from the soil moisture active-passive (SMAP) satellite has been assimilated into a globally consistent data product with a latency less than 3 days, known as SMAP Level 4. In LHASA, thesedata representthe antecedent conditions prior to landslide-triggering rainfall. In some cases, soil moisture may have accumulated over aperiod of many months. The model behind SMAP Level 4 also estimates the amount of snow on the ground, which is an important factor in some landslide events. LHASA also incorporates this information as an antecedent condition that modulates the response torainfall. Slope, lithology, and active faults were also used as predictor variables. These factors can have a strong influence on where landslides initiate.LHASA relies on precipitation estimates from the Global Precipitation Measurement mission to identify the locations where landslides are most probable. The low latency and consistent global coverage of these data make them ideal for real-time applications at continental to global scales. LHASA relies primarily on rainfall from the last 24 hours to spothazardous sites, which is rescaled by the local 99thpercentile rainfall.However, the multi-day latency of SMAP requires the use of a 2-day antecedent rainfall variable to represent the accumulation of rain between the antecedent soil moisture and current rainfall. LHASA merges these predictors with XGBoost, a commonly used machine-learning tool, relying on historical landslide inventories to develop the relationship between landslide occurrence and various risk factors. The resulting model relies heavily on current daily rainfall, but other factors also play an important role. LHASA outputsthe probability oflandslide occurrence ona grid of roughly one kilometer over all continents from 60 North to 60 South latitude. Evaluation over the period 2019-2020 showsthat LHASA version 2 doubles the accuracy of the global landslide nowcast without increasing the global false alarm rate. LHASA also identifies the areas where the human exposure to landslide hazard is most intense. Landslide hazard is divided into 4 levels: minimal, low, moderate, and high. Next, the number of persons and the length of major roads (primary and secondary roads)within each of these areas is calculated for every second-level administrative district (county). These results can be viewedthrough a web portal hosted at the Goddard Space Flight Center. In addition, users can download daily hazard and exposure data.LHASAversion 2uses machine learning and satellite data to identify areas of probable landslide hazard within hours of heavy rainfall. Itsglobal maps are significantly more accurate, and it now includes rapid estimates of exposed populations and infrastructure. In addition, a forecast mode will be implemented soon.

Thomas Stanley↗