Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Elucidating Abnormal Grain Growth in Thermomagnetic Processed Materials with Transfer Learning and Reinforcement Learning

The goal of this research program is to establish the mechanism governing local grain boundary motion, which is needed to design and process desirable microstructures for better performance, by identifying the relative contributions of grain boundary (GB) energy and mobility to grain growth. Classical models for grain growth assume that the primary mechanism for reducing the total interfacial energy is area reduction and that GB restructuring is not significant. This assumption implies that grain growth is locally driven by curvature. However, recent experimental observations using new non-destructive 3D x-ray diffraction microscopy techniques (3D-XRM) reveal that classic descriptors (i.e., curvature, number of neighbors, grain size) do not predict real grain growth. Instead, local GB motion appears to be governed by its energy relative to its neighbors such that low-energy boundaries replace those of higher energy. However, simulations that incorporate GB energy anisotropy still fail to reproduce these observations. These discrepancies suggest that the common assumption for grain growth theory must be re-examined to predict and, thus, control microstructure evolution in real polycrystals. A significant challenge to testing this assumption is due to anisotropic GB mobility. Mobility may cause abnormal grain growth or affect the final grain shapes or growth rate but its true contributions are unknown because it is difficult to measure. For example, observations in Fe have found that grains associated with high energy and high mobility boundaries tend to experience abnormal grain growth, whereas abnormal grain growth is associated with low energy and high mobility boundaries in alumina. As mobility and energy both control GB motion, it is challenging to isolate the local driving forces necessary to test the common assumption that the primary mechanism is area reduction. The novelty of this work is the use of machine learning tools to capture GB mobility and energy from 3D-XRM measurements in polycrystals to test the common assumption used in grain growth models. Machine learning can capture high-order correlations in dynamic systems like those found in the evolving GB topology. The PIs have developed a physics-regularized interpretable machine learning microstructure evolution (PRIMME) model that accurately replicates the grain growth behavior of its trained data set.

36 MATERIALS SCIENCE↗

Human Factors for Machine Learning in Astronomy

In this work, we present a collection of human-centered pitfalls that can occur when using machine learning tools and techniques in modern astronomical research, and we recommend best practices in order to mitigate these pitfalls. Human concerns affect the adoption and evolution of machine learning (ML) techniques in both existing workflows and work cultures. We use current and future surveys such as ZTF and LSST, the data that they collect, and the techniques implemented to process that data as examples of these challenges and the potential application of these best practices, with the ultimate goal of maximizing the discovery potential of these surveys.

astronomy, machine learning, human factors, trust,↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

End-to-End Automated Segmentation Framework for Four-Dimensional Scanning Transmission Electron Microscopy Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) is powerful for rapidly characterizing arrays of nanoparticles produced via high-throughput synthesis. However, such 4D-STEM datasets typically contain thousands of nanoparticles, each characterized by thousands of diffraction patterns spatially distributed across the nanoparticle, necessitating efficient and comprehensive analysis. We propose an end-to-end segmentation framework to automatically segment each nanoparticle into regions with distinct composition/orientation of crystal grains, using only the 4D-STEM data. Bragg disk information is extracted in a physics-informed manner from the diffraction patterns at each spatial location and combined with the real space coordinates to form feature vectors. These feature vectors are then used as inputs to a Gaussian mixture model (GMM) to segment the nanoparticle into distinct regions. We also develop two visualization tools based on the GMM outputs to infer the interface transition and the degree of superposition. Our framework comprehensively integrates machine learning tools and physics knowledge, and provides a basis for substantially compressing enormous 4D-STEM datasets, e.g., by replacing the full 4D-STEM dataset for each nanoparticle with only a single set of Bragg disk features for each distinct crystal grain identified in the nanoparticle. In this article, we demonstrate the power of our framework by presenting results for real, complex datasets.

47 OTHER INSTRUMENTATION↗

A machine learning based approach to online electron reconstruction at CLAS12

Online reconstruction is key for monitoring purposes and real time analysis in High Energy and Nuclear Physics experiments. A necessary component of reconstruction algorithms is particle identification that combines information left by a particle passing through several detector components to identify the particle’s type. Of particular interest to electro-production Nuclear Physics experiments such as CLAS12 is electron identification which is used to trigger data recording. A machine learning approach was developed for CLAS12 to reconstruct and identify electrons by combining raw signals at the data acquisition level from several detector components. Here, this approach achieves an electron identification purity above 75% whilst retaining an efficiency close to 100%. The machine learning tools are capable of running at high rates exceeding the data acquisition rates and will allow electron reconstruction in real-time. This work enhances online analyses and monitoring and can contribute to improved triggering at CLAS12. This machine learning driven approach will also be crucial for experiments aiming to transition to streaming readout operations where online reconstruction will be a key component of the data taking paradigm.

Artificial intelligence↗

Searching for Short-Lived Particles Using the FASER Neutrino Detector

Short-lived particles such as $\tau$ leptons and charm hadrons produced from high-energy neutrino interactions can provide key insights into the Standard Model and beyond. In this study, initial results of the performance of dedicated search tools developed for FASERν are presented. For charged charm searches, all daughter tracks are identified for ~3/4 of the events. Advanced yet reliable machine learning tools are used for D0 detection. For $\tau$ searches, a signal / background ratio of 3.8 is achieved. These results are a major step towards discovering the first charm hadrons and $\tau$ leptons produced from collider neutrinos.

Thor, Simon [CERN] (ORCID:000000029183526X)↗

iRF v2.0

A predictive, stable, and interpretable machine learning tool: the iterative random forest algorithm (iRF). iRF discovers high-order interactions among variables with the same order of computational cost as random forests (RF). We have demonstrated the utility of iRF in several applications in the biological and environmental sciences. It is a general purpose machine learning framework for building "explainable" predictive engines.

Brown, JamesB.↗

Integrated Design of Ultradurable, Low CO 2 Alternative Binder Systems via Machine Learning

This ARPA-E project developed a machine learning tool to use in formulation design of cementitious binders for concrete having 50% less embodied CO 2 and possessing twice the durability compared to concrete based on ordinary portland cement (OPC) binders. The technical focus was on limestone/calcined clay cement (LC3), the leading replacement for OPC. Here, hierarchical machine learning (HML) was used to model the flowability, set time, strength, and durability of LC3 concrete. This methodology identifies latent variables derived from domain knowledge and empirical models that develop an accurate model for a response surface from small datasets. For the flowability metric, particle packing was a dominant factor, while strength and durability were both strongly determined by the fraction of metakaolin and the water:solids ratio. Under constraints of water:binder ratio, material performance metrics, embodied CO 2 , and cost per tonne of OPC, multi-objective optimization was used to design binders parameterized by the mineral composition replacing OPC, particle size distributions, and water:solids ratio. The trained algorithm was able to predict multiple mixes met these performance criteria, and experimental testing validated the predictions. The model demonstrated here is relevant for North America, where pure kaolin deposits are found broadly. The approach is being taken forward into commercial application by Ansatz AI, a materials informatics company founded by PI Washburn and co-PI Poczos. Through collaborations with the cement and concrete industry, and funding from SBIR programs, a commercial software will be developed in future research.

36 MATERIALS SCIENCE↗

Identification of carbohydrate gene clusters obtained from in vitro fermentations as predictive biomarkers of prebiotic responses

Prebiotic fibers are non-digestible substrates that modulate the gut microbiome by promoting expansion of microbes having the genetic and physiological potential to utilize those molecules. Although several prebiotic substrates have been consistently shown to provide health benefits in human clinical trials, responder and non-responder phenotypes are often reported. These observations had led to interest in identifying, a priori, prebiotic responders and non-responders as a basis for personalized nutrition. In this study, we conducted in vitro fecal enrichments and applied shotgun metagenomics and machine learning tools to identify microbial gene signatures from adult subjects that could be used to predict prebiotic responders and non-responders. Using short chain fatty acids as a targeted response, we identified genetic features, consisting of carbohydrate active enzymes, transcription factors and sugar transporters, from metagenomic sequencing of in vitro fermentations for three prebiotic substrates: xylooligosacharides, fructooligosacharides, and inulin. A machine learning approach was then used to select substrate-specific gene signatures as predictive features. These features were found to be predictive for XOS responders with respect to SCFA production in an in vivo trial. Our results confirm the bifidogenic effect of commonly used prebiotic substrates along with inter-individual microbial responses towards these substrates. We successfully trained classifiers for the prediction of prebiotic responders towards XOS and inulin with robust accuracy (≥ AUC 0.9) and demonstrated its utility in a human feeding trial. Overall, the findings from this study highlight the practical implementation of pre-intervention targeted profiling of individual microbiomes to stratify responders and non-responders.

59 BASIC BIOLOGICAL SCIENCES↗

Parameters, Properties, and Process: Conditional Neural Generation of Realistic SEM Imagery Toward ML-Assisted Advanced Manufacturing

Abstract The research and development cycle of advanced manufacturing processes traditionally requires a large investment of time and resources. Experiments can be expensive and are hence conducted on relatively small scales. This poses problems for typically data-hungry machine learning tools which could otherwise expedite the development cycle. We build upon prior work by applying conditional generative adversarial networks (GANs) to scanning electron microscope (SEM) imagery from an emerging advanced manufacturing process, shear-assisted processing and extrusion (ShAPE). We generate realistic images conditioned on temper and either experimental parameters or material properties. In doing so, we are able to integrate machine learning into the development cycle, by allowing a user to immediately visualize the microstructure that would arise from particular process parameters or properties. This work forms a technical backbone for a fundamentally new approach for understanding manufacturing processes in the absence of first-principle models. By characterizing microstructure from a topological perspective, we are able to evaluate our models’ ability to capture the breadth and diversity of experimental scanning electron microscope (SEM) samples. Our method is successful in capturing the visual and general microstructural features arising from the considered process, with analysis highlighting directions to further improve the topological realism of our synthetic imagery.

36 MATERIALS SCIENCE↗

A representation-independent electronic charge density database for crystalline materials

Abstract In addition to being the core quantity in density-functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With the increasing prevalence of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

A representation-independent electronic charge density database for crystalline materials

In addition to being the core quantity in density functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With increasing utilization of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

Manufacturing of Fabric Electrodes using a High-Throughput Screening Platform for Redox Flow Batteries

The objective of this project is to establish a new manufacturing methodology with machine learning- based high-throughput screening for the design and development of hierarchical structured, high-performance fabric electrodes for redox flow batteries (RFBs). The end goal of the project is to design and manufacture fabric electrodes for RFB applications that can provide 250 mA/cm2 current density operation for 100-cycles with 80% average energy efficiency. This was accomplished by first examining the structure-performance-property linkages of the electrodes provided by our partner, AvCarb. The electrodes’ microstructure was characterized by determining their pore size distribution, tortuosity, specific surface area, and porosity. The ohmic, charge transfer and mass transfer resistances were then calculated using electrochemical impedance spectroscopy. Carbon cloth electrodes showed the greatest resistance, which was dominated by charge transfer resistance, which we believe is related to the surface functionalization. Full cell cycling was used in order to determine the area specific resistance and energy efficiency of the cells. All of this experimental data and the results of the mathematical model (to increase the amount of inputs with parametric sweeping) were used to develop a machine learning-based model for the design of high-performance fabric electrodes. Using the results from the machine learning tool, optimized electrodes were fabricated by AvCarb. The ohmic, charge transfer and mass transfer resistances for these new electrodes were measured, and both performed better than any of the initial samples which had been provided by AvCarb.

25 ENERGY STORAGE↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Forensic Analysis of SOHO Router Binaries

Small Office/Home Office (SOHO) routers are used by millions of consumers across the United States, and are commensurately vulnerable. Forensic analysis of SOHO router firmware helps to understand and mitigate those vulnerabilities. This poster focused particularly on analysis of BusyBox executables, a software suite that provides several Unix utilities in a single file. Three main tools were used to analyze the binaries. BinWalk was used to extract the files, but also to build entropy graphs, extract Linux kernel images, and identify CPU architectures; WiiBin processed the binaries to find endianness, architecture, the percent compressed/encrypted, and compiler data; and @DisCo, a machine learning tool used to determine function similarity in disassembled binaries, analyzed similarities and determined versions of extracted BusyBox files from each router. These tools found that venders from all five routers utilized the same version of the BusyBox software across different firmware updates, demonstrating the importance of constant firmware scrutiny to protect against security vulnerabilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Coupling 1D xRAGE simulations with machine learning for graded inner shell design optimization in double shell capsules

Advances in machine learning provide the ability to leverage data from expensive simulations of high-energy-density experiments to significantly cut down on computational time and costs associated with the search for optimal target designs. This study presents an application of cutting-edge Bayesian optimization methods to the one-dimensional (1D) design optimization of double shell graded layer targets for inertial confinement fusion experiments. This investigation attempts to reduce hydrodynamic instabilities while retaining high yields for future NIF experiments. Machine learning methods can use predictive physics simulations to identify graded layer designs from within the vast design space that demonstrate high predicted performance, including novel designs with high uncertainty in performance that may hold unexpected promise. By applying machine learning tools to the simulation design, we map the trade-off between 1D yield and instability, specifically isolating parameter ranges, which maintain high performance while showing significantly improved Rayleigh–Taylor stability over the point design. Furthermore, the groundwork laid in this study will be a useful design tool for future NIF experiments with graded layer targets.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗