Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

NASA Pilot-Engaged Expert Response Using IBM Watson Technology: Prototype Evaluation of Knowledge Retrieval System

NASA Langley Research Center and IBM have been investigating the use of IBM Watson technology in aerospace research and development. One application of Watson technology is the Pilot-Engaged Expert Response (PEER) use case. The PEER system is envisioned as an in-cockpit advisor that will act as a source of situationally-relevant information for pilots and other flight crew members to assist in decision making about real-time events and situations that arise in the course of aircraft operations. PEER will make available vast stores of knowledge and information quickly and directly, putting important informational resources where they are needed most. IBM has worked with NASA to develop an architecture and articulate a roadmap for the development of the PEER system. That vision is built around Watson Discovery Advisor (WDA) software solution, derived from IBM's Jeopardy!-winning automatic question answering system. PEER makes use of WDA's sophisticated question-answering capabilities as its core, adding important User Interface components and other customizations for the cockpit environment, including communication with flight systems and other external data sources. The development plan for PEER includes four development stages, with the current project constituting the first phase. In this project, a prototype instance of PEER was successfully adapted to the aviation domain, enabling users to ask questions about aviation topics and receive useful and accurate answers to these questions. Major tasks accomplished include the development of procedures for domain adaptation through automatic lexicon extraction from domain glossaries; generation of question-answer training data which was used to train the system; and assessment of the effectiveness of domain adaptation, which showed a dramatic improvement in the ability of the PEER system to answer domain-relevant questions. In addition, the vision for the PEER system was pushed forward by the articulation of a plan for the automatic enhancement of question-answering with contextual information. This initial phase focused on two main goals: 1) the targeted domain adaptation of the underlying WDA system to the aviation domain; and, 2) the design of the software systems needed to leverage flight-contextual data. Domain adaptation of the WDA system proceeds via three main activities: Domain data ingestion, lexical customization and model training. A textual corpus consisting of 1,147 individual documents with more than 7.5 million words of text was ingested into the system and this served as the basis of all further development. A domain lexicon of over 3,500 aviation-domain terms was semi-automatically generated from domain documents and used to train the system. In addition, a set of over 500 question-answer (QA) pairs relevant to the PEER use case was developed; these were used to train and assess the system. These important first steps established the basis for the PEER system. In addition, steps were taken towards the integration of the PEER system into the cockpit environment with the development of a functional design for the Contextual Data Augmentation (CDA) subsystem. This subsystem brings to bear contextual data to improve system responses. It has three main submodules: the Contextual Data Collection module, the Contextual Data Selection module, and the Contextual QA Augmentation module. These modules form a processing pipeline that addresses the problems associated with automatically integrating information from external resources into the knowledge-retrieval mechanism.

Machine learning↗

Polymer informatics: Current status and critical next steps

Artificial intelligence (AI) based approaches are beginning to impact several domains of human life, science and technology. Polymer informatics is one such domain where AI and machine learning (ML) tools are being used in the efficient development, design and discovery of polymers. Surrogate models are trained on available polymer data for instant property prediction, allowing screening of promising polymer candidates with specific target property requirements. Questions regarding synthesizability, and potential (retro)synthesis steps to create a target polymer, are being explored using statistical means. Data-driven strategies to tackle unique challenges resulting from the extraordinary chemical and physical diversity of polymers at small and large scales are being explored. Other major hurdles for polymer informatics are the lack of widespread availability of curated and organized data, and approaches to create machine-readable representations that capture not just the structure of complex polymeric situations but also synthesis and processing conditions. Methods to solve inverse problems, wherein polymer recommendations are made using advanced AI algorithms that meet application targets, are being investigated. As various parts of the burgeoning polymer informatics ecosystem mature and become integrated, efficiency improvements, accelerated discoveries and increased productivity can result. Here in this paper, we review emergent components of this polymer informatics ecosystem and discuss imminent challenges and opportunities.

36 MATERIALS SCIENCE↗

A Universal Machine Learning Model for Elemental Grain Boundary Energies

The grain boundary (GB) energy has a profound influence on the grain growth and properties of polycrystalline metals. Here, we show that the energy of a GB, normalized by the bulk cohesive energy, can be described purely by four geometric features. By machine learning on a large computed database of 361 small Σ (Σ<10) GBs of more than 50 metals, we develop a model that can predict the grain boundary energies to within a mean absolute error of 0.13 J m –2 . More importantly, this universal GB energy model can be extrapolated to the energies of high Σ GBs without loss in accuracy. These results highlight the importance of capturing fundamental scaling physics and domain knowledge in the design of interpretable, extrapolatable machine learning models for materials science.

36 MATERIALS SCIENCE↗

Scientific machine learning benchmarks

Deep learning has transformed the use of machine learning technologies for the analysis of large experimental datasets. In science, such datasets are typically generated by large-scale experimental facilities, and machine learning focuses on the identification of patterns, trends and anomalies to extract meaningful scientific insights from the data. In upcoming experimental facilities, such as the Extreme Photonics Application Centre (EPAC) in the UK or the international Square Kilometre Array (SKA), the rate of data generation and the scale of data volumes will increasingly require the use of more automated data analysis. Furthermore, at present, identifying the most appropriate machine learning algorithm for the analysis of any given scientific dataset is a challenge due to the potential applicability of many different machine learning frameworks, computer architectures and machine learning models. Historically, for modelling and simulation on high-performance computing systems, these issues have been addressed through benchmarking computer applications, algorithms and architectures. Extending such a benchmarking approach and identifying metrics for the application of machine learning methods to open, curated scientific datasets is a new challenge for both scientists and computer scientists. Here, we introduce the concept of machine learning benchmarks for science and review existing approaches. As an example, we describe the SciMLBench suite of scientific machine learning benchmarks.

42 ENGINEERING↗

A workflow for segmenting soil and plant X-ray computed tomography images with deep learning in Google’s Colaboratory

X-ray micro-computed tomography (X-ray μCT) has enabled the characterization of the properties and processes that take place in plants and soils at the micron scale. Despite the widespread use of this advanced technique, major limitations in both hardware and software limit the speed and accuracy of image processing and data analysis. Recent advances in machine learning, specifically the application of convolutional neural networks to image analysis, have enabled rapid and accurate segmentation of image data. Yet, challenges remain in applying convolutional neural networks to the analysis of environmentally and agriculturally relevant images. Specifically, there is a disconnect between the computer scientists and engineers, who build these AI/ML tools, and the potential end users in agricultural research, who may be unsure of how to apply these tools in their work. Additionally, the computing resources required for training and applying deep learning models are unique, more common to computer gaming systems or graphics design work, than to traditional computational systems. To navigate these challenges, we developed a modular workflow for applying convolutional neural networks to X-ray μCT images, using low-cost resources in Google’s Colaboratory web application. Here we present the results of the workflow, illustrating how parameters can be optimized to achieve best results using example scans from walnut leaves, almond flower buds, and a soil aggregate. We expect that this framework will accelerate the adoption and use of emerging deep learning techniques within the plant and soil sciences.

59 BASIC BIOLOGICAL SCIENCES↗

Physics Discovery in Nanoplasmonic Systems via Autonomous Experiments in Scanning Transmission Electron Microscopy

Abstract Physics‐driven discovery in an autonomous experiment has emerged as a dream application of machine learning in physical sciences. Here, this work develops and experimentally implements a deep kernel learning (DKL) workflow combining the correlative prediction of the target functional response and its uncertainty from the structure, and physics‐based selection of acquisition function, which autonomously guides the navigation of the image space. Compared to classical Bayesian optimization (BO) methods, this approach allows to capture the complex spatial features present in the images of realistic materials, and dynamically learn structure–property relationships. In combination with the flexible scalarizer function that allows to ascribe the degree of physical interest to predicted spectra, this enables physical discovery in automated experiment. Here, this approach is illustrated for nanoplasmonic studies of nanoparticles and experimentally implemented in a truly autonomous fashion for bulk‐ and edge plasmon discovery in MnPS 3 , a lesser‐known beam‐sensitive layered 2D material. This approach is universal, can be directly used as‐is with any specimen, and is expected to be applicable to any probe‐based microscopic techniques including other STEM modalities, scanning probe microscopies, chemical, and optical imaging.

42 ENGINEERING↗

Neural network-based control of an ultrafast laser

With the recent advances in machine learning (ML) and data science (DS), the control, modeling, and analysis of these complex systems continues to improve. In this work, we report on the optimization of the intensity of a femtosecond laser using feedforward neural networks (FFNN) that model the input–output relationships of the data. The input parameters of the system were optimized to achieve the required performance of the femtosecond laser. We propose a neural network-based control system to model the relationship between the spectral amplitude and phase of the input laser pulse at the amplifier input and the shape of the output pulse. Low-jitter laser parameter inputs and the resulting laser pulse duration were modeled, and the resulting correlation between the input and output data was used to optimize the laser pulse. Here, we demonstrate improved processing and laser control performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SOMAS: a platform for data-driven material discovery in redox flow battery development

Abstract Aqueous organic redox flow batteries offer an environmentally benign, tunable, and safe route to large-scale energy storage. The energy density is one of the key performance parameters of organic redox flow batteries, which critically depends on the solubility of the redox-active molecule in water. Prediction of aqueous solubility remains a challenge in chemistry. Recently, machine learning models have been developed for molecular properties prediction in chemistry and material science. The fidelity of a machine learning model critically depends on the diversity, accuracy, and abundancy of the training datasets. We build a comprehensive open access organic molecular database “Solubility of Organic Molecules in Aqueous Solution” (SOMAS) containing about 12,000 molecules that covers wider chemical and solubility regimes suitable for aqueous organic redox flow battery development efforts. In addition to experimental solubility, we also provide eight distinctive quantum descriptors including optimized geometry derived from high-throughput density functional theory calculations along with six molecular descriptors for each molecule. SOMAS builds a critical foundation for future efforts in artificial intelligence-based solubility prediction models.

25 ENERGY STORAGE↗

Accelerating discrete dislocation dynamics simulations with graph neural networks

Discrete dislocation dynamics (DDD) is a widely employed computational method to study plasticity at the mesoscale that connects the motion of dislocation lines to the macroscopic response of crystalline materials. However, the computational cost of DDD simulations remains a bottleneck that limits its range of applicability. Here, we introduce a new DDD-GNN framework in which the expensive time-integration of dislocation motion is entirely substituted by a graph neural network (GNN) model trained on DDD trajectories. As a first application, we demonstrate the feasibility and potential of our method on a simple yet relevant model of a dislocation line gliding through an array of obstacles. We show that the DDD-GNN model is stable and reproduces very well unseen ground-truth DDD simulation responses for a range of straining rates and obstacle densities, without the need to explicitly compute nodal forces or dislocation mobilities during time-integration. Our approach opens new promising avenues to accelerate DDD simulations and to incorporate more complex dislocation motion behaviors.

36 MATERIALS SCIENCE↗

Ambient-temperature liquid jet targets for high-repetition-rate HED discovery science

High-power lasers can generate energetic particle beams and astrophysically relevant pressure and temperature states in the high-energy-density (HED) regime. Recently-commissioned high-repetition-rate (HRR) laser drivers are capable of producing these conditions at rates exceeding 1 Hz. However, experimental output from these systems is often limited by the difficulty of designing targets that match these repetition rates. To overcome this challenge, we have developed tungsten microfluidic nozzles, which produce a continuously replenishing jet that operates at flow speeds of approximately 10 m/s and can sustain shot frequencies up to 1 kHz. The ambient-temperature planar liquid jets produced by these nozzles can have thicknesses ranging from hundreds of nanometers to tens of micrometers. In this work, we illustrate the operational principle of the microfluidic nozzle and describe its implementation in a vacuum environment. Further, we provide evidence of successful laser-driven ion acceleration using this target and discuss the prospect of optimizing the ion acceleration performance through an in situ jet thickness scan. Future applications for the jet throughout HED science include shock compression and studies of strongly heated nonequilibrium plasmas. When fielded in concert with HRR-compatible laser, diagnostic, and active feedback technology, this target will facilitate advanced automated studies in HRR HED science, including machine learning-based optimization and high-dimensional statistical analysis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Application of machine learning for the estimation of electron energy distribution from optical emission spectra

Abstract This paper discusses the use of probabilistic deep neural networks for the prediction of the electron energy probability function in low-temperature non-thermal plasmas. The neural networks are trained using optical emission spectroscopy and Langmuir probe measurements, with the goal of providing a reliable estimate of the electron energy probability function solely from optical emission data. The performance of both non-Bayesian and Bayesian networks is evaluated. It is found that Bayesian models are preferable as they assign a higher level of uncertainty to their prediction especially when the dataset used to train them is small. This work describes one of the many potential applications of machine learning in plasma science and technology.

Physics↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Database schema design: Energy Flexibility Environmental Tradeoffs Tool

The Energy Flexibility-Environment Tradeoff Toolset is designed using Streamlit framework, with Python as the programming language. Data storage is facilitated through the use of SQLite. Streamlit is an open-source Python framework for machine learning and data science teams, and SQLite is the most used database engine.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗

The mGA1.0: A common LISP implementation of a messy genetic algorithm

Genetic algorithms (GAs) are finding increased application in difficult search, optimization, and machine learning problems in science and engineering. Increasing demands are being placed on algorithm performance, and the remaining challenges of genetic algorithm theory and practice are becoming increasingly unavoidable. Perhaps the most difficult of these challenges is the so-called linkage problem. Messy GAs were created to overcome the linkage problem of simple genetic algorithms by combining variable-length strings, gene expression, messy operators, and a nonhomogeneous phasing of evolutionary processing. Results on a number of difficult deceptive test functions are encouraging with the mGA always finding global optima in a polynomial number of function evaluations. Theoretical and empirical studies are continuing, and a first version of a messy GA is ready for testing by others. A Common LISP implementation called mGA1.0 is documented and related to the basic principles and operators developed by Goldberg et. al. (1989, 1990). Although the code was prepared with care, it is not a general-purpose code, only a research version. Important data structures and global variations are described. Thereafter brief function descriptions are given, and sample input data are presented together with sample program output. A source listing with comments is also included.

Goldberg, David E.↗

Neural Network Machine Learning and Dimension Reduction for Data Visualization

Neural network machine learning in computer science is a continuously developing field of study. Although neural network models have been developed which can accurately predict a numeric value or nominal classification, a general purpose method for constructing neural network architecture has yet to be developed. Computer scientists are often forced to rely on a trial-and-error process of developing and improving accurate neural network models. In many cases, models are constructed from a large number of input parameters. Understanding which input parameters have the greatest impact on the prediction of the model is often difficult to surmise, especially when the number of input variables is very high. This challenge is often labeled the "curse of dimensionality" in scientific fields. However, techniques exist for reducing the dimensionality of problems to just two dimensions. Once a problem's dimensions have been mapped to two dimensions, it can be easily plotted and understood by humans. The ability to visualize a multi-dimensional dataset can provide a means of identifying which input variables have the highest effect on determining a nominal or numeric output. Identifying these variables can provide a better means of training neural network models; models can be more easily and quickly trained using only input variables which appear to affect the outcome variable. The purpose of this project is to explore varying means of training neural networks and to utilize dimensional reduction for visualizing and understanding complex datasets.

Liles, Charles A.↗

Research and Technology Challenges for Human Data Analysts in Future Safety Management Systems

Enabling new and novel concepts of operations for Advanced Air Mobility poses an important need to evolve current safety management systems (SMS) and is posited to be realized through advances in Machine Learning (ML) Data Sciences and Artificial Intelligence. The “In-time Aviation Safety Management System” (IASMS) concept of operations supports the need to evolve today’s SMS to become more tailorable, scalable, and interoperable in response to forecasted changes expected for the future airspace system. Key to IASMS is integration of proactive and predictive ML algorithms trained to provide “in time” detection and mitigation of hazards and emergent risks through new methods and novel data types. IASMS research and technology development includes human factors design considerations for these systems to include human-system teaming, innovations in human interfaces and management of complex digital data information, human-system interaction/model-based system engineering, and verification and validation for data assurance and trust.

Chad L Stephens↗

ICME for NASA Aerospace Applications: Batteries for Electric Aviation

NASA’s approach to computational materials modeling is detailed in the NASA Vision 2040 Roadmap for Multiscale Modeling and Simulation of Materials and Systems. This report is in the spirit of national initiatives such as the Material Genome Initiative (MGI), Integrated Computational Materials Engineering (ICME), and others. We utilize a combination of fundamental modeling, computational high-throughput screening, and data science methods, e.g., machine learning, are used to find innovative solutions to NASA or national technology challenges. Applications of interest are wide ranging from advanced alloys to batteries to coatings, among others. In this talk, we present three examples for recent work related to NASA applications. First, doping advanced sulfur battery cathodes with selenium boosts electrical conductivity important for electric aircraft applications. First principles calculations will be discussed that result in compositional design maps for these materials. Second, development of icephobic coatings is important to mitigate safety hazards associated with icing for aircraft. Molecular dynamics simulations are reported for ice-surface interfaces to understand adhesion mechanisms and help screen optimal ice-phobic coatings. Third, shape memory alloys have numerous applications as actuators, superelastic materials, etc. for aerospace. We report machine learning models that predict martensitic transition temperatures across a broad swath of compositional space.

John Lawson↗