Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Single Bimodular Sensor for Differentiated Detection of Multiple Oxidative Gases

Semiconductive metal-oxide sensors suffer from cross-sensitivities under mixed chemical condition, specifically upon mixture of multiple oxidative or reductive gases. Herein, a single bimodular sensor is demonstrated for smart differentiation of multiple oxidative analytes by relating the resistance-metric mode to impedance-metric mode. The sensor construct based on ZnO nanorods readily outputs three response datasets upon exposure of oxidative-gas mixture including O 2 , SO 2 , and NO 2 , the resistance, real part impedance, and imaginary part impedance. The differentiative and correlated nature between these response signals allows such a single sensor platform to differentiate these oxidative gases accurately and robustly. Linear and non-linear decision boundaries are established over a large gas-concentration range from 2 ppm to 3% through a combination of principal component analysis and artificial neural network training. A facile user interface is demonstrated for recognition and measurement of unknown gas analytes, with the error of the predicted analyte-concentration as low as 2%.

36 MATERIALS SCIENCE↗

Unraveling the size fluctuation and shrinkage of nanovoids during in situ radiation of Cu by automatic pattern recognition and phase field simulation

Void formation is an important aspect of irradiation response of metals. In situ transmission electron microscopy observation for void evolution during irradiation is an effective technique for studying void evolution. However, the amount of data collected during in situ studies drastically overwhelm the current capability for manual data analyses. Here, we used a data-driven approach where a convolutional neural network combined with greedy matching to detect and track nanovoid evolutions and migrations. This approach was able to discover the surprising phenomena of void size fluctuation and shrinkage during irradiation of Cu with pre-existing nanovoids. Phase–field simulations revealed the fundamental mechanism behind this in situ observed phenomenon of void size fluctuation.

36 MATERIALS SCIENCE↗

Brain Connectivity Workshop Series Report

The National Institutes of Health (NIH) BRAIN Initiative aims to accelerate development and application of technologies that show how individual brain cells in complex neural circuits interact at the speed of thought and action. Complete wiring diagrams of the mammalian brain will revolutionize the capabilities of researchers to formulate and test models of how activity within brain circuits drives coordinated function and behavior. A top priority for the BRAIN Initiative, motivated by the strategic guidance of the BRAIN 2.0 Working Group report, is to muster the resources and foster the collaborations needed to generate these wiring diagrams at the level of long-range projections (“projectomes”) and synapses (“connectomes”) in whole mammalian brains, including those from rodents, humans, and other large-brained mammals. Achieving the scale and resolution necessary to characterize and use these projectomes and connectomes will not only push the boundaries of current imaging and computational capabilities to generate faster, cheaper, and more scalable technologies, but also drive broad innovation in data science, artificial intelligence, and machine learning.

42 ENGINEERING↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Frontiers in the Simulation of Dislocations

Dislocations play a vital role in the mechanical behavior of crystalline materials during deformation. To capture dislocation phenomena across all relevant scales, a multiscale modeling framework of plasticity has emerged, with the goal of reaching a quantitative understanding of microstructure–property relations, for instance, to predict the strength and toughness of metals and alloys for engineering applications. This review describes the state of the art of the major dislocation modeling techniques, and then discusses how recent progress can be leveraged to advance the frontiers in simulations of dislocations. Furthermore, the frontiers of dislocation modeling include opportunities to establish quantitative connections between the scales, validate models against experiments, and use data science methods (e.g., machine learning) to gain an understanding of and enhance the current predictive capabilities.

36 MATERIALS SCIENCE↗

Machine-Learning for Excited-State Dynamics

The primary objective of this computational chemistry sciences team is to design a machine learning NAMD environment that will utilize current petascale and future exascale computational capabilities to advance understanding of charge and energy flow in materials. Our machine-learning NAMD environment will 1) integrate advanced NAMD capabilities directly into electronic structure software (e.g., ABINIT, Quantum Espresso, VASP, etc.); 2) merge the preparatory tools of Pychemia into PYXAID and Avogadro environments so that massive data collection from NAMD simulations.

36 MATERIALS SCIENCE↗

Toward autonomous laboratories: Convergence of artificial intelligence and experimental automation

The ever-increasing demand for novel materials with superior properties inspires retrofitting traditional research paradigms in the era of artificial intelligence and automation. An autonomous experimental platform (AEP) has emerged as an exciting research frontier that achieves full autonomy via integrating data-driven algorithms such as machine learning (ML) with experimental automation in the material development loop from synthesis, characterization, and analysis, to decision making. In this review, we started with a primer to describe how to develop data-driven algorithms for solving material problems. Then, we systematically summarized recent progress on automated material synthesis, ML-enabled data analysis, and decision-making. Finally, we discussed the challenges and opportunities in an endeavor to develop the next-generation AEP for ultimately realizing an autonomous or self-driving laboratory. In conclusion, this review will provide insights for researchers aiming to learn the frontier of ML in materials science and deploy AEP in their labs for accelerating material development.

36 MATERIALS SCIENCE↗

A new paradigm in electron microscopy: Automated microstructure analysis utilizing a dynamic segmentation convolutional neutral network

Over the past half century, the transmission electron microscope enabled insight into the fundamental arrangements and structures of materials. State-of-the-art electron microscopes can acquire large image datasets across multiple imaging modalities. However, the manual annotation process for feature or defect quantification may not be feasible with the modern microscope. Convolutional neural networks emerged to characterize individual microstructural features from an image in a cost-effective, consistent manner. However, many of these neural network approaches rely on thousands to hundreds of thousands of manual annotations of each feature type across hundreds of images to train the network for adequate performance. This work focused on the development and application of a pixel-wise defect detection machine-learning dynamic segmentation convolutional neural network with associated automated acquisition and postprocessing to identify microstructural features rapidly and quantitatively from a small initial dataset incorporating multiple imaging modes. The approach was demonstrated for characterization of superalloy 718 from both single image acquisition on multiple detectors to in-situ evolution captured with a single detector on a standard desktop computer to demonstrate the low barrier to entry required for widespread adoption. Pixel-by-pixel class identification was excellent with strong identification of chemically distinct phases, structurally distinct phases, and defect structures, thus demonstrating the new paradigm of machine learning-assisted characterization.

36 MATERIALS SCIENCE↗

GridDS: Data Science Toolkit for Energy Grid Data

According to the U.S. Energy Information Administration (EIA), the demand for energy is expected to increase 50% by the year 20501. While energy standards, such as the Institute of Electrical and Electronics Engineers (IEEE) Standard 1547, (Basso 2015) and monitoring with wide area management systems (WAMS) (Liu 2017, Zhou 2016) have enabled large scale data collection and storage, the application of this data in mitigating costs associated with increased consumer demand is an ongoing focus for energy research. This ubiquitous data collection presents a promising opportunity for machine learning and data science to improve efficiency of distributed energy resources (DERs). The GridDS software toolkit is designed to leverage advanced metering infrastructure (AMI), outage management systems data (OMS), Supervisory control Data Acquisition (SCADA), and geographic information systems (GIS) to forecast future energy demands and detect incipient grid failures. GridDS is a python software library designed to be modular and generalizable to data recorded by DERs. In adapting to disparate datasets recorded by various WAMS, GridDS provides a range of unique functionality not presently implemented in current WAMS which have highly specific software infrastructure by design. GridDS functionality ranges from data specification and preparation, to training and validation for state of the art machine learning, to interactive data visualization. For data intake, GridDS combines: Pandera: a library for creating data specifications. TimeScaleDB: a postgresSQL database infrastructure for efficient storage of timeseries data. Dataset class: A custom dataset class / interface that ensures modularity between a range of synthetic and live recorded datasets. Is

Ladd, Alexander↗

Subtleties in the trainability of quantum machine learning models

A new paradigm for data science has emerged, with quantum data, quantum models, and quantum computational devices. This field, called quantum machine learning (QML), aims to achieve a speedup over traditional machine learning for data analysis. However, its success usually hinges on efficiently training the parameters in quantum neural networks, and the field of QML is still lacking theoretical scaling results for their trainability. Some trainability results have been proven for a closely related field called variational quantum algorithms (VQAs). While both fields involve training a parametrized quantum circuit, there are crucial differences that make the results for one setting not readily applicable to the other. In this work, we bridge the two frameworks and show that gradient scaling results for VQAs can also be applied to study the gradient scaling of QML models. Our results indicate that features deemed detrimental for VQA trainability can also lead to issues such as barren plateaus in QML. Consequently, our work has implications for several QML proposals in the literature. In addition, we provide theoretical and numerical evidence that QML models exhibit further trainability issues not present in VQAs, arising from the use of a training dataset. We refer to these as dataset-induced barren plateaus. These results are most relevant when dealing with classical data, as here the choice of embedding scheme (i.e., the map between classical data and quantum states) can greatly affect the gradient scaling.

97 MATHEMATICS AND COMPUTING↗

Increasing the Reproducibility and Replicability of Supervised AI/ML in the Earth Systems Science by Leveraging Social Science Methods

Artificial intelligence (AI) and machine learning (ML) pose a challenge for achieving science that is both reproducible and replicable. The challenge is compounded in supervised models that depend on manually labeled training data, as they introduce additional decision-making and processes that require thorough documentation and reporting. We address these limitations by providing an approach to hand labeling training data for supervised ML that integrates quantitative content analysis (QCA)—a method from social science research. The QCA approach provides a rigorous and well-documented hand labeling procedure to improve the replicability and reproducibility of supervised ML applications in Earth systems science (ESS), as well as the ability to evaluate them. Specifically, the approach requires (a) the articulation and documentation of the exact decision-making process used for assigning hand labels in a “codebook” and (b) an empirical evaluation of the reliability” of the hand labelers. In this paper, we outline the contributions of QCA to the field, along with an overview of the general approach. We then provide a case study to further demonstrate how this framework has and can be applied when developing supervised ML models for applications in ESS. With this approach, we provide an actionable path forward for addressing ethical considerations and goals outlined by recent AGU work on ML ethics in ESS.

58 GEOSCIENCES↗

FY21 Proxy App Suite Release: Report for ECP Proxy App Project Milestone ADCD504-12

The FY21 Proxy App Suite Release milestone includes the following activities: Curate a collection of proxy applications that represents the breadth of ECP applications, including application domains, programming models, supporting libraries, numerical methods, etc. Identify gaps in coverage and work with application teams to commission or develop proxies to cover gaps. From within this collection, designate the "ECP Proxy Application Suite" of 12-15 proxies that balance breadth of coverage with ease of use and quality of implementation. Also designate approximately 8-10 proxies to form the \ECP Machine Learning Proxy Suite". The ML suite will represent algorithms, use cases, and programming methods typically used by ECP science workloads to incorporate machine learning into their workflows.

97 MATHEMATICS AND COMPUTING↗

FY22 Proxy App Suite Release

The FY22 Proxy App Suite Release milestone includes the following activities: Curate a collection of proxy applications that represents the breadth of ECP applications, including application domains, programming models, supporting libraries, numerical methods, etc. Identify gaps in coverage and work with application teams to commission or develop proxies to cover gaps. From within this collection, designate the ”ECP Proxy Application Suite” of 10–15 proxies that balance breadth of coverage with ease of use and quality of implementation. Also designate approximately 6–10 proxies to form the “ECP Machine Learning Proxy Suite”. The ML suite will represent algorithms, use cases, and programming methods typically used by ECP science workloads to incorporate machine learning into their workflows.

97 MATHEMATICS AND COMPUTING↗

Laboratory earthquake forecasting: A machine learning competition

Earthquake prediction, the long-sought holy grail of earthquake science, continues to confound Earth scientists. Could we make advances by crowdsourcing, drawing from the vast knowledge and creativity of the machine learning (ML) community? We used Google’s ML competition platform, Kaggle, to engage the worldwide ML community with a competition to develop and improve data analysis approaches on a forecasting problem that uses laboratory earthquake data. The competitors were tasked with predicting the time remaining before the next earthquake of successive laboratory quake events, based on only a small portion of the laboratory seismic data. The more than 4,500 participating teams created and shared more than 400 computer programs in openly accessible notebooks. Complementing the now well-known features of seismic data that map to fault criticality in the laboratory, the winning teams employed unexpected strategies based on rescaling failure times as a fraction of the seismic cycle and comparing input distribution of training and testing data. In addition to yielding scientific insights into fault processes in the laboratory and their relation with the evolution of the statistical properties of the associated seismic data, the competition serves as a pedagogical tool for teaching ML in geophysics. The approach may provide a model for other competitions in geosciences or other domains of study to help engage the ML community on problems of significance.

58 GEOSCIENCES↗

NASA Pilot-Engaged Expert Response Using IBM Watson Technology: Prototype Evaluation of Knowledge Retrieval System

NASA Langley Research Center and IBM have been investigating the use of IBM Watson technology in aerospace research and development. One application of Watson technology is the Pilot-Engaged Expert Response (PEER) use case. The PEER system is envisioned as an in-cockpit advisor that will act as a source of situationally-relevant information for pilots and other flight crew members to assist in decision making about real-time events and situations that arise in the course of aircraft operations. PEER will make available vast stores of knowledge and information quickly and directly, putting important informational resources where they are needed most. IBM has worked with NASA to develop an architecture and articulate a roadmap for the development of the PEER system. That vision is built around Watson Discovery Advisor (WDA) software solution, derived from IBM's Jeopardy!-winning automatic question answering system. PEER makes use of WDA's sophisticated question-answering capabilities as its core, adding important User Interface components and other customizations for the cockpit environment, including communication with flight systems and other external data sources. The development plan for PEER includes four development stages, with the current project constituting the first phase. In this project, a prototype instance of PEER was successfully adapted to the aviation domain, enabling users to ask questions about aviation topics and receive useful and accurate answers to these questions. Major tasks accomplished include the development of procedures for domain adaptation through automatic lexicon extraction from domain glossaries; generation of question-answer training data which was used to train the system; and assessment of the effectiveness of domain adaptation, which showed a dramatic improvement in the ability of the PEER system to answer domain-relevant questions. In addition, the vision for the PEER system was pushed forward by the articulation of a plan for the automatic enhancement of question-answering with contextual information. This initial phase focused on two main goals: 1) the targeted domain adaptation of the underlying WDA system to the aviation domain; and, 2) the design of the software systems needed to leverage flight-contextual data. Domain adaptation of the WDA system proceeds via three main activities: Domain data ingestion, lexical customization and model training. A textual corpus consisting of 1,147 individual documents with more than 7.5 million words of text was ingested into the system and this served as the basis of all further development. A domain lexicon of over 3,500 aviation-domain terms was semi-automatically generated from domain documents and used to train the system. In addition, a set of over 500 question-answer (QA) pairs relevant to the PEER use case was developed; these were used to train and assess the system. These important first steps established the basis for the PEER system. In addition, steps were taken towards the integration of the PEER system into the cockpit environment with the development of a functional design for the Contextual Data Augmentation (CDA) subsystem. This subsystem brings to bear contextual data to improve system responses. It has three main submodules: the Contextual Data Collection module, the Contextual Data Selection module, and the Contextual QA Augmentation module. These modules form a processing pipeline that addresses the problems associated with automatically integrating information from external resources into the knowledge-retrieval mechanism.

Machine learning↗

A Universal Machine Learning Model for Elemental Grain Boundary Energies

The grain boundary (GB) energy has a profound influence on the grain growth and properties of polycrystalline metals. Here, we show that the energy of a GB, normalized by the bulk cohesive energy, can be described purely by four geometric features. By machine learning on a large computed database of 361 small Σ (Σ<10) GBs of more than 50 metals, we develop a model that can predict the grain boundary energies to within a mean absolute error of 0.13 J m –2 . More importantly, this universal GB energy model can be extrapolated to the energies of high Σ GBs without loss in accuracy. These results highlight the importance of capturing fundamental scaling physics and domain knowledge in the design of interpretable, extrapolatable machine learning models for materials science.

36 MATERIALS SCIENCE↗

Scientific machine learning benchmarks

Deep learning has transformed the use of machine learning technologies for the analysis of large experimental datasets. In science, such datasets are typically generated by large-scale experimental facilities, and machine learning focuses on the identification of patterns, trends and anomalies to extract meaningful scientific insights from the data. In upcoming experimental facilities, such as the Extreme Photonics Application Centre (EPAC) in the UK or the international Square Kilometre Array (SKA), the rate of data generation and the scale of data volumes will increasingly require the use of more automated data analysis. Furthermore, at present, identifying the most appropriate machine learning algorithm for the analysis of any given scientific dataset is a challenge due to the potential applicability of many different machine learning frameworks, computer architectures and machine learning models. Historically, for modelling and simulation on high-performance computing systems, these issues have been addressed through benchmarking computer applications, algorithms and architectures. Extending such a benchmarking approach and identifying metrics for the application of machine learning methods to open, curated scientific datasets is a new challenge for both scientists and computer scientists. Here, we introduce the concept of machine learning benchmarks for science and review existing approaches. As an example, we describe the SciMLBench suite of scientific machine learning benchmarks.

42 ENGINEERING↗