Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Accelerating manufacturing for biomass conversion via integrated process and bench digitalization: a perspective

We present a perspective for accelerating biomass manufacturing via digitalization. We summarize the challenges for manufacturing and identify areas where digitalization can help. A profound potential in using lignocellulosic biomass and renewable feedstocks, in general, is to produce new molecules and products with unmatched properties that have no analog in traditional refineries. Discovering such performance-advantaged molecules and the paths and processes to make them rapidly and systematically can transform manufacturing practices. Furthermore, we discuss retrosynthetic approaches, text mining, natural language processing, and modern machine learning methods to enable digitalization. Laboratory and multiscale computation automation via active learning are crucial to complement existing literature and expedite discovery and valuable data collection without a human in the loop. Such data can help process simulation and optimization select the most promising processes and molecules according to economic, environmental, and societal metrics. We propose the close integration between bench and process scale models and data to exploit the low dimensionality of the data and transform the manufacturing for renewable feedstocks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven Strategies for Accelerated Materials Design

The ongoing revolution of the natural sciences by the advent of machine learning and artificial intelligence sparked significant interest in the material science community in recent years. The intrinsically high dimensionality of the space of realizable materials makes traditional approaches ineffective for large-scale explorations. Modern data science and machine learning tools developed for increasingly complicated problems are an attractive alternative. An imminent climate catastrophe calls for a clean energy transformation by overhauling current technologies within only several years of possible action available. Tackling this crisis requires the development of new materials at an unprecedented pace and scale. For example, organic photovoltaics have the potential to replace existing silicon-based materials to a large extent and open up new fields of application. In recent years, organic light-emitting diodes have emerged as state-of-the-art technology for digital screens and portable devices and are enabling new applications with flexible displays. Reticular frameworks allow the atom-precise synthesis of nanomaterials and promise to revolutionize the field by the potential to realize multifunctional nanoparticles with applications from gas storage, gas separation, and electrochemical energy storage to nanomedicine. In the recent decade, significant advances in all these fields have been facilitated by the comprehensive application of simulation and machine learning for property prediction, property optimization, and chemical space exploration enabled by considerable advances in computing power and algorithmic efficiency. In this Account, we review the most recent contributions of our group in this thriving field of machine learning for material science. We start with a summary of the most important material classes our group has been involved in, focusing on small molecules as organic electronic materials and crystalline materials. Specifically, we highlight the data-driven approaches we employed to speed up discovery and derive material design strategies. Subsequently, our focus lies on the data-driven methodologies our group has developed and employed, elaborating on high-throughput virtual screening, inverse molecular design, Bayesian optimization, and supervised learning. We discuss the general ideas, their working principles, and their use cases with examples of successful implementations in data-driven material discovery and design efforts. Furthermore, we elaborate on potential pitfalls and remaining challenges of these methods. Finally, we provide a brief outlook for the field as we foresee increasing adaptation and implementation of large scale data-driven approaches in material discovery and design campaigns.

36 MATERIALS SCIENCE↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Breeding Realistic D‐Brane Models

Abstract Intersecting branes provide a useful mechanism to construct particle physics models from string theory with a wide variety of desirable characteristics. The landscape of such models can be enormous, and navigating towards regions which are most phenomenologically interesting is potentially challenging. Machine learning techniques can be used to efficiently construct large numbers of consistent and phenomenologically desirable models. In this work we phrase the problem of finding consistent intersecting D‐brane models in terms of genetic algorithms, which mimic natural selection to evolve a population collectively towards optimal solutions. For a four‐dimensional supersymmetric type IIA orientifold with intersecting D6‐branes, we demonstrate that unique, fully consistent models can be easily constructed, and, by a judicious choice of search environment and hyper‐parameters, of the found models contain the desired Standard Model gauge group factor. Having a sizable sample allows us to draw some preliminary landscape statistics of intersecting brane models both with and without the restriction of having the Standard Model gauge factor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Reinforced double-threaded slide-ring networks for accelerated hydrogel discovery and 3D printing

Traditionally, slide-ring gels are stretchable but soft as a result of an elasticity-stretchability trade-off. Herein, we introduce a new approach to breaking this trade-off and creating reinforced slide-ring networks with mobile crosslinkers. Our approach involves the construction of a polyethylene glycol double-threaded γ-cyclodextrin-based pro-slide-ring crosslinker that serves as a modular component for 3D printing and copolymerization. The resulting crystalline-domain-reinforced slide-ring hydrogels, or CrysDoS-gels, exhibit both high elasticity and high stretchability. The modular synthesis allows for high-throughput synthesis of CrysDoS-gels, generating a large amount of data for structure-property analysis. Here, by employing data science techniques, such as machine learning and linear regression, not only were we able to identify which chemical components influence the mechanical properties of CrysDoS-gels, but this analysis also aided in the discovery of better-performing CrysDoS-gels. Finally, we demonstrate the potential application of the newly discovered CrysDoS-gels as sensing devices by 3D printing them as stress sensors with high sensitivity and a broad detection range.

3D-printing↗

Scalable 3D reconstruction for X-ray single particle imaging with online machine learning

X-ray free-electron lasers offer unique capabilities for measuring the structure and dynamics of biomolecules, helping us understand the basic building blocks of life. Notably, high-repetition-rate free-electron lasers enable single particle imaging, where individual, weakly scattering biomolecules are imaged under near-physiological conditions with the opportunity to access fleeting states that cannot be captured in cryogenic or crystallized conditions. Existing X-ray single particle reconstruction algorithms, which estimate the particle orientation for each image independently, are slow and memory-intensive when handling the massive datasets generated by emerging free-electron lasers. Here, we introduce X-RAI (X-Ray single particle imaging with Amortized Inference), an online reconstruction framework that estimates the structure of 3D macromolecules from large X-ray single particle datasets. X-RAI consists of a convolutional encoder, which amortizes pose estimation over large datasets, as well as a physics-based decoder, which employs an implicit neural representation to enable high-quality 3D reconstruction in an end-to-end, self-supervised manner. We demonstrate that X-RAI achieves state-of-the-art performance for small-scale datasets in simulation and challenging experimental settings and demonstrate its unprecedented ability to process large datasets containing millions of diffraction images in an online fashion. These abilities signify a paradigm shift in X-ray single particle imaging towards real-time reconstruction.

Computer science↗

DIPS-Plus: The enhanced database of interacting protein structures for interface prediction

Abstract In this work, we expand on a dataset recently introduced for protein interface prediction (PIP), the Database of Interacting Protein Structures (DIPS), to present DIPS-Plus, an enhanced, feature-rich dataset of 42,112 complexes for machine learning of protein interfaces. While the original DIPS dataset contains only the Cartesian coordinates for atoms contained in the protein complex along with their types, DIPS-Plus contains multiple residue-level features including surface proximities, half-sphere amino acid compositions, and new profile hidden Markov model (HMM)-based sequence features for each amino acid, providing researchers a curated feature bank for training protein interface prediction methods. We demonstrate through rigorous benchmarks that training an existing state-of-the-art (SOTA) model for PIP on DIPS-Plus yields new SOTA results, surpassing the performance of some of the latest models trained on residue-level and atom-level encodings of protein complexes to date.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset of tensile properties for sub-sized specimens of nuclear structural materials

Mechanical testing with sub-sized specimens plays an important role in the nuclear industry, facilitating tests in confined experimental spaces with lower irradiation levels and accelerating the qualification of new materials. The reduced size of specimens results in different material behavior at the microscale, mesoscale, and macroscale, in comparison to standard-sized specimens, which is referred to as the “specimen size effect.” Although analytical models have been proposed to correlate the properties of sub-sized specimens to standard-sized specimens, these models lack broad applicability across different materials and testing conditions. The objective of this study is to create the first large public dataset of tensile properties for sub-sized specimens used in nuclear structural materials. We performed an extensive literature review of relevant publications and extracted over 1,000 tensile testing records comprising 55 columns including material type and composition, manufacturing information, irradiation conditions, specimen dimensions, and tensile properties. The dataset can serve as a valuable resource to investigate the specimen size effect and develop computational methods to correlate the tensile properties of sub-sized specimens.

36 MATERIALS SCIENCE↗

UOWDetection

The code provides machine learning, data collection, and computer science tools used to detect undocumented orphaned wells in United States from aerial imagery.

Kim, Anastasiia↗

Integrated parameter and process learning for hydrologic and biogeochemical modules in Earth System Models

Focus area: Primary focal area #2; secondary focal area #3: Learning about parameters and processes of land surface hydrologic and biogeochemical models in Earth System models by integrating machine learning, physics, and big data. Science challenges: How do we maximally leverage big-data observations to improve hydrobiogeochemical process description and parameterization so that such modules more realistically capture hydrologic and vegetation responses and feedbacks under the future climate? For example, how can we leverage physics, limited observations of vegetation and streamflow to better estimate evapotranspiration, and, relatedly, net primary productivity, especially for drought areas? Vegetation plays a critical role in regional and global water cycles; however, existing vegetation models have failed to predict vegetation response to droughts (McDowell & Xu, 2017) , arctic greening (Keenan & Riley, 2018) , and critical transitions between forest and savanna (Hirota et al., 2011) . These studies suggest that when we build process-based models (PBM) parameterized from regional and global plant traits, we tend to poorly describe plant adaptation and local-scale competition processes. The models and their associated parameters assigned for different regions in the world are not capturing essential heterogeneity in vegetation responses at finer spatial scales. Many parameters of the land surface models control hydrology and vegetation dynamics at the same time. The heterogeneity in vegetation response is a function of (i) plant type, (ii) plant size, (iii) competition and succession, (iv) environmental controls, and (v) local variations due to the unique ecological community that are very difficult to describe (e.g., the size of gaps resulting from fire that facilitated the coexistence of pioneering species). In the demographic models, only factors (i) and (iv) were captured, and plant types were generally described only by leaf phenology and climate zones. With current demographic models, we generally consider more traits to define plant types (i) and calibrate these traits to consider factors (ii), (iii) and (iv); however, it is substantially challenging to scale to regional and global simulations due to trait variations across space (Ali et al., 2016). Moreover, it has been noted that hillslope processes, including ridge-to-valley flow and sunny vs. shady slopes are primary organizers of water, energy, and vegetation (Clark et al., 2015; Fan et al., 2019) . Although gradual improvements in the hydrologic model component in earth system models may reduce this error (at a remarkably slow pace), the long-term, gradual impact of hydrology on plant traits are not well captured. Recent work showed that the hydrologic controls exerted by groundwater and lateral flow are primary regulators of rooting depth (Fan et al., 2017) . Such hydrologic controls have seldom been reflected in vegetation model parameterizations.

54 ENVIRONMENTAL SCIENCES↗

Brain Connectivity Workshop Series Report

The National Institutes of Health (NIH) BRAIN Initiative aims to accelerate development and application of technologies that show how individual brain cells in complex neural circuits interact at the speed of thought and action. Complete wiring diagrams of the mammalian brain will revolutionize the capabilities of researchers to formulate and test models of how activity within brain circuits drives coordinated function and behavior. A top priority for the BRAIN Initiative, motivated by the strategic guidance of the BRAIN 2.0 Working Group report, is to muster the resources and foster the collaborations needed to generate these wiring diagrams at the level of long-range projections (“projectomes”) and synapses (“connectomes”) in whole mammalian brains, including those from rodents, humans, and other large-brained mammals. Achieving the scale and resolution necessary to characterize and use these projectomes and connectomes will not only push the boundaries of current imaging and computational capabilities to generate faster, cheaper, and more scalable technologies, but also drive broad innovation in data science, artificial intelligence, and machine learning.

42 ENGINEERING↗

Cyber-Attack Detection and Accommodation for the Energy Delivery System

The goals of this project were to create a software system with a suite of key algorithms for cyber-attack detection and accommodation providing domain layer protection for critical power generation assets. Example assets included gas and steam turbines, heat recovery steam generators, and electrical generators. The aggressive algorithm goals were aimed at reducing the false positive rates in threat detection to <1% using learnings from many evolving disciplines (power turbine and generator physics, power system modeling, modern control theory, system identification, machine learning, deep learning, mathematics and data science). Additional goals for the algorithms involved localizing threats on-the-fly to know in which monitoring node the effects of attacks are present, and then providing accommodation to keep the system running uninterrupted much of the time in the presence of the attack. Accommodation had a performance goal of providing resiliency when up to 50% of monitoring nodes are in an attack state.

cybersecurity, cyber-physical↗

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields, in part due to a culture of open data sharing and reuse. AI/ML methodology is well-suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Inexperienced researchers can produce models that perform poorly outside of the training dataset. Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Casaletto↗

Methods for rapid identification of anomalous layers in laser powder bed fusion

In situ process monitoring is a key requirement for increased industry acceptance of powder bed Additive Manufacturing. As sensing technologies increase in maturity, attention must also be given to effective data exploration techniques. These data are often high-resolution and multi-modal, with each build consisting of thousands of layers. Here, in this paper, the authors propose two methods enabling users to rapidly identify layers of interest within a build. Both methods leverage results from deep learning based segmentations of in situ powder bed images. The first method is an unsupervised “reverse layer search” algorithm while the second method uses supervised machine learning.

36 MATERIALS SCIENCE↗

Opportunities for human factors in machine learning

Introduction The field of machine learning and its subfield of deep learning have grown rapidly in recent years. With the speed of advancement, it is nearly impossible for data scientists to maintain expert knowledge of cutting-edge techniques. This study applies human factors methods to the field of machine learning to address these difficulties. Methods Using semi-structured interviews with data scientists at a National Laboratory, we sought to understand the process used when working with machine learning models, the challenges encountered, and the ways that human factors might contribute to addressing those challenges. Results Results of the interviews were analyzed to create a generalization of the process of working with machine learning models. Issues encountered during each process step are described. Discussion Recommendations and areas for collaboration between data scientists and human factors experts are provided, with the goal of creating better tools, knowledge, and guidance for machine learning scientists.

97 MATHEMATICS AND COMPUTING↗

A Machine Learning Approach to Predict Martensitic Transition Temperatures for Shape Memory Alloys

Shape memory alloys (SMAs) are a unique class of materials with several remarkable properties including shape recovery, superelasticity, etc. Especially important for many NASA applications is the ability to tune the martensitic phase transition temperature by varying the alloy composition. Nickel-titanium (NiTi) based alloys are the most widely studied of this class, with compositions involving ternary, quaternary, or higher additions being considered. Over the past several years, a significant database of SMA properties has been assembled by NASA researchers. Such a database is ideal for data science-based approaches including machine learning. We present results from a developed machine learning model capable of accurately predicting the transition temperature of SMAs across a wide range of compositions. Our model has the added benefit of interpretability and even provides confidence intervals for our predictions. This model will make rapid screening and design of new SMA materials possible. Predictions from the machine learning model can be validated by empirical and/or atomistic scale modeling.

Shreyas Honrao↗

Multiscale Modeling Meets Machine Learning: What Can We Learn?

Machine learning is increasingly recognized as a promising technology in the biological, biomedical, and behavioral sciences. There can be no argument that this technique is incredibly successful in image recognition with immediate applications in diagnostics including electrophysiology, radiology, or pathology, where we have access to massive amounts of annotated data. However, machine learning often performs poorly in prognosis, especially when dealing with sparse data. This is a field where classical physics-based simulation seems to remain irreplaceable. In this review, we identify areas in the biomedical sciences where machine learning and multiscale modeling can mutually benefit from one another: Machine learning can integrate physics-based knowledge in the form of governing equations, boundary conditions, or constraints to manage ill-posted problems and robustly handle sparse and noisy data; multiscale modeling can integrate machine learn- ing to create surrogate models, identify system dynamics and parameters, analyze sensitivities, and quantify uncertainty to bridge the scales and understand the emergence of function. With a view towards applications in the life sciences, we discuss the state of the art of combining machine learning and multiscale modeling, identify applications and opportunities, raise open questions, and address potential challenges and limitations. We anticipate that it will stimulate discussion within the community of computational mechanics and reach out to other disciplines including mathematics, statistics, computer science, artificial intelligence, biomedicine, systems biology, and precision medicine to join forces towards creating robust and efficient models for biological systems.

machine learning, multiscale modeling, physics-bas↗