Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Breeding Realistic D‐Brane Models

Abstract Intersecting branes provide a useful mechanism to construct particle physics models from string theory with a wide variety of desirable characteristics. The landscape of such models can be enormous, and navigating towards regions which are most phenomenologically interesting is potentially challenging. Machine learning techniques can be used to efficiently construct large numbers of consistent and phenomenologically desirable models. In this work we phrase the problem of finding consistent intersecting D‐brane models in terms of genetic algorithms, which mimic natural selection to evolve a population collectively towards optimal solutions. For a four‐dimensional supersymmetric type IIA orientifold with intersecting D6‐branes, we demonstrate that unique, fully consistent models can be easily constructed, and, by a judicious choice of search environment and hyper‐parameters, of the found models contain the desired Standard Model gauge group factor. Having a sizable sample allows us to draw some preliminary landscape statistics of intersecting brane models both with and without the restriction of having the Standard Model gauge factor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Reinforced double-threaded slide-ring networks for accelerated hydrogel discovery and 3D printing

Traditionally, slide-ring gels are stretchable but soft as a result of an elasticity-stretchability trade-off. Herein, we introduce a new approach to breaking this trade-off and creating reinforced slide-ring networks with mobile crosslinkers. Our approach involves the construction of a polyethylene glycol double-threaded γ-cyclodextrin-based pro-slide-ring crosslinker that serves as a modular component for 3D printing and copolymerization. The resulting crystalline-domain-reinforced slide-ring hydrogels, or CrysDoS-gels, exhibit both high elasticity and high stretchability. The modular synthesis allows for high-throughput synthesis of CrysDoS-gels, generating a large amount of data for structure-property analysis. Here, by employing data science techniques, such as machine learning and linear regression, not only were we able to identify which chemical components influence the mechanical properties of CrysDoS-gels, but this analysis also aided in the discovery of better-performing CrysDoS-gels. Finally, we demonstrate the potential application of the newly discovered CrysDoS-gels as sensing devices by 3D printing them as stress sensors with high sensitivity and a broad detection range.

3D-printing↗

Scalable 3D reconstruction for X-ray single particle imaging with online machine learning

X-ray free-electron lasers offer unique capabilities for measuring the structure and dynamics of biomolecules, helping us understand the basic building blocks of life. Notably, high-repetition-rate free-electron lasers enable single particle imaging, where individual, weakly scattering biomolecules are imaged under near-physiological conditions with the opportunity to access fleeting states that cannot be captured in cryogenic or crystallized conditions. Existing X-ray single particle reconstruction algorithms, which estimate the particle orientation for each image independently, are slow and memory-intensive when handling the massive datasets generated by emerging free-electron lasers. Here, we introduce X-RAI (X-Ray single particle imaging with Amortized Inference), an online reconstruction framework that estimates the structure of 3D macromolecules from large X-ray single particle datasets. X-RAI consists of a convolutional encoder, which amortizes pose estimation over large datasets, as well as a physics-based decoder, which employs an implicit neural representation to enable high-quality 3D reconstruction in an end-to-end, self-supervised manner. We demonstrate that X-RAI achieves state-of-the-art performance for small-scale datasets in simulation and challenging experimental settings and demonstrate its unprecedented ability to process large datasets containing millions of diffraction images in an online fashion. These abilities signify a paradigm shift in X-ray single particle imaging towards real-time reconstruction.

Computer science↗

DIPS-Plus: The enhanced database of interacting protein structures for interface prediction

Abstract In this work, we expand on a dataset recently introduced for protein interface prediction (PIP), the Database of Interacting Protein Structures (DIPS), to present DIPS-Plus, an enhanced, feature-rich dataset of 42,112 complexes for machine learning of protein interfaces. While the original DIPS dataset contains only the Cartesian coordinates for atoms contained in the protein complex along with their types, DIPS-Plus contains multiple residue-level features including surface proximities, half-sphere amino acid compositions, and new profile hidden Markov model (HMM)-based sequence features for each amino acid, providing researchers a curated feature bank for training protein interface prediction methods. We demonstrate through rigorous benchmarks that training an existing state-of-the-art (SOTA) model for PIP on DIPS-Plus yields new SOTA results, surpassing the performance of some of the latest models trained on residue-level and atom-level encodings of protein complexes to date.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset of tensile properties for sub-sized specimens of nuclear structural materials

Mechanical testing with sub-sized specimens plays an important role in the nuclear industry, facilitating tests in confined experimental spaces with lower irradiation levels and accelerating the qualification of new materials. The reduced size of specimens results in different material behavior at the microscale, mesoscale, and macroscale, in comparison to standard-sized specimens, which is referred to as the “specimen size effect.” Although analytical models have been proposed to correlate the properties of sub-sized specimens to standard-sized specimens, these models lack broad applicability across different materials and testing conditions. The objective of this study is to create the first large public dataset of tensile properties for sub-sized specimens used in nuclear structural materials. We performed an extensive literature review of relevant publications and extracted over 1,000 tensile testing records comprising 55 columns including material type and composition, manufacturing information, irradiation conditions, specimen dimensions, and tensile properties. The dataset can serve as a valuable resource to investigate the specimen size effect and develop computational methods to correlate the tensile properties of sub-sized specimens.

36 MATERIALS SCIENCE↗

UOWDetection

The code provides machine learning, data collection, and computer science tools used to detect undocumented orphaned wells in United States from aerial imagery.

Kim, Anastasiia↗

Integrated parameter and process learning for hydrologic and biogeochemical modules in Earth System Models

Focus area: Primary focal area #2; secondary focal area #3: Learning about parameters and processes of land surface hydrologic and biogeochemical models in Earth System models by integrating machine learning, physics, and big data. Science challenges: How do we maximally leverage big-data observations to improve hydrobiogeochemical process description and parameterization so that such modules more realistically capture hydrologic and vegetation responses and feedbacks under the future climate? For example, how can we leverage physics, limited observations of vegetation and streamflow to better estimate evapotranspiration, and, relatedly, net primary productivity, especially for drought areas? Vegetation plays a critical role in regional and global water cycles; however, existing vegetation models have failed to predict vegetation response to droughts (McDowell & Xu, 2017) , arctic greening (Keenan & Riley, 2018) , and critical transitions between forest and savanna (Hirota et al., 2011) . These studies suggest that when we build process-based models (PBM) parameterized from regional and global plant traits, we tend to poorly describe plant adaptation and local-scale competition processes. The models and their associated parameters assigned for different regions in the world are not capturing essential heterogeneity in vegetation responses at finer spatial scales. Many parameters of the land surface models control hydrology and vegetation dynamics at the same time. The heterogeneity in vegetation response is a function of (i) plant type, (ii) plant size, (iii) competition and succession, (iv) environmental controls, and (v) local variations due to the unique ecological community that are very difficult to describe (e.g., the size of gaps resulting from fire that facilitated the coexistence of pioneering species). In the demographic models, only factors (i) and (iv) were captured, and plant types were generally described only by leaf phenology and climate zones. With current demographic models, we generally consider more traits to define plant types (i) and calibrate these traits to consider factors (ii), (iii) and (iv); however, it is substantially challenging to scale to regional and global simulations due to trait variations across space (Ali et al., 2016). Moreover, it has been noted that hillslope processes, including ridge-to-valley flow and sunny vs. shady slopes are primary organizers of water, energy, and vegetation (Clark et al., 2015; Fan et al., 2019) . Although gradual improvements in the hydrologic model component in earth system models may reduce this error (at a remarkably slow pace), the long-term, gradual impact of hydrology on plant traits are not well captured. Recent work showed that the hydrologic controls exerted by groundwater and lateral flow are primary regulators of rooting depth (Fan et al., 2017) . Such hydrologic controls have seldom been reflected in vegetation model parameterizations.

54 ENVIRONMENTAL SCIENCES↗

Brain Connectivity Workshop Series Report

The National Institutes of Health (NIH) BRAIN Initiative aims to accelerate development and application of technologies that show how individual brain cells in complex neural circuits interact at the speed of thought and action. Complete wiring diagrams of the mammalian brain will revolutionize the capabilities of researchers to formulate and test models of how activity within brain circuits drives coordinated function and behavior. A top priority for the BRAIN Initiative, motivated by the strategic guidance of the BRAIN 2.0 Working Group report, is to muster the resources and foster the collaborations needed to generate these wiring diagrams at the level of long-range projections (“projectomes”) and synapses (“connectomes”) in whole mammalian brains, including those from rodents, humans, and other large-brained mammals. Achieving the scale and resolution necessary to characterize and use these projectomes and connectomes will not only push the boundaries of current imaging and computational capabilities to generate faster, cheaper, and more scalable technologies, but also drive broad innovation in data science, artificial intelligence, and machine learning.

42 ENGINEERING↗

Cyber-Attack Detection and Accommodation for the Energy Delivery System

The goals of this project were to create a software system with a suite of key algorithms for cyber-attack detection and accommodation providing domain layer protection for critical power generation assets. Example assets included gas and steam turbines, heat recovery steam generators, and electrical generators. The aggressive algorithm goals were aimed at reducing the false positive rates in threat detection to <1% using learnings from many evolving disciplines (power turbine and generator physics, power system modeling, modern control theory, system identification, machine learning, deep learning, mathematics and data science). Additional goals for the algorithms involved localizing threats on-the-fly to know in which monitoring node the effects of attacks are present, and then providing accommodation to keep the system running uninterrupted much of the time in the presence of the attack. Accommodation had a performance goal of providing resiliency when up to 50% of monitoring nodes are in an attack state.

cybersecurity, cyber-physical↗

Methods for rapid identification of anomalous layers in laser powder bed fusion

In situ process monitoring is a key requirement for increased industry acceptance of powder bed Additive Manufacturing. As sensing technologies increase in maturity, attention must also be given to effective data exploration techniques. These data are often high-resolution and multi-modal, with each build consisting of thousands of layers. Here, in this paper, the authors propose two methods enabling users to rapidly identify layers of interest within a build. Both methods leverage results from deep learning based segmentations of in situ powder bed images. The first method is an unsupervised “reverse layer search” algorithm while the second method uses supervised machine learning.

36 MATERIALS SCIENCE↗

Opportunities for human factors in machine learning

Introduction The field of machine learning and its subfield of deep learning have grown rapidly in recent years. With the speed of advancement, it is nearly impossible for data scientists to maintain expert knowledge of cutting-edge techniques. This study applies human factors methods to the field of machine learning to address these difficulties. Methods Using semi-structured interviews with data scientists at a National Laboratory, we sought to understand the process used when working with machine learning models, the challenges encountered, and the ways that human factors might contribute to addressing those challenges. Results Results of the interviews were analyzed to create a generalization of the process of working with machine learning models. Issues encountered during each process step are described. Discussion Recommendations and areas for collaboration between data scientists and human factors experts are provided, with the goal of creating better tools, knowledge, and guidance for machine learning scientists.

97 MATHEMATICS AND COMPUTING↗

Multiscale Modeling Meets Machine Learning: What Can We Learn?

Machine learning is increasingly recognized as a promising technology in the biological, biomedical, and behavioral sciences. There can be no argument that this technique is incredibly successful in image recognition with immediate applications in diagnostics including electrophysiology, radiology, or pathology, where we have access to massive amounts of annotated data. However, machine learning often performs poorly in prognosis, especially when dealing with sparse data. This is a field where classical physics-based simulation seems to remain irreplaceable. In this review, we identify areas in the biomedical sciences where machine learning and multiscale modeling can mutually benefit from one another: Machine learning can integrate physics-based knowledge in the form of governing equations, boundary conditions, or constraints to manage ill-posted problems and robustly handle sparse and noisy data; multiscale modeling can integrate machine learn- ing to create surrogate models, identify system dynamics and parameters, analyze sensitivities, and quantify uncertainty to bridge the scales and understand the emergence of function. With a view towards applications in the life sciences, we discuss the state of the art of combining machine learning and multiscale modeling, identify applications and opportunities, raise open questions, and address potential challenges and limitations. We anticipate that it will stimulate discussion within the community of computational mechanics and reach out to other disciplines including mathematics, statistics, computer science, artificial intelligence, biomedicine, systems biology, and precision medicine to join forces towards creating robust and efficient models for biological systems.

machine learning, multiscale modeling, physics-bas↗

A framework for materials informatics education through workshops

The burgeoning field of materials informatics necessitates a focus on educating the next generation of materials scientists in the concepts of data science, artificial intelligence (AI), and machine learning (ML). In addition to incorporating these topics in undergraduate and graduate curricula, regular hands-on workshops present the most effective medium to initiate researchers to informatics and have them start applying the best AI/ML tools to their own research. With the help of the Materials Research Society (MRS), members of the MRS AI Staging Committee, and a dedicated team of instructors, we successfully conducted workshops covering the essential concepts of AI/ML as applied to materials data, at both the Spring and Fall Meetings in 2022, with plans to make this a regular feature in future meetings. Here, in this article, we discuss the importance of materials informatics education via the lens of these workshops, including details such as learning and implementing specific algorithms, the crucial nuts and bolts of ML, and using competitions to increase interest and participation.

36 MATERIALS SCIENCE↗

Applications and Techniques for Fast Machine Learning in Science

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science—the concept of integrating powerful ML methods into the real-time experimental data processing loop to accelerate scientific discovery. The material for the report builds on two workshops held by the Fast ML for Science community and covers three main areas: applications for fast ML across a number of scientific domains; techniques for training and implementing performant and resource-efficient ML algorithms; and computing architectures, platforms, and technologies for deploying these algorithms. We also present overlapping challenges across the multiple scientific domains where common solutions can be found. This community report is intended to give plenty of examples and inspiration for scientific discovery through integrated and accelerated ML solutions. This is followed by a high-level overview and organization of technical advances, including an abundance of pointers to source material, which can enable these breakthroughs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗