Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

AMIA KDDM Working Group Collaborative Workshop: Enriching Electronic Health Records with Social Determinants of Health to Improve Outcomes and Health Equity

Prior research has demonstrated that social determinants of health (SDoH) are major drivers of health outcomes and contributors to widespread health inequities. It was estimated that, in the United States, SDoH could be responsible for up to 40% of all preventable deaths, significantly higher than the 10-15% for which better medical care is responsible. Public health interventions that target SDoH are instrumental for improving health outcomes and reducing long-standing health inequities. Currently, most mainstream EHR vendors have implemented SDoH screeners in their EHR systems. However, the utility of the screeners is low, rendering patient-level SDoH still widely unavailable in the structured fields. SDoH are sometimes mentioned in free-text clinical notes (e.g., social context section) where natural language processing (NLP) can be applied to extract relevant information. Contextual-level SDoH can be identified from multiple data sources, many of which are publicly available and spatiotemporally linked to EHR data. As such, there is an opportunity for the KDDM research community to create innovative solutions to draw meaningful insights by creating and using rich data with SDoH to improve health outcomes while reducing disparities. In this workshop organized by AMIA Knowledge Discovery and Data Mining Working Group (AMIA KDDM WG), we will invite world-leading experts from academia, national laboratories, and life science industry with varied backgrounds in biomedical informatics, epidemiology, data science, machine learning, natural language processing, and pediatric cardiology to discuss the best practice of capturing, standardizing, and using SDoH information in various applications aiming at improving outcomes and health equity.

He, Zhe↗

Citizen science coupled with machine learning to quantify green-blue infrastructure cooling potential in Maricopa County, Arizona

Here, this study investigates the spatiotemporal cooling performance of green and blue infrastructure (GBI) in the Dobson Ranch urban neighborhood in Phoenix, Arizona. We leveraged citizen science near-surface (2 m) air temperature (Tair) measurements to train a highly accurate Tair predicting LightGBM machine learning model (R 2 : 0.986, MAE: 0.251 °C, RMSE: 0.585 °C). On June 16, 2024, the park area exhibited approximately 1 °C cooling effect (relative to the neighborhood mean) during both day and night. In contrast, the nearby artificial lake exhibited a stronger cooling effect of 2.4 °C during the day but a slight warming of 0.3 °C at night. At 00:00, locations 50 m downwind of the park were 0.3 °C warmer than the park, while locations 50 m upwind were 0.8 °C warmer. At 11:00, we observed that the downwind area is 0.8 °C cooler and the upwind area is 0.6 °C warmer—at the same 50 m distances relative to the park. We also observed 1 °C cooler and warmer effects respectively at the same 50 m downwind and upwind locations at 19:00 on June 17, 2024. Our data-driven analysis highlights potential limitations of car-traverse measurements, showing that failure to account for temporal variations during the traverse can lead to overestimation of Tair at night and underestimation during the day. Our analysis also showed only a weak correlation (coefficient: 0.48) between Landsat-derived land surface temperature (LST) and model predicted Tair at the time of the local Landsat overpass (∼11.00). This highlights the potential error of relying solely on LST for human thermal exposure analysis—particularly within the heterogenous built-environment.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning and Data Science to Advance Laboratory Earthquake Prediction and Illuminate the Mechanics of Precursors to Failure

Earthquakes represent one of our greatest natural hazards and in recent years human induced seismicity is adding to the threat. Even a modest improvement in the ability to forecast devastating large earthquakes or smaller shallow events associated with fluid injection could save thousands of lives and billions of dollars. Current efforts to forecast earthquakes are limited by knowledge of earthquake physics and hampered by a lack of reliable lab or field observations. However, recent work has provided a critical opportunity for advancement. We have found: 1) clear and consistent precursors prior to earthquake-like failure in the laboratory and 2) that lab earthquakes can be predicted using machine learning (ML). These works show that stick-slip failure events –the lab equivalent of earthquakes– are preceded by a cascade of micro-failure events that radiate elastic energy in a manner that foretells catastrophic failure. Remarkably, ML predicts the fault zone stress state, the failure time and in some cases the magnitude of lab earthquakes. In addition, the observations include clear precursors to failure in the form of changes in fault zone properties prior to lab earthquakes. Precursors have been observed in previous laboratory studies but their origin is poorly understood and their possible connection to ML based earthquake prediction is unknown. The work conducted under our project has dramatically expanded these efforts. We have developed an integrated data science approach to illuminate the physics of earthquake precursors and lab earthquake prediction. Our work has accelerated the development of ML, artificial intelligence (AI), and related data science approaches by providing massive data sets that are tightly connected to critical scientific problems and by bringing together leading subject matter experts and data scientists. Earthquake physics involves phenomena that are far from equilibrium. Our work has leveraged data science methods to illuminate these phenomena and investigate how they relate to earthquake prediction. In addition to a large database with many types of labeled events that is available to everyone, our work has advanced the fundamental understanding of seismic forecasting, earthquake physics, and fault rheology

58 GEOSCIENCES↗

Machine Learning of Plasma Science for Next Generation Microelectronics (Project Final Report)

Low temperature plasmas (LTPs) are an enabling technology behind reducing device dimensions and the continuation of Moore’s Law. It is estimated that 40-45% of all process steps necessary to manufacture semiconductor devices involve LTPs. However, challenges in plasma process design and continuous incorporation of novel materials for new device architectures are pushing the limits of what is possible with current plasma technology. For example, creating higher aspect ratio structures and etching features at the atomic scale both require finer control of the ion energy/velocity at wafer surfaces. To support these types of future innovations in the plasma processing systems that Sandia and the DOE rely upon, we have developed novel diagnostics, simulations, and machine learning capabilities to discover, characterize, and predict plasma phenomena affecting the ion energy/velocity distribution function (IEDF). These efforts also supported research program development and external collaboration with industry and academia through Sandia’s Plasma Research Facility (PRF). This report will focus on the following topics and accomplishments of this three year LDRD project, briefly summarized.

42 ENGINEERING↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Data-centric machine learning in quantum information science

Abstract We propose a series of data-centric heuristics for improving the performance of machine learning systems when applied to problems in quantum information science. In particular, we consider how systematic engineering of training sets can significantly enhance the accuracy of pre-trained neural networks used for quantum state reconstruction without altering the underlying architecture. We find that it is not always optimal to engineer training sets to exactly match the expected distribution of a target scenario, and instead, performance can be further improved by biasing the training set to be slightly more mixed than the target. This is due to the heterogeneity in the number of free variables required to describe states of different purity, and as a result, overall accuracy of the network improves when training sets of a fixed size focus on states with the least constrained free variables. For further clarity, we also include a ‘toy model’ demonstration of how spurious correlations can inadvertently enter synthetic data sets used for training, how the performance of systems trained with these correlations can degrade dramatically, and how the inclusion of even relatively few counterexamples can effectively remedy such problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A prospective on machine learning challenges, progress, and potential in polymer science

Abstract Artificial intelligence and machine learning (ML) continue to see increasing interest in science and engineering every year. Polymer science is no different, though implementation of data-driven algorithms in this subfield has unique challenges barring widespread application of these techniques to the study of polymer systems. In this Prospective, we discuss several critical challenges to implementation of ML in polymer science, including polymer structure and representation, high-throughput techniques and limitations, and limited data availability. Promising studies targeting resolution of these issues are explored, and contemporary research demonstrating the potential of ML in polymer science despite existing obstacles are discussed. Finally, we present an outlook for ML in polymer science moving forward. Graphical Abstract

Struble, Daniel C. (ORCID:0009000093410612)↗

Physics in the Machine: Integrating Physical Knowledge in Autonomous Phase-Mapping

Application of artificial intelligence (AI), and more specifically machine learning, to the physical sciences has expanded significantly over the past decades. In particular, science-informed AI, also known as scientific AI or inductive bias AI, has grown from a focus on data analysis to now controlling experiment design, simulation, execution and analysis in closed-loop autonomous systems. The CAMEO (closed-loop autonomous materials exploration and optimization) algorithm employs scientific AI to address two tasks: learning a material system’s composition-structure relationship and identifying materials compositions with optimal functional properties. By integrating these, accelerated materials screening across compositional phase diagrams was demonstrated, resulting in the discovery of a best-in-class phase change memory material. Key to this success is the ability to guide subsequent measurements to maximize knowledge of the composition-structure relationship, or phase map. In this work we investigate the benefits of incorporating varying levels of prior physical knowledge into CAMEO’s autonomous phase-mapping. This includes the use of ab-initio phase boundary data from the AFLOW repositories, which has been shown to optimize CAMEO’s search when used as a prior.

97 MATHEMATICS AND COMPUTING↗

Multi-Attribute Subset Selection enables prediction of representative phenotypes across microbial populations

The interpretation of complex biological datasets requires the identification of representative variables that describe the data without critical information loss. This is particularly important in the analysis of large phenotypic datasets (phenomics). Here we introduce Multi-Attribute Subset Selection (MASS), an algorithm which separates a matrix of phenotypes (e.g., yield across microbial species and environmental conditions) into predictor and response sets of conditions. Using mixed integer linear programming, MASS expresses the response conditions as a linear combination of the predictor conditions, while simultaneously searching for the optimally descriptive set of predictors. We apply the algorithm to three microbial datasets and identify environmental conditions that predict phenotypes under other conditions, providing biologically interpretable axes for strain discrimination. MASS could be used to reduce the number of experiments needed to identify species or to map their metabolic capabilities. The generality of the algorithm allows addressing subset selection problems in areas beyond biology.

59 BASIC BIOLOGICAL SCIENCES↗

A unifying Bayesian framework for merging X-ray diffraction data

Novel X-ray methods are transforming the study of the functional dynamics of biomolecules. Key to this revolution is detection of often subtle conformational changes from diffraction data. Diffraction data contain patterns of bright spots known as reflections. To compute the electron density of a molecule, the intensity of each reflection must be estimated, and redundant observations reduced to consensus intensities. Systematic effects, however, lead to the measurement of equivalent reflections on different scales, corrupting observation of changes in electron density. Here, we present a modern Bayesian solution to this problem, which uses deep learning and variational inference to simultaneously rescale and merge reflection observations. We successfully apply this method to monochromatic and polychromatic single-crystal diffraction data, as well as serial femtosecond crystallography data. We find that this approach is applicable to the analysis of many types of diffraction experiments, while accurately and sensitively detecting subtle dynamics and anomalous scattering.

59 BASIC BIOLOGICAL SCIENCES↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

SMART – A Comprehensive Research and Development Program to Demonstrate Application of Machine Learning for Supporting CCS Deployment

The objective of the US Department of Energy’s SMART Initiative, i.e., Science-informed Machine Learning (ML) for Accelerating Real-Time Decisions in Subsurface Applications, is to showcase how the utilization of ML can significantly improve efficiency and effectiveness of field-scale commercial carbon storage operations. This paper will present the results from the current phase of SMART (field deployment) for demonstrating the applicability of ML-based tools and workflows for: (a) virtual learning during the pre-injection permitting phase, (b) advanced storage reservoir imaging to better characterize fractures and faults, and (c) dynamic storage reservoir modelling and optimization to inform operational decision making and visualization of system evolution.

Siriwardane, Hema↗

2022 American Conference on Neutron Scattering (ACNS 2022)

The 11th American Conference on Neutron Scattering (ACNS 2022) will be held on June 5-9, 2022, in Boulder, CO. The Conference will provide essential information on the breadth and depth of current neutron-related research worldwide. Hosted by the Neutron Scattering Society of America, this year’s Conference will feature a combination of invited and contributed talks, poster sessions, and tutorials. Topics of the conference are: Advances in Neutron Facilities, Instrumentation and Software: Developments in sources, instrumentation, sample environments and control software. Hard Condensed Matter: Magnetism, correlated metals, quantum/topological materials, superconductors, ferroelectrics, multiferroics, glasses, and disorder phenomena. Submissions outlining examples of neutron scattering in industrial and engineering applications involving hard condensed matter systems are also encouraged. Soft Matter: Neutron studies of soft materials and related fields including in situ and in operando studies. Polymers, surfactants, emulsions, gels, nanoparticles, colloidal suspensions and more. Submissions of computational studies or applications of machine learning beneficial to neutron scattering experiments, as well as examples of neutron scattering in industrial and engineering applications are strongly encouraged. Biology, Biophysics and Biotechnology: Neutron studies of biological and biologically relevant systems. Proteins, bio membranes, biological assemblies, natural materials, nucleic acids, drug-delivery platforms and biomedical systems. Submissions of computational studies or applications of machine learning beneficial to biological neutron scattering experiments, as well as examples of neutron scattering in applied research involving biological systems, are strongly encouraged. Materials Chemistry and Energy: Neutron-based studies of functional materials and materials for energy applications. Examples include porous materials such as metal organic frameworks (MOFs), zeolites; phosphors; novel pigments; electrolytes; catalysts; ionic conductors/cathode materials; photovoltaic materials (hybrid perovskites); thermoelectrics; magnetocalorics/electrocalorics. Structural Materials and Engineering: Neutron scattering studies of materials and engineering processes including structural materials, concrete and metals, as well as engineering processes including combustion, corrosion, additive manufacturing, and others. Neutron Physics: Fundamental physical studies of the neutron and related areas. Emerging Applications in Neutron Scattering: Machine Learning and Data Science: Advances in computing power have contributed to rapidly evolving machine learning and data science fields that can be leveraged to the benefit of the neutron scattering community. The purpose of this session is to highlight recent advances in machine learning and data science and to serve as the foundation of a parallel data and computation track highlighting computation advances and applications in neutron scattering throughout the conference.

36 MATERIALS SCIENCE↗