SEARCH · Engineering Papers
Results for “machine learning for science”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Self-Driving Microscopy for AI/ML-Enabled Physics Discovery and Materials Optimization
Materials are the bedrock of economy and foundation for all real-world technologies. The viability of space travel, grid energy storage, solar to fuels conversion, methane removal, and photovoltaic energy solutions hinge on the discovery and optimization of novel materials and rapid scaling toward manufacturing. The last 20 years have seen an exponential growth in the theoretical predictive capability for crystalline materials and small molecules. However, it is only in the last five years that we have seen the rapid expansion of high-throughput synthesis enabled by laboratory robotics and microfluidics, as well as a resurgence of combinatorial synthesis (Abolhasani and Kumacheva 2023; Epps and Abolhasani 2021; Jiang et al. 2022; Rajan 2008; Soldatov et al. 2021; Szymanski et al. 2023). Combinatorial synthesis, microfluidics, and ultimately dip-pen megalibraries have demonstrated the ability to “write” multicomponent nanomaterials at high throughput scale, generating millions of material examples in the 3D, 4D, and 5D composition spaces (Chen et al. 2016, 2019; Jibril et al. 2022).
Deploying Machine Learning Workflows into HPC environment
Outline: Workflows Overview; Common Workflow Language (CWL), Example of a CWL; BEE Overview; Machine Learning Components; Machine Learning Scientific Workflow using CWL → A new test case for BEE; Discussion: Benefits and Caveats of current ML workflow; Conclusion.
Ultrafast radiographic imaging and tracking: An overview of instruments, methods, data, and applications
Ultrafast radiographic imaging and tracking (U-RadIT) use state-of-the-art ionizing particle and light sources to experimentally study sub-nanosecond transients or dynamic processes in physics, chemistry, biology, geology, materials science and other fields. These processes are fundamental to modern technologies and applications, such as nuclear fusion energy, advanced manufacturing, communication, and green transportation, which often involve one mole or more atoms and elementary particles, and thus are challenging to compute by using the first principles of quantum physics or other forward models. One of the central problems in U-RadIT is to optimize information yield through, e.g. high-luminosity X-ray and particle sources, efficient imaging and tracking detectors, novel methods to collect data, and large-bandwidth online and offline data processing, regulated by the underlying physics, statistics, and computing power. We review and highlight recent progress in: (a.) Detectors such as high-speed complementary metal-oxide semiconductor (CMOS) cameras, hybrid pixelated array detectors integrated with Timepix4 and other application-specific integrated circuits (ASICs), and digital photon detectors; (b.) U-RadIT modalities such as dynamic phase contrast imaging, dynamic diffractive imaging, and four-dimensional (4D) particle tracking; (c.) U-RadIT data and algorithms such as neural networks and machine learning, and (d.) Applications in ultrafast dynamic material science using XFELs, synchrotrons and laser-driven sources. Hardware-centric approaches to U-RadIT optimization are constrained by detector material properties, low signal-to-noise ratio, high cost and long development cycles of critical hardware components such as ASICs. Interpretation of experimental data, including comparisons with forward models, is frequently hindered by sparse measurements, model and measurement uncertainties, and noise. Alternatively, U-RadIT make increasing use of data science and machine learning algorithms, including experimental implementations of compressed sensing. Machine learning and artificial intelligence approaches, refined by physics and materials information, may also contribute significantly to data interpretation, uncertainty quantification and U-RadIT optimization.
Linking repeat lidar with Landsat products for large scale quantification of fire-induced permafrost thaw settlement in interior Alaska
The permafrost–fire–climate system has been a hotspot in research for decades under a warming climate scenario. Surface vegetation plays a dominant role in protecting permafrost from summer warmth, thus, any alteration of vegetation structure, particularly following severe wildfires, can cause dramatic top–down thaw. A challenge in understanding this is to quantify fire-induced thaw settlement at large scales (>1000 km 2 ). In this study, we explored the potential of using Landsat products for a large-scale estimation of fire-induced thaw settlement across a well-studied area representative of ice-rich lowland permafrost in interior Alaska. Six large fires have affected ~1250 km 2 of the area since 2000. We first identified the linkage of fires, burn severity, and land cover response, and then developed an object-based machine learning ensemble approach to estimate fire-induced thaw settlement by relating airborne repeat lidar data to Landsat products. The model delineated thaw settlement patterns across the six fire scars and explained ~65% of the variance in lidar-detected elevation change. Our results indicate a combined application of airborne repeat lidar and Landsat products is a valuable tool for large scale quantification of fire-induced thaw settlement.
Single Cell RNA-Seq and Machine Learning Reveal Novel Subpopulations in Low-Grade Inflammatory Monocytes With Unique Regulatory Circuits
Subclinical doses of LPS (SD-LPS) are known to cause low-grade inflammatory activation of monocytes, which could lead to inflammatory diseases including atherosclerosis and metabolic syndrome. Sodium 4-phenylbutyrate is a potential therapeutic compound which can reduce the inflammation caused by SD-LPS. To understand the gene regulatory networks of these processes, we have generated scRNA-seq data from mouse monocytes treated with these compounds and identified 11 novel cell clusters. We have developed a machine learning method to integrate scRNA-seq, ATAC-seq, and binding motifs to characterize gene regulatory networks underlying these cell clusters. Using guided regularized random forest and feature selection, our method achieved high performance and outperformed a traditional enrichment-based method in selecting candidate regulatory genes. Our method is particularly efficient in selecting a few candidate genes to explain observed expression pattern. In particular, among 531 candidate TFs, our method achieves an auROC of 0.961 with only 10 motifs. Finally, we found two novel subpopulations of monocyte cells in response to SD-LPS and we confirmed our analysis using independent flow cytometry experiments. Our results suggest that our new machine learning method can select candidate regulatory genes as potential targets for developing new therapeutics against low grade inflammation.
De novo atomic protein structure modeling for cryoEM density maps using 3D transformer and HMM
Accurately building 3D atomic structures from cryo-EM density maps is a crucial step in cryo-EM-based protein structure determination. Converting density maps into 3D atomic structures for proteins lacking accurate homologous or predicted structures as templates remains a significant challenge. Here, we introduce Cryo2Struct, a fully automated de novo cryo-EM structure modeling method. Cryo2Struct utilizes a 3D transformer to identify atoms and amino acid types in cryo-EM density maps, followed by an innovative Hidden Markov Model (HMM) to connect predicted atoms and build protein backbone structures. Cryo2Struct produces substantially more accurate and complete protein structural models than the widely used ab initio method Phenix. Additionally, its performance in building atomic structural models is robust against changes in the resolution of density maps and the size of protein structures.
Using machine learning to jointly harness the strength of microscopic, fundamental-science driven and macroscopic, application-driven experiments
The PARADIGM project aims at accelerating progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application driven experiments to maximally reduce pertinent data uncertainties? Hence, we are bridging between microscopic experiments and data, and macroscopic simulations and experiments. Answering this question entails solving a high-dimensional and complex optimization problem which we solve with machine learning techniques.
Machine Learning-Based Atmospheric Phenomena Detection Platform
As the number of Earth pointing satellites has increased over the last several decades, the data volume retrieved from instruments onboard these satellites has also increased. It is expected that this trend will continue as more data intensive missions and small satellite constellations are launched. Currently, feature detection - namely atmospheric phenomena - in these datasets is performed manually and is thus not scalable with the growing data archives. Recent advancements in computational efficiency allow for the Earth science community to leverage machine learning to identify interesting atmospheric phenomena. Given the wide range of distinctive features in various atmospheric phenomena, a specialized machine learning model is required for accurate detection of these phenomena independently. The Phenomena Portal, developed at NASA IMPACT, is designed to provide visualization for the output from these machine learning models. In addition, detected events for each atmospheric phenomena are stored in a database that can be used to more easily use/subset larger spatiotemporal datasets. The user interface also incorporates additional features to enhance the user experience including spatiotemporal analysis, multiple base layer images, and a slider to filter events with lower probabilities of positive detection. Each detection supports user feedback on whether the detection is true or false that can then be stored and used to improve the machine learning model performance.
Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions
A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth independence and autonomy of mission operations. Here we present an overview of AI/ML architecture to support deep space mission goals, developed with leaders in the field. First, we focus on the fundamental biological research that supports our understanding of physiological responses to spaceflight, and we describe current efforts to support AI/ML research including data standardization and data engineering through maximally open and FAIR (findable, accessible, interoperable, reusable) databases and the generation of AI-ready datasets for reuse and analysis. We also discuss remote data management frameworks for research data as well as environmental and health data that are generated during deep space missions. We highlight several research projects that leverage data standardization and management for fundamental biological discovery to uncover the complex effects of space travel on living systems. Next, we provide an overview of cutting-edge AI/ML approaches that can be integrated to support remote monitoring and analysis during deep space missions, including generative models and large language models to learn the underlying biomedical patterns and predict outcomes or answer questions during off world medical scenarios. We also describe current AI/ML methods to support this research and monitoring through automated cloud-based labs which enable limited human intervention and closed-loop experimentation in remote settings. These labs could support mission autonomy by analyzing environmental data streams, and would be facilitated through in situ analytics capabilities to avoid sending large raw data files through low bandwidth communications. Finally, in the context of deep space missions with limited communications or access to medical advice from Earth, we describe a solution for integrated, real-time mission biomonitoring across hierarchical levels from continuous environmental monitoring, to wearables and point-of-care devices, to molecular and physiological monitoring. We introduce a precision space health system that will ensure that the future of space health is predictive, preventative, participatory and personalized.
Data from: Learning coagulation processes with combinatorially-invariant neural networks
This dataset contains all the necessary information to recreate the study presented in the paper entitled "Learning coagulation processes with combinatorially-invariant neural networks". This consists of (1) the aggregated output files used for machine learning, (2) the machine learning codes used to learn the presented models, (3) the PartMC model source code that was used to generate the simulation data and (4) the Python scripts used construct the scenario library for training and testing simulations. This data was used to investigate a method (combinatorally-invariant neural network) for learning the aerosol process of coagulation. This data may be useful for application of other methods.
Hydrology in the Age of Artificial Intelligence: From Fragmentation to Coherent Terrestrial Hydrosphere Science
The rapid rise of machine learning (ML) in hydrology has prompted debate about the discipline's scientific relevance. While ML often outperforms traditional models in streamflow prediction, we argue that this reflects a deeper limitation: persistent fragmentation of hydrological science itself. Narrow focus on isolated components has hindered the development of coherent, scale‐relevant understanding of the integrated terrestrial hydrosphere. This is illustrated, for example, by widely divergent estimates of groundwater–streamflow interactions and of water balance‐implied ongoing storage changes. We argue that hydrology's future lies not in choosing between ML and physics, but in integrating data‐driven and process‐based approaches to advance consistent, realistic, and societally relevant understanding of the terrestrial hydrosphere and its multifaceted roles in the Earth System.
A Machine Learning Approach to Predict Martensitic Transition Temperatures for Shape Memory Alloys
Shape memory alloys (SMAs) are a unique class of materials with several remarkable properties including shape recovery, superelasticity, etc. Especially important for many NASA applications is the ability to tune the martensitic phase transition temperature by varying the alloy composition. Nickel-titanium (NiTi) based alloys are the most widely studied of this class, with compositions involving ternary, quaternary, or higher additions being considered. Over the past several years, a significant database of SMA properties has been assembled by NASA researchers. Such a database is ideal for data science-based approaches including machine learning. We present results from a developed machine learning model capable of accurately predicting the transition temperature of SMAs across a wide range of compositions. Our model has the added benefit of interpretability and even provides confidence intervals for our predictions. This model will make rapid screening and design of new SMA materials possible. Predictions from the machine learning model can be validated by empirical and/or atomistic scale modeling.
Multiscale Modeling Meets Machine Learning: What Can We Learn?
Machine learning is increasingly recognized as a promising technology in the biological, biomedical, and behavioral sciences. There can be no argument that this technique is incredibly successful in image recognition with immediate applications in diagnostics including electrophysiology, radiology, or pathology, where we have access to massive amounts of annotated data. However, machine learning often performs poorly in prognosis, especially when dealing with sparse data. This is a field where classical physics-based simulation seems to remain irreplaceable. In this review, we identify areas in the biomedical sciences where machine learning and multiscale modeling can mutually benefit from one another: Machine learning can integrate physics-based knowledge in the form of governing equations, boundary conditions, or constraints to manage ill-posted problems and robustly handle sparse and noisy data; multiscale modeling can integrate machine learn- ing to create surrogate models, identify system dynamics and parameters, analyze sensitivities, and quantify uncertainty to bridge the scales and understand the emergence of function. With a view towards applications in the life sciences, we discuss the state of the art of combining machine learning and multiscale modeling, identify applications and opportunities, raise open questions, and address potential challenges and limitations. We anticipate that it will stimulate discussion within the community of computational mechanics and reach out to other disciplines including mathematics, statistics, computer science, artificial intelligence, biomedicine, systems biology, and precision medicine to join forces towards creating robust and efficient models for biological systems.
Nuclear Science User Facilities High Performance Computing: Artificial Intelligence and Machine Learning Software Training for Researchers
Artificial Intelligence and Machine Learning Software Training for Researchers
Beyond Human Vision: Exploring Materials with Machine Intelligence
Machine intelligence has the potential to revolutionize materials science, enabling autonomous synthesis, self-driving characterization, and accelerated modeling. However, despite the promise, successful implementation of these methods in day-to-day research remains a challenge. This talk will delve into the reasons behind this, exploring how truly intelligent experiments are hindered by opaque experiment control, a lack of domain-specific models, and human-centric design. Through a focus on the characterization of next-generation microelectronics and energy storage materials, I will share insights from both successful and failed attempts to implement machine intelligence. We will then explore the next steps necessary to unlock the full potential of machine intelligence in materials science, creating a future where intelligent systems work seamlessly alongside researchers to drive innovation and discovery.
Artificial Intelligence and QM/MM with a Polarizable Reactive Force Field for Next-Generation Electrocatalysts
Not Available