Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine Learning in Network Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Results of the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC)

Abstract Next-generation surveys like the Legacy Survey of Space and Time (LSST) on the Vera C. Rubin Observatory (Rubin) will generate orders of magnitude more discoveries of transients and variable stars than previous surveys. To prepare for this data deluge, we developed the Photometric LSST Astronomical Time-series Classification Challenge (PLAsTiCC), a competition that aimed to catalyze the development of robust classifiers under LSST-like conditions of a nonrepresentative training set for a large photometric test set of imbalanced classes. Over 1000 teams participated in PLAsTiCC, which was hosted in the Kaggle data science competition platform between 2018 September 28 and 2018 December 17, ultimately identifying three winners in 2019 February. Participants produced classifiers employing a diverse set of machine-learning techniques including hybrid combinations and ensemble averages of a range of approaches, among them boosted decision trees, neural networks, and multilayer perceptrons. The strong performance of the top three classifiers on Type Ia supernovae and kilonovae represent a major improvement over the current state of the art within astronomy. This paper summarizes the most promising methods and evaluates their results in detail, highlighting future directions both for classifier development and simulation needs for a next-generation PLAsTiCC data set.

79 ASTRONOMY AND ASTROPHYSICS↗

Processing NASA Earth Science Data on Nebula Cloud

Three applications were successfully migrated to Nebula, including S4PM, AIRS L1/L2 algorithms, and Giovanni MAPSS. Nebula has some advantages compared with local machines (e.g. performance, cost, scalability, bundling, etc.). Nebula still faces some challenges (e.g. stability, object storage, networking, etc.). Migrating applications to Nebula is feasible but time consuming. Lessons learned from our Nebula experience will benefit future Cloud Computing efforts at GES DISC.

Chen, Aijun↗

Survey of Deep Learning and Physics-Based Approaches in Computational Wave Imaging

Computational wave imaging (CWI) extracts hidden structure and physical properties of a volume of material by analyzing wave signals that traverse that volume. Applications include seismic exploration of the Earth’s subsurface, acoustic imaging and nondestructive testing (NDT) in material science, and ultrasound computed tomography (USCT) in medicine. Current approaches for solving CWI problems can be divided into two categories: those rooted in traditional physics and those based on deep learning. Physics-based methods stand out for their ability to provide high-resolution and quantitatively accurate estimates of acoustic properties within the medium. However, they can be computationally intensive and are susceptible to ill-posedness and nonconvexity typical of CWI problems. Machine learning (ML)-based computational methods have recently emerged, offering a different perspective to address these challenges. Diverse scientific communities have independently pursued the integration of deep learning in CWI. This review discusses how contemporary scientific ML techniques, and deep neural networks in particular, have been developed to enhance and integrate with traditional physics-based methods for solving CWI problems. We present a structured framework that consolidates existing research spanning multiple domains, including computational imaging, wave physics, and data science. This study concludes with important lessons learned from existing ML-based methods and identifies technical hurdles and emerging trends through a systematic analysis of the extensive literature on this topic.

42 ENGINEERING↗

Advancing Methodologies for Applying Machine Learning and Evaluating Spatiotemporal Models of Fine Particulate Matter (PM 2.5 ) Using Satellite Data Over Large Regions

Reconstructing the distribution of fine particulate matter (PM 2.5 ) in space and time, even far from ground monitoring sites, is an important exposure science contribution to epidemiologic analyses of PM 2.5 health impacts. Flexible statistical methods for prediction have demonstrated the integration of satellite observations with other predictors, yet these algorithms are susceptible to overfitting the spatiotemporal structure of the training datasets. We present a new approach for predicting PM 2.5 using machine-learning methods and evaluating prediction models for the goal of making predictions where they were not previously available. We apply extreme gradient boosting (XGBoost) modeling to predict daily PM 2.5 on a 1 x 1 km 2 resolution for a 13 state region in the Northeastern USA for the years 2000–2015 using satellite-derived aerosol optical depth and implement a recursive feature selection to develop a parsimonious model. We demonstrate excellent predictions of withheld observations but also contrast an RMSE of 3.11 μg/m 3 in our spatial cross-validation withholding nearby sites versus an overfit RMSE of 2.10 μg/m 3 using a more conventional random ten-fold splitting of the dataset. As the field of exposure science moves forward with the use of advanced machine-learning approaches for spatiotemporal modeling of air pollutants, our results show the importance of addressing data leakage in training, overfitting to spatiotemporal structure, and the impact of the predominance of ground monitoring sites in dense urban sub-networks on model evaluation. The strengths of our resultant modeling approach for exposure in epidemiologic studies of PM 2.5 include improved efficiency, parsimony, and interpretability with robust validation while still accommodating complex spatiotemporal relationships.

air pollution↗

Next-Generation Optical Sensing Technologies for Exploring Ocean Worlds - NASA FluidCam, MiDAR, and NeMO-Net

We highlight three emerging NASA optical technologies that enhance our ability to remotely sense, analyze, and explore ocean worlds–FluidCam and fluid lensing, MiDAR, and NeMO-Net. Fluid lensing is the first remote sensing technology capable of imaging through ocean waves without distortions in 3D at sub-cm resolutions. Fluid lensing and the purpose-built FluidCam CubeSat instruments have been used to provide refraction-corrected 3D multispectral imagery of shallow marine systems from unmanned aerial vehicles (UAVs). Results from repeat 2013 and 2016 airborne fluid lensing campaigns over coral reefs in American Samoa present a promising new tool for monitoring fine-scale ecological dynamics in shallow aquatic systems tens of square kilometers in area. MiDAR is a recently-patented active multispectral remote sensing and optical communications instrument which evolved from FluidCam. MiDAR is being tested on UAVs and autonomous underwater vehicles (AUVs) to remotely sense living and non-living structures in light-limited and analog planetary science environments. MiDAR illuminates targets with high-intensity narrowband structured optical radiation to measure an object’s spectral reflectance while simultaneously transmitting data. MiDAR is capable of remotely sensing reflectance at fine spatial and temporal scales, with a signal-to-noise ratio 10-10(exp 3) times higher than passive airborne and spaceborne remote sensing systems, enabling high-framerate multispectral sensing across the ultraviolet, visible, and near-infrared spectrum. Preliminary results from a 2018 mission to Guam show encouraging applications of MiDAR to imaging coral from airborne and underwater platforms whilst transmitting data across the air-water interface. Finally, we share NeMO-Net, the Neural Multi-Modal Observation & Training Network for Global Coral Reef Assessment. NeMO-Net is a machine learning technology under development that exploits high-resolution data from FluidCam and MiDAR for augmentation of low-resolution airborne and satellite remote sensing. NeMO-Net is intended to harmonize the growing diversity of 2D and 3D remote sensing with in situ data into a single open-source platform for assessing shallow marine ecosystems globally using active learning for citizen-science based training. Preliminary results from four-class Q17 coral classification have an accuracy of 94.4%. Together, these maturing technologies present promising scalable, practical, and cost-efficient innovations that address current observational and technological challenges in optical sensing of marine systems.

Ved Chirayath↗

Micropulse Lidar Cloud Mask Machine-Learning Value-Added Product Report

Cloud detection algorithms of various techniques have been developed and applied to atmospheric ground-based lidar data to identify cloud boundaries and produce clouds masks. While these algorithms are able to identify a wide variety of cloud types and conditions, it is often observed that the algorithms can still fail to accurately detect clouds that are readily discernible when inspecting the lidar imagery. Based on this observation, an alternative approach for cloud detection is to take advantage of machine-learning capabilities and the trained human eye as an interpreter of lidar images, and in turn, to train a neural network to recognize the desired features in the lidar data.

54 ENVIRONMENTAL SCIENCES↗

Determination of Infinite Dilution Activity Coefficients of Molecular Solutes in Ionic Liquids and Deep Eutectic Solvents by Factorization-Machine-Based Neural Networks

Widely known as “green solvents,” ionic liquids (ILs) and deep eutectic solvents (DESs) have been used as substitutes for traditional organic solvents in separation science. To achieve better separations using ILs and DESs, this work aimed to predict the infinite dilution activity coefficients (IDACs) of molecular solutes in these solvents with the state-of-the-art factorization-machine-based neural network (DeepFM). DeepFM combines the benefits of factorization machines and deep neural networks to learn both low-order and high-order interactions among features. The IDAC prediction model was established with 52,372 experimental IDAC datapoints including 260 solvents (252 ILs and 8 DESs) and 112 molecular solutes collected at various temperatures from 288.15 to 428.15 K. Chemical information describing the ILs, DESs, and molecular solutes was included in the IDAC prediction model, including chemical functional groups, molecular weights, and Abraham solvation parameters. The IDAC prediction model showed an improved accuracy compared with alternative models; additionally, DESs were included in the IDAC prediction model for the first time. Here, the model will reduce the energy and resources needed to optimize the selection of ILs and DESs for specific separations, which will promote the development of these green solvents for sustainable chemical processes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Protein model quality assessment using rotation–equivariant transformations on point clouds

Machine learning research concerning protein structure has seen a surge in popularity over the last years with promising advances for basic science and drug discovery. Working with macromolecular structure in a machine learning context requires an adequate numerical representation, and researchers have extensively studied representations such as graphs, discretized 3D grids, and distance maps. As part of CASP14, we explored a new and conceptually simple representation in a blind experiment: atoms as points in 3D, each with associated features. These features—initially just the basic element type of each atom—are updated through a series of neural network layers featuring rotation-equivariant convolutions. Starting from all atoms, we further aggregate information at the level of alpha carbons before making a prediction at the level of the entire protein structure. We find that this approach yields competitive results in protein model quality assessment despite its simplicity and despite the fact that it incorporates minimal prior information and is trained on relatively little data. As a result, its performance and generality are particularly noteworthy in an era where highly complex, customized machine learning methods such as AlphaFold 2 have come to dominate protein structure prediction.

59 BASIC BIOLOGICAL SCIENCES↗

A Wrapper to Use a Machine-Learning-Based Algorithm for Earthquake Monitoring

Seismology is one of the main sciences used to monitor volcanic activity worldwide. Fast, efficient, and accurate seismicity detectors are crucial to assess the activity level of a volcano in near–real time and to issue timely warnings. Traditional real–time seismic processing software uses phase onset pickers followed by a phase association algorithm to declare an event and estimate its location. The pickers typically do not identify whether the detected phase is a P or S arrival, which can have a negative impact on hypocentral location quality and complicates phase association. We implemented the deep–neural–network–based method PhaseNet to identify in real time P and S seismic waves on data from one– and three–component seismometers. We tuned the Earthworm binder_ew associator module to use the phase identification from PhaseNet to detect and locate the events, which we archive in a SeisComP3 database. We assessed the performance of the algorithm by comparing the results with existing catalogs built to monitor seismic and volcanic activity in Mayotte and the Lesser Antilles region. Our algorithm, which we refer to as PhaseWorm, showed promising results in both contexts and clearly outperformed the previous automatic method implemented in Mayotte. As a result, this innovative real–time processing system is now operational for seismicity monitoring in Mayotte and Martinique.

58 GEOSCIENCES↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data Science and Machine Learning for Genome Security

This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.

59 BASIC BIOLOGICAL SCIENCES↗

Machine-Learning Microstructure for Inverse Material Design

Metallurgy and material design have thousands of years’ history and have played a critical role in the civilization process of humankind. The traditional trial-and-error method has been unprecedentedly challenged in the modern era when the number of components and phases in novel alloys keeps increasing, with high-entropy alloys as the representative. New opportunities emerge for alloy design in the artificial intelligence era. Here a successful machine-learning (ML) method is developed to identify the microstructure images with eye-challenging morphology for a number of martensitic and ferritic steels. Assisted by it, a new neural-network method is proposed for the inverse design of alloys with 20 components, which can accelerate the design process based on microstructure. The method is also readily applied to other material systems given sufficient microstructure images. This work lays the foundation for inverse alloy design based on microstructure images with extremely similar features.

36 MATERIALS SCIENCE↗

Machine learning magnetism classifiers from atomic coordinates

The determination of magnetic structure poses a long-standing challenge in condensed matter physics and materials science. Experimental techniques such as neutron diffraction are resource-limited and require complex structure refinement protocols, while computational approaches such as first-principles density functional theory (DFT) need additional semi-empirical correction, and reliable prediction is still largely limited to collinear magnetism. Here, we present a machine learning model that aims to classify the magnetic structure by inputting atomic coordinates containing transition metal and rare earth elements. By building a Euclidean equivariant neural network that preserves the crystallographic symmetry, the magnetic structure (ferromagnetic, antiferromagnetic, and nonmagnetic) and magnetic propagation vector (zero or non-zero) can be predicted with an average accuracy of 77.8% and 73.6%. In particular, a 91% accuracy is reached when predicting no magnetic ordering even if the structure contains magneticelement(s). Ourworkrepresents onestepforwardtosolvingthegrand challenge of full magnetic structure determination.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Exploring NaCl-PuCl 3 molten salts with machine learning interatomic potentials and graph theory

Actinide molten salts are the basis of the liquid fuels used in molten salt reactors. Due to the inherent difficulties associated with high temperature and hazardous conditions, experimental investigations of fundamental properties of these materials are usually challenging. In this work, we describe the structure and transport of NaCl-PuCl 3 mixtures using computational techniques. Three compositions were considered (16, 25, and 36 mol% PuCl 3 ) over a temperature range (730 – 1257K) using ab initio molecular dynamics, which provided the necessary data sets for training machine learned interatomic potentials. Further, molecular dynamics simulations based on these potentials were then used to determine structure and transport properties. A substantial change was noted in the structure factor when increasing the PuCl 3 content from 25 to 36 mol%. This change is linked to the aggregation of larger Pu 3+ clusters. In addition, the similarity of the atomic environments of metal cations in molten salt systems to their solid states counterparts was investigated using an unsupervised learning technique. Finally, graph theory was employed to explore the structure and size of actinide networks. Consistent with the structure factor, a dense Pu 3+ intermolecular structure is observed within the 36 mol% PuCl 3 mixture. The structure of cation-cation inter-junctions is also discussed. In all cases, the diffusion of Pu 3+ is significantly lower than that of Na + and Cl - .

36 MATERIALS SCIENCE↗

Genomic fingerprints of the world’s soil ecosystems

Despite the explosion of soil metagenomic data, we lack a synthesized understanding of patterns in the distribution and functions of soil microorganisms. These patterns are critical to predictions of soil microbiome responses to climate change and resulting feedbacks that regulate greenhouse gas release from soils. To address this gap, we assay 1,512 manually curated soil metagenomes using complementary annotation databases, read-based taxonomy, and machine learning to extract multidimensional genomic fingerprints of global soil microbiomes. Our objective is to uncover novel biogeographical patterns of soil microbiomes across environmental factors and ecological biomes with high molecular resolution. We reveal shifts in the potential for (i) microbial nutrient acquisition across pH gradients; (ii) stress-, transport-, and redox-based processes across changes in soil bulk density; and (iii) greenhouse gas emissions across biomes. We also use an unsupervised approach to reveal a collection of soils with distinct genomic signatures, characterized by coordinated changes in soil organic carbon, nitrogen, and cation exchange capacity and in bulk density and clay content that may ultimately reflect soil environments with high microbial activity. Genomic fingerprints for these soils highlight the importance of resource scavenging, plant-microbe interactions, fungi, and heterotrophic metabolisms. Across all analyses, we observed phylogenetic coherence in soil microbiomes—more closely related microorganisms tended to move congruently in response to soil factors. Collectively, the genomic fingerprints uncovered here present a basis for global patterns in the microbial mechanisms underlying soil biogeochemistry and help beget tractable microbial reaction networks for incorporation into process-based models of soil carbon and nutrient cycling.

59 BASIC BIOLOGICAL SCIENCES↗

Developing a Machine-Learning-Based Processing Framework for Twitter and Other Crowdsourced Data

Crowdsourced data streams such as Twitter and other social media are important sources of real-time and historical global information for Earth science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we have been exploring the Twitter data stream for its potential in augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. To realize this potential, we need to increase the information density and enhance the quality of filtered precipitation tweets. We have implemented various components of a machine learning (ML)-based processing infrastructure for crowdsourced data that outputs, in this instance, useful and usable information derived from precipitation tweets. We have test enriched the Twitter stream with higher quality active tweets from those knowingly contributing to our effort and from existing crowdsourced programs (e.g., mPING, CoCoRaHS). We have experimented with various algorithms for processing tweets, including Naà ve Bayes, Convolutional Neural Network (CNN), Hierarchical Attention Network (HAN), and semi-supervised learning (with tri-training). Our current work focuses on (1) automated review of Earth science-related publications to determine relationships between discipline research needs and ML algorithms; (2) investigating Sequential Generative Adversarial Network (SeqGAN) for processing precipitation tweets for anomaly detection; and (3) managing crowdsourced data in a way that is compatible with existing NASA satellite data archives and using the data for ML applications. Key results include (1) network visualization of NLP-processed publications in various Earth science disciplines; (2) difference between GPM-linked, generated tweets and collected actual tweets that is small for GPM-determined light to moderate rain cases and high for GPM-determined heavy rain cases; and (3) identification of MongoDB for storing raw tweets and Zarr format for gridded tweets (compatible with GPM data). Our results have taken us a step closer to an operational ML-based tweet processing infrastructure and have already demonstrated that tweet-derived precipitation information is potentially useful for validation of Earth science satellite data.

Teng, William↗

Long-time integration of parametric evolution equations with physics-informed DeepONets

Ordinary and partial differential equations (ODEs/PDEs) play a paramount role in analyzing and simulating complex dynamic processes across all corners of science and engineering. In recent years machine learning tools are aspiring to introduce new effective ways of simulating such equations, however existing approaches are not able to reliably return stable and accurate predictions across long temporal horizons. We aim to address this challenge by introducing an effective framework for learning evolution operators that map random initial conditions to associated ODE/PDE solutions within a short time interval. Such operators can be parametrized by deep neural networks that are trained in an entirely self-supervised manner without requiring one to generate any paired input-output observations. Global long-time predictions across a range of initial conditions can be then obtained by iteratively evaluating the trained model using each prediction as the initial condition for the next evaluation step. Here, this introduces a new approach to temporal domain decomposition that is shown to be effective in performing accurate long-time simulations for a wide range of parametric ODE and PDE systems, from wave propagation, to reaction-diffusion dynamics and stiff chemical kinetics, introducing a new way of rapidly emulating non-equilibrium processes in science and engineering.

97 MATHEMATICS AND COMPUTING↗

Visualizing Uranium Crystallization from Melt: Experiment-Informed Phase Field Modeling and Machine Learning

The focus of this project was to observe and simulate the solidification of uranium metal at the crystallographic level from its molten state. Melting experiments were conducted at two different scales to observe microstructural evolution using either a laboratory-scale induction furnace (hundreds of grams of metal) or a microscope heating stage (hundreds of milligrams of metal), respectively. Experimental parameters and characterization data were then used to inform a phase field model of gamma-U crystal growth as dendrites with or without secondary phase impurities in the form of uranium carbide particles. Finally, training datasets were generated by the phase field model as inputs to a neural network, developed with the aim of providing a faster, cheaper surrogate model for microstructural simulations within a given parameter space. Progress is reported herein for each of these task areas. Ultimately, 1) an optical microscope heating stage capability has been stood-up for uranium metal solidification studies, 2) a phase field model was advanced to simulate multiple uranium grains growing in the presence of carbide impurity particles and 3) a neural network was constructed and optimized to predict the microstructure features of individually growing uranium crystals.

36 MATERIALS SCIENCE↗