Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Anomalous electroweak physics unraveled via evidential deep learning

The ever-growing ecosystem of beyond standard model (BSM) calculations and parametrizations has motivated the development of systematic methods for making quantitative cross-comparisons over the wide range of possible models, especially with controllable uncertainties. In this setting, the language of uncertainty quantification (UQ) furnishes useful metrics for assessing statistical overlaps and discrepancies among BSM and related models. In this study, we leverage recent machine learning (ML) developments in evidential deep learning (EDL) for UQ to separate data (aleatoric) and knowledge (epistemic) uncertainties in a model-discrimination setting. We construct several potentially BSM-motivated scenarios for the anomalous electroweak interaction (AEWI) of neutrinos with nucleons in deep inelastic scattering ( v DIS). These scenarios are then quantitatively mapped, as a demonstration, alongside Monte Carlo replicas of the CT18 PDFs used to calculate the $\varDelta \chi ^{2}$ statistic for a typical multi-GeV v DIS experiment, CDHSW. Our framework effectively highlights areas of model agreement and provides a classification of out-of-distribution (OOD) samples. By offering the opportunity to quantitatively understand model overlaps, the approach presented in this work can help facilitate efficient BSM model exploration and exclusion for future New Physics searches.

AI↗

Spatiotemporal Super-Resolution with Generative Machine Learning for Creating Renewable Energy Resource Data Under Climate Change Scenarios

As we plan for a future with higher penetrations of renewables and increasing electrification, it becomes more important to understand how the electricity grid will operate under a variety of weather events. We must also consider that the weather our future grid will experience will be different and possibly more extreme than the historical weather that we have extensive data for. We can use data from global climate models (GCMs) to help understand how our climate may change over the next several decades, but there is often a significant gap between the low-resolution GCM data and the high-resolution weather data required to study power systems under specific weather events. Therefore, our objective in this work is to develop tools that can bridge this gap by using low-resolution GCM data to create realistic high-resolution weather datasets that can be used to study renewable energy generation and electricity demand. To accomplish this objective, we have developed a set of generative machine learning models that can rapidly downscale GCM daily average output data at an approximate grid resolution of 100km to hourly data at an approximate 4 km grid resolution. The models can be used to create high resolution data from nearly any GCM included in the Coupled Model Intercomparison Project (CMIP) Phase 5 or 6. Our methods include all datasets regularly used to study the integration of wind and solar power plants as well as changes in electricity demand due to heating and cooling loads. These models and datasets enable power systems modelers to study climate change-influenced weather events and their impact on the grid. We have downscaled and validated wind, solar, temperature, and humidity data with very promising results. The generative machine learning methods are computationally efficient and produce data that has similar statistical characteristics to current state-of-the-art historical datasets. We have trained initial generative models and produced an initial dataset collectively referred to as Sup3rCC: Super-Resolved Renewable Energy Resource Data with Climate Change Impacts. The data covers a (mostly) historical period from 2015-2025 and a future period from 2050-2059. We have also taken hypothetical high-electrification load data and scaled the heating and cooling loads with respect to the 2050-2059 high-resolution Sup3rCC meteorology. The results show how future levels of renewable energy generation and electrified load may be impacted by climate change, setting the stage for capacity expansion models to consider a dynamic climate through model years.

climate change↗

Learning likelihood ratios with neural network classifiers

The likelihood ratio is a crucial quantity for statistical inference in science that enables hypothesis testing, construction of confidence intervals, reweighting of distributions, and more. Many modern scientific applications, however, make use of data- or simulation-driven models for which computing the likelihood ratio can be very difficult or even impossible. By applying the so-called “likelihood ratio trick,” approximations of the likelihood ratio may be computed using clever parametrizations of neural network-based classifiers. A number of different neural network setups can be defined to satisfy this procedure, each with varying performance in approximating the likelihood ratio when using finite training data. We present a series of empirical studies detailing the performance of several common loss functionals and parametrizations of the classifier output in approximating the likelihood ratio of two univariate and multivariate Gaussian distributions as well as simulated high-energy particle physics datasets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Modern Senicide in the Face of a Pandemic: An Examination of Public Discourse and Sentiment About Older Adults and COVID-19 Using Machine Learning

Objectives This study examined public discourse and sentiment regarding older adults and COVID-19 on social media and assessed the extent of ageism in public discourse. Methods Twitter data (N = 82,893) related to both older adults and COVID-19 and dated from January 23 to May 20, 2020, were analyzed. We used a combination of data science methods (including supervised machine learning, topic modeling, and sentiment analysis), qualitative thematic analysis, and conventional statistics. Results The most common category in the coded tweets was “personal opinions” (66.2%), followed by “informative” (24.7%), “jokes/ridicule” (4.8%), and “personal experiences” (4.3%). The daily average of ageist content was 18%, with the highest of 52.8% on March 11, 2020. Specifically, more than 1 in 10 (11.5%) tweets implied that the life of older adults is less valuable or downplayed the pandemic because it mostly harms older adults. A small proportion (4.6%) explicitly supported the idea of just isolating older adults. Almost three-quarters (72.9%) within “jokes/ridicule” targeted older adults, half of which were “death jokes.” Also, 14 themes were extracted, such as perceptions of lockdown and risk. A bivariate Granger causality test suggested that informative tweets regarding at-risk populations increased the prevalence of tweets that downplayed the pandemic. Discussion Ageist content in the context of COVID-19 was prevalent on Twitter. Information about COVID-19 on Twitter influenced public perceptions of risk and acceptable ways of controlling the pandemic. Finaly, public education on the risk of severe illness is needed to correct misperceptions.

60 APPLIED LIFE SCIENCES↗

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES↗

Actinide Molten Salts: A Machine-Learning Potential Molecular Dynamics Study

We know that actinide molten salts represent a class of important materials in nuclear energy. Understanding them at a molecular level is critical to proper and optimal design of relevant technological applications. Yet, owing to the complexity of electronic structure due to the 5f orbitals, computational studies of heavy elements in condensed phases using ab initio potentials to study the structure and dynamics of these elements embedded in molten salts are difficult. This lack of efficient computational protocols makes it difficult to obtain information on properties that require extensive statistical sampling like transport. To tackle this problem, we adopted a machine-learning approach to study ThCl 4 -NaCl and UCl 3 -NaCl binary systems. The machine-learning potential, with the density functional theory accuracy, allows us to obtain long molecular dynamics trajectories (ns) for large systems (10 3 atoms) at a considerably low computing cost, thereby efficiently gaining information about their bonding structures, thermodynamics, and dynamics at a range of temperature. We observed a considerable change in the coordination environments of actinide elements and their characteristic coordination-sphere lifetime. Our study also suggests that actinides in molten salts may not follow well known entropy-scaling laws.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

pnnl/frequency_sensitivity

This research code base includes functions to compute statistics of image distributions in (spatial DFT) frequency space, train deep learning image classifiers on these distributions with variable depth and weight decay, and finally measure the sensitivity of the trained models to perturbations along (spatial DFT) frequency components.

Central, PNNL Developer↗

Full event particle-level unfolding with variable-length latent variational diffusion

The measurements performed by particle physics experiments must account for the imperfect response of the detectors used to observe the interactions. One approach, unfolding, statistically adjusts the experimental data for detector effects. Recently, generative machine learning models have shown promise for performing unbinned unfolding in a high number of dimensions. However, all current generative approaches are limited to unfolding a fixed set of observables, making them unable to perform full-event unfolding in the variable dimensional environment of collider data. A novel modification to the variational latent diffusion model (VLD) approach to generative unfolding is presented, which allows for unfolding of high- and variable-dimensional feature spaces. The performance of this method is evaluated in the context of semi-leptonic t\bar{t} t t ‾ production at the Large Hadron Collider.

Shmakov, Alexander↗

NASA Tech Briefs, June 2014

Topics include: Real-Time Minimization of Tracking Error for Aircraft Systems; Detecting an Extreme Minority Class in Hyperspectral Data Using Machine Learning; KSC Spaceport Weather Data Archive; Visualizing Acquisition, Processing, and Network Statistics Through Database Queries; Simulating Data Flow via Multiple Secure Connections; Systems and Services for Near-Real-Time Web Access to NPP Data; CCSDS Telemetry Decoder VHDL Core; Thermal Response of a High-Power Switch to Short Pulses; Solar Panel and System Design to Reduce Heating and Optimize Corridors for Lower-Risk Planetary Aerobraking; Low-Cost, Very Large Diamond-Turned Metal Mirror; Very-High-Load-Capacity Air Bearing Spindle for Large Diamond Turning Machines; Elevated-Temperature, Highly Emissive Coating for Energy Dissipation of Large Surfaces; Catalyst for Treatment and Control of Post-Combustion Emissions; Thermally Activated Crack Healing Mechanism for Metallic Materials; Subsurface Imaging of Nanocomposites; Self-Healing Glass Sealants for Solid Oxide Fuel Cells and Electrolyzer Cells; Micromachined Thermopile Arrays with Novel Thermo - electric Materials; Low-Cost, High-Performance MMOD Shielding; Head-Mounted Display Latency Measurement Rig; Workspace-Safe Operation of a Force- or Impedance-Controlled Robot; Cryogenic Mixing Pump with No Moving Parts; Seal Design Feature for Redundancy Verification; Dexterous Humanoid Robot; Tethered Vehicle Control and Tracking System; Lunar Organic Waste Reformer; Digital Laser Frequency Stabilization via Cavity Locking Employing Low-Frequency Direct Modulation; Deep UV Discharge Lamps in Capillary Quartz Tubes with Light Output Coupled to an Optical Fiber; Speech Acquisition and Automatic Speech Recognition for Integrated Spacesuit Audio Systems, Version II; Advanced Sensor Technology for Algal Biotechnology; High-Speed Spectral Mapper; "Ascent - Commemorating Shuttle" - A NASA Film and Multimedia Project DVD; High-Pressure, Reduced-Kinetics Mechanism for N-Hexadecane Oxidation; Method of Error Floor Mitigation in Low-Density Parity-Check Codes; X-Ray Flaw Size Parameter for POD Studies; Large Eddy Simulation Composition Equations for Two-Phase Fully Multicomponent Turbulent Flows; Scheduling Targeted and Mapping Observations with State, Resource, and Timing Constraints;

Source record↗

At Risk Population Estimates for Belarus, Poland and Slovakia with Machine Learning

High-resolution gridded population modeling is crucial for various applications, including disaster response planning, infectious disease spread modeling, climate change impact estimation, policy development, and more. Multiple gridded population datasets have been developed, each tailored to meet specific objectives. Among them, LandScan Global dataset is designed to represent ambient and unwarned population distributions. However, this dataset relies on a statistical approach that requires manual adjustments, making it time consuming and labour intensive. Existing machine learning (ML) methods often train and test at different spatial resolutions, potentially leading to inflated results, and they rely on Census population totals for disaggregation. To address these limitations, in this study we developed population estimates using ML models trained and tested at a consistent 30 arc-second resolution (≈1 square kilometer), specifically using Random Forest (RF) and XGBoost. These models were trained on 2020 datum to predict for 2021 for three countries: Belarus, Poland, and Slovakia. Our findings show that both RF (MAE varies from 5.75 to 13.25) and XGBoost (MAE varies from 8.15 to 23.44) model performance is close to LandScan Global estimates. Furthermore, neither of the models performed the best across all grid cells: the RF model was more effective in areas with lower populations, while XGBoost excelled in more densely populated regions. The proposed approach can be used for countries where the Census data is not available.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term. The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission (TRMM) Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters↗

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission TRMM Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters↗

Machine learning to accelerate screening for Marcus reorganization energies

Understanding and predicting the charge transport properties of π-conjugated materials is an important challenge for designing new organic electronic devices, such as solar cells, plastic transistors, light-emitting devices, and chemical sensors. A key component of the hopping mechanism of charge transfer in these materials is the Marcus reorganization energy which serves as an activation barrier to hole or electron transfer. While modern density functional methods have proven to accurately predict trends in intramolecular reorganization energy, such calculations are computationally expensive. In this work, we outline active machine learning methods to predict computed intramolecular reorganization energies of a wide range of polythiophenes and their use toward screening new compounds with low internal reorganization energies. Our models have an overall root mean square error (RMSE) of ±0.113 eV, but a much smaller RMSE of only ±0.036 eV on the new screening set. Since the larger error derives from high-reorganization energy compounds, the new method is highly effective to screen for compounds with potentially efficient charge transport parameters.

14 SOLAR ENERGY↗

Reducing Southern Ocean Shortwave Radiation Errors in the ERA5 Reanalysis with Machine Learning and 25 Years of Surface Observations

Earth system models struggle to simulate clouds and their radiative effects over the Southern Ocean, partly due to a lack of measurements and targeted cloud microphysics knowledge. We have evaluated biases of downwelling shortwave radiation in the ERA5 climate reanalysis using 25 years (1995–2019) of summertime surface measurements, collected on the Research and Supply Vessel (RSV) Aurora Australis, the Research Vessel (R/V) Investigator, and at Macquarie Island. During October–March daylight hours, the ERA5 simulation of SW down exhibited large errors (mean bias = 54 W m -2 , mean absolute error = 82 W m -2 , root-mean-square error = 132 W m -2 , and R 2 = 0.71). To determine whether we could improve these statistics, we bypassed ERA5’s radiative transfer model for SW down with machine learning–based models using a number of ERA5’s gridscale meteorological variables as predictors. These models were trained and tested with the surface measurements of SW down using a 10-fold shuffle split. An extreme gradient boosting (XGBoost) and a random forest–based model setup had the best performance relative to ERA5, both with a near complete reduction of the mean bias error, a decrease in the mean absolute error and root-mean-square error by 25% ± 3%, and an increase in the R 2 value of 5% ± 1% over the 10 splits. Large improvements occurred at higher latitudes and cyclone cold sectors, where ERA5 performed most poorly. We further interpret our methods using Shapley additive explanations. Our results indicate that data-driven techniques could have an important role in simulating surface radiation fluxes and in improving reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross-dataset tasks. Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Supervised machine learning-based multivariate regression of parallel closures for a high-collisionality deuterium-carbon plasma

Many plasmas of interest in laboratory experiments and space consist of multiple ion species. In tokamak edge plasmas, for instance, ionized impurities expelled from the vessel wall influence plasma transport. When describing multi-species plasmas using fluid equations, we need accurate closure relations to close the set of fluid equations. In this study, we introduce the development of fitting formulas for parallel closures using supervised machine learning, in conjunction with the recent closure theory, considering multi-ion collisions and arbitrary ion temperatures. We apply this approach to a high-collisionality deuterium-carbon plasma and demonstrate its effectiveness. As a result, the machine learning-based method for developing practical and accurate closures can be extended to a wider range of plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Transboundary determinants of avian zoonotic infectious diseases: challenges for strengthening research capacity and connecting surveillance networks

As the climate changes, global systems have become increasingly unstable and unpredictable. This is particularly true for many disease systems, including subtypes of highly pathogenic avian influenzas (HPAIs) that are circulating the world. Ecological patterns once thought stable are changing, bringing new populations and organisms into contact with one another. Wild birds continue to be hosts and reservoirs for numerous zoonotic pathogens, and strains of HPAI and other pathogens have been introduced into new regions via migrating birds and transboundary trade of wild birds. With these expanding environmental changes, it is even more crucial that regions or counties that previously did not have surveillance programs develop the appropriate skills to sample wild birds and add to the understanding of pathogens in migratory and breeding birds through research. For example, little is known about wild bird infectious diseases and migration along the Mediterranean and Black Sea Flyway (MBSF), which connects Europe, Asia, and Africa. Focusing on avian influenza and the microbiome in migratory wild birds along the MBSF, this project seeks to understand the determinants of transboundary disease propagation and coinfection in regions that are connected by this flyway. Through the creation of a threat reduction network for avian diseases (Avian Zoonotic Disease Network, AZDN) in three countries along the MBSF (Georgia, Ukraine, and Jordan), this project is strengthening capacities for disease diagnostics; microbiomes; ecoimmunology; field biosafety; proper wildlife capture and handling; experimental design; statistical analysis; and vector sampling and biology. Here, we cover what is required to build a wild bird infectious disease research and surveillance program, which includes learning skills in proper bird capture and handling; biosafety and biosecurity; permits; next generation sequencing; leading-edge bioinformatics and statistical analyses; and vector and environmental sampling. Creating connected networks for avian influenzas and other pathogen surveillance will increase coordination and strengthen biosurveillance globally in wild birds.

54 ENVIRONMENTAL SCIENCES↗