Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Accurate models of the added mass force of a uniform random distribution of spherical particles or bubbles

The added mass force resulting from the acceleration of a body in a fluid is of fundamental and practical interest in dispersed multiphase flows. Euler–Lagrange (EL) and Euler–Euler (EE) simulations require closure terms for the added mass force in order to accurately couple the conserved variables between phases. Presently, a more thorough understanding of the added mass force in a multi-particle system is developed based on potential flow resulting in a resistance matrix formulation analogous to Stokesian dynamics. This formulation is then used to generate a dataset of added mass resistance matrices for large systems of randomly generated particles. This methodology is used to create a volume fraction corrected binary model for predicting the added mass force in large systems as well as generate statistics of the added mass force in such systems. This work provides clarification to the theory of the added mass force for particle clouds, and modelling options that may be implemented in existing EL and EE codes.

42 ENGINEERING↗

Remote Sensing and GIS data at 1km-grid over Chesapeake Bay used in “He et al. 2024, Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem”

The package contains the data layers used in “He et al. 2024, Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem”. The study aims to use multi-source remote sensing and GIS datasets to investigate the spatial heterogeneity and identify spatial zones with similar environmental characteristics and understand the primary driving factors affecting soil respiration within sub-ecosystems of the coastal ecosystem. We employed unsupervised hierarchical clustering analysis to identify spatial regions with distinct environmental characteristics, then determined the main driving factors using Random Forest regression and SHapley Additive exPlanations (SHAP). Spatial data layers include soil respiration, kernel Normalized Difference Vegetation Index (kNDVI) computed from Harmonized Landsat 8 and Sentinel-2 time series, climate variables from the Daymet dataset, land cover, biodiversity, topographical metrics, soil property, and tidal elevation.

54 ENVIRONMENTAL SCIENCES↗

Modeling Spatial Distribution of Snow Water Equivalent by Combining Meteorological and Satellite Data with Lidar Maps

Abstract An accurate characterization of the water content of snowpack, or snow water equivalent (SWE), is necessary to quantify water availability and constrain hydrologic and land surface models. Recently, airborne observations (e.g., lidar) have emerged as a promising method to accurately quantify SWE at high resolutions (scales of ∼100 m and finer). However, the frequency of these observations is very low, typically once or twice per season in the Rocky Mountains of Colorado. Here, we present a machine learning framework that is based on random forests to model temporally sparse lidar-derived SWE, enabling estimation of SWE at unmapped time points. We approximated the physical processes governing snow accumulation and melt as well as snow characteristics by obtaining 15 different variables from gridded estimates of precipitation, temperature, surface reflectance, elevation, and canopy. Results showed that, in the Rocky Mountains of Colorado, our framework is capable of modeling SWE with a higher accuracy when compared with estimates generated by the Snow Data Assimilation System (SNODAS). The mean value of the coefficient of determination R 2 using our approach was 0.57, and the root-mean-square error (RMSE) was 13 cm, which was a significant improvement over SNODAS (mean R 2 = 0.13; RMSE = 20 cm). We explored the relative importance of the input variables and observed that, at the spatial resolution of 800 m, meteorological variables are more important drivers of predictive accuracy than surface variables that characterize the properties of snow on the ground. This research provides a framework to expand the applicability of lidar-derived SWE to unmapped time points. Significance Statement Snowpack is the main source of freshwater for close to 2 billion people globally and needs to be estimated accurately. Mountainous snowpack is highly variable and is challenging to quantify. Recently, lidar technology has been employed to observe snow in great detail, but it is costly and can only be used sparingly. To counter that, we use machine learning to estimate snowpack when lidar data are not available. We approximate the processes that govern snowpack by incorporating meteorological and satellite data. We found that variables associated with precipitation and temperature have more predictive power than variables that characterize snowpack properties. Our work helps to improve snowpack estimation, which is critical for sustainable management of water resources.

54 ENVIRONMENTAL SCIENCES↗

Karhunen–Loève deep learning method for surrogate modeling and approximate Bayesian parameter estimation

We evaluate the performance of the Karhunen-Loève Deep Neural Network (KL-DNN) framework for surrogate modeling and approximate Bayesian parameter estimation in partial differential equation models. In the surrogate model, the Karhunen-Loève (KL) expansions are used for the dimensionality reduction of the number of unknown parameters and variables, and a deep neural network is employed to relate the reduced space of parameters to that of the state variables. The KL-DNN surrogate model is used to formulate a maximum-a-posteriori-like least-squares problem, which is randomized to draw samples of the posterior distribution of the parameters. We test the proposed framework for a hypothetical unconfined aquifer via comparison with the forward MODFLOW and inverse PEST++ iterative ensemble smoother (IES) solutions as well as the state-of-the-art Fourier neural operator (FNO) and deep operator networks (DeepONets) operator learning surrogate models. Our results show that the KL-DNN surrogate model outperforms FNO and DeepONet for forward predictions. For solving inverse problems, the randomized algorithm provides the same or more accurate Bayesian predictions of the parameters than IES as evidenced by the higher log-predictive probability of both the estimated parameter field and the forecast hydraulic head. The posterior mean obtained from the randomized algorithm is closer to the reference parameter field than that obtained with FNO as the maximum a posteriori estimate.

Approximate Bayesian inference↗

State-Level Trends in Renewable Energy Procurement via Solar Installation versus Green Electricity

In recent years, options for procuring renewable energy have increased, ranging from rooftop solar installation to utility green pricing to Community Choice Aggregation. These options vary in terms of costs and benefits to the consumer as well as grid integration implications. However, little is known regarding how the presence of a wide range of voluntary utility-scale renewable procurement options as well as their growth could affect adoption of distributed residential solar. To examine this relationship, we fit a two-stage least squares random effects regression model on panel data from 2016 to 2019 for all fifty US states plus the District of Columbia, controlling for variables that measure state-level policies, economic factors, and resource availability. Although there was no evidence of a strong relationship between demand for utility-scale and distributed options across all states, the state-level correlations suggest a wide variation between states including a positive, zero or negative relationship between utility-scale and distributed generation.

consumer demand↗

A Probabilistic Model for Global EMIC Wave Activity Using Van Allen Probes Observations

Electromagnetic ion cyclotron (EMIC) waves play a key role in radiation belt dynamics through resonant interactions. However, their low occurrence probability, high variability, and spatial intermittency pose challenges for accurate modeling. In this study, we present a machine learning (ML)-based global EMIC wave model built on the entire data set from the Van Allen Probes mission. To capture the distinct statistical characteristics of wave occurrence and amplitude, the model is separated into two modules: an occurrence model trained using ML techniques, and a wave amplitude model sampled from observed probability distributions. The input parameters are limited to real-time or predictable variables to ensure practical applicability. Our model shows strong performance across the entire test set and demonstrates improved predictive capability over a baseline random occurrence model, particularly during quiet geomagnetic conditions. Evaluation during both quiet and active periods confirms the model's ability to represent the clustered and intermittent nature of EMIC wave activity. Furthermore, the model provides global estimates of wave power, enabling integration with radiation belt electron data and showing signatures consistent with wave-induced scattering. We found a good correlation between the global wave activity from the model and relativistic electron observation by Van Allen Probes, regardless of the availability of in situ wave observations. The modular structure of the model also allows for straightforward expansion for additional wave properties, such as wave frequency, which can be modeled independently. This flexible, event-sensitive approach offers a promising framework for data-driven radiation belt simulations and space weather applications.

79 ASTRONOMY AND ASTROPHYSICS↗

Electricity Pricing aware Deep Reinforcement Learning based Intelligent HVAC Control

Recently, deep reinforcement learning (DRL) based intelligent control of Heating, Ventilation, and Air Conditioning (HVAC) has gained a lot of attention due to DRL's ability to optimally control HVAC for minimizing operational cost while maintaining resident's comfort. The success of such DRL-based techniques largely depends on the articulation of the problem in terms of states, actions, and reward function. Inclusion of the electricity pricing information in the problem formulation can play an important role in saving the cost of HVAC operation. However, less attention has been given in the literature on formulating well-crafted state features based on electricity pricing. In this work, we propose an approach for training the DRL model with a specific focus on feature engineering based on electricity pricing. During training, we generate random but sufficiently realistic electricity price signals so that the pre-trained DRL model is robust and adaptive to the dynamic and variable electricity prices. The validation results are encouraging and show the potential of ≈12%-15% savings in the one day cost of HVAC operation, proving the usefulness of including electricity pricing related features as state features.

Kurte, Kuldeep↗

Physical Insights From the Multidecadal Prediction of North Atlantic Sea Surface Temperature Variability Using Explainable Neural Networks

Abstract North Atlantic sea surface temperatures (NASST), particularly in the subpolar region, are among the most predictable in the world's oceans. However, the relative importance of atmospheric and oceanic controls on their variability at multidecadal timescales remain uncertain. Neural networks (NNs) are trained to examine the relative importance of oceanic and atmospheric predictors in predicting the NASST state in the Community Earth System Model 1 (CESM1). In the presence of external forcings, oceanic predictors outperform atmospheric predictors, persistence, and random chance baselines out to 25‐year leadtimes. Layer‐wise relevance propagation is used to unveil the sources of predictability, and reveal that NNs consistently rely upon the Gulf Stream‐North Atlantic Current region for accurate predictions. Additionally, CESM1‐trained NNs successfully predict the phasing of multidecadal variability in an observational data set, suggesting consistency in physical processes driving NASST variability between CESM1 and observations.

Geology↗

Predicting Rare Earth Element Potential in Produced and Geothermal Waters of the United States via Emergent Self-Organizing Maps

This work applies emergent self-organizing map (ESOM) techniques, a form of machine learning, in the multidimensional interpretation and prediction of rare earth element (REE) abundance in produced and geothermal waters in the United States. Visualization of the variables in the ESOM trained using the input data shows that each REE, with the exception of Eu, follows the same distribution patterns and that no single parameter appears to control their distribution. Cross-validation, using a random subsample of the starting data and only using major ions, shows that predictions are generally accurate to within an order of magnitude. Using the same approach, an abridged version of the U.S. Geological Survey Produced Waters Database, Version 2.3 (which includes both data from produced and geothermal waters) was mapped to the ESOM and predicted values were generated for samples that contained enough variables to be effectively mapped. Results show that in general, produced and geothermal waters are predicted to be enriched in REEs by an order of magnitude or more relative to seawater, with maximum predicted enrichments in excess of 1000-fold. Cartographic mapping of the resulting predictions indicates that maximum REE concentrations exceed values in seawater across the majority of geologic basins investigated and that REEs are typically spatially co-associated. The factors causing this co-association were not determined from ESOM analysis, but based on the information currently available, REE content in produced and geothermal waters is not directly controlled by lithology, reservoir temperature, or salinity.

Engle, Mark A. (ORCID:0000000152587374)↗

Computational multiphysics modeling of radioactive aerosol deposition in diverse human respiratory tract geometries

The evaluation of aerosol exposure relies on generic mathematical models that assume uniform particle deposition profiles over the human respiratory tract and do not account for subject-specific characteristics. Here we introduce a hybrid-automated computational workflow that generates personalized particle deposition profiles in 3D reconstructed human airways from computed tomography scans using Computational Fluid and Particle Dynamics simulations. This is the first large-scale study to consider realistic airways variability, where 380 lower and 40 upper human respiratory tract 3D geometries are reconstructed and parameterized. The data is clustered into nine groups using random forest regression. Computational fluid and particle dynamics simulations are conducted on these representative geometries using a realistic heavy-breathing respiratory cycle and radioactive iodine-131 as a source term. Monte Carlo radiation transport simulations are performed to obtain detailed energy deposition maps. Our findings emphasize the importance of personalized studies, as minor respiratory tract variations notably influence deposition patterns rather than global parameters of the lower airways, observing more than 30% variance in the mass deposition fraction.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Finding simplicity: unsupervised discovery of features, patterns, and order parameters via shift-invariant variational autoencoders *

Abstract Recent advances in scanning tunneling and transmission electron microscopies (STM and STEM) have allowed routine generation of large volumes of imaging data containing information on the structure and functionality of materials. The experimental data sets contain signatures of long-range phenomena such as physical order parameter fields, polarization, and strain gradients in STEM, or standing electronic waves and carrier-mediated exchange interactions in STM, all superimposed onto scanning system distortions and gradual changes of contrast due to drift and/or mis-tilt effects. Correspondingly, while the human eye can readily identify certain patterns in the images such as lattice periodicities, repeating structural elements, or microstructures, their automatic extraction and classification are highly non-trivial and universal pathways to accomplish such analyses are absent. We pose that the most distinctive elements of the patterns observed in STM and (S)TEM images are similarity and (almost-) periodicity, behaviors stemming directly from the parsimony of elementary atomic structures, superimposed on the gradual changes reflective of order parameter distributions. However, the discovery of these elements via global Fourier methods is non-trivial due to variability and lack of ideal discrete translation symmetry. To address this problem, we explore the shift-invariant variational autoencoders (shift-VAEs) that allow disentangling characteristic repeating features in the images, their variations, and shifts that inevitably occur when randomly sampling the image space. Shift-VAEs balance the uncertainty in the position of the object of interest with the uncertainty in shape reconstruction. This approach is illustrated for model 1D data, and further extended to synthetic and experimental STM and STEM 2D data. We further introduce an approach for training shift-VAEs that allows finding the latent variables that comport to known physical behavior. In this specific case, the condition is that the latent variable maps should be smooth on the length scale of the atomic lattice (as expected for physical order parameters), but other conditions can be imposed. The opportunities and limitations of the shift VAE analysis for pattern discovery are elucidated.

97 MATHEMATICS AND COMPUTING↗

New Capabilities for Sampling Tools

This report will discuss new capabilities that have been added to the sample.py and uniform_sampler.py codes. The codes have been updated to allow for both log-uniform sampling and categorical variables. The categorical variables do not have to be numeric. The different values for a categorical variable are specified in a limits file by having spaces between them. The difference between sample.py and uniform_sampler.py is that sample.py is for generating random samples and uniform_sampler.py is for generating samples or points on a fixed grid. After sourcing a file to set the environment, execute the codes with the commands: sample.py and uniform sampler.py .

97 MATHEMATICS AND COMPUTING↗

Investigating performance and variability of NIF ICF experiments with deep learning

The parameter space involved in designing an inertial confinement fusion shot at the National Ignition Facility (NIF) is massively multi-dimensional and the cost of a single shot makes a comprehensive set of sensitivity studies in the laboratory impractical. The use of machine learning to overcome these challenges has gained popularity and has had several successful applications by the scientific community. We extend on these efforts by training a neural network (NN) on information about the experimental design, engineering elements, and drive asymmetry to predict with uncertainty the neutron yield of an experiment. We find the measured and model predicted values are in good agreement, with an R 2 value of 0.91 for a randomly selected test dataset. Almost all the predicted 95% credible intervals contain the corresponding measured value for both training and test datasets. We identify correlations picked up by the NN between the shot design, yield, and variability and use them to motivate shot sensitivity studies. The first shot to exceed the Lawson-like ignition criteria (N210808) was conducted at the NIF and subsequent shots studied the design’s robustness. In a follow-up shot to N210808, our model predicts capsule quality to be the main performance degradation mechanism that prevented the shot from repeating previous performance levels. Shot N221204 was the first shot to exceed a target energy gain of 1. Our model predicts increased yield with reduced coast time for a N221204 study and greater variability for designs with lower peak powers at constant yield. The model’s fast prediction speed and uncertainty prediction are useful for identifying interesting design paths that could warrant further investigation with conventional simulations to search for robust high yield designs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Adaptive Interface-PINNs (AdaI-PINNs) for inverse problems: Determining material properties for heterogeneous systems

Here, we determine spatially varying discontinuous material properties using a domain-decomposition based physics-informed neural networks (PINNs) framework named the Adaptive Interface-PINNs or AdaI-PINNs (Roy et al., 2024). We propose the use of distinct neural networks for the field variables and material properties within each material, utilizing adaptive activation functions. While the neural networks across different materials share the same weights and biases, their activation functions are uniquely tailored using a hyperparameter that influences the slope of the activation function. The proposed framework is tested on several one-dimensional and two-dimensional benchmark examples, and its performance is compared with conventional PINNs and existing domain-decomposition PINNs frameworks, namely, the Multi-domain physics-informed neural network (M-PINN), and the eXtended physics-informed neural networks (XPINNs). The results demonstrate that the proposed approach can determine randomly distributed discontinuous material properties with an L 2 error of $\mathscr{O}$ (10 -3 ) for the material property and the root-mean-square error of $\mathscr{O}$ (10 -3 ) for the primary variable while the other approaches yield errors that are approximately two orders of magnitude larger (that is, $\mathscr{O}$ (10 -1 )). Moreover, the spatial distribution of material properties obtained using the proposed framework is in close agreement with the true distribution, whereas the other approaches fare much worse. Additionally, the proposed approach is approximately 40% faster than its competitors, indicating its potential as a robust alternative for solving inverse problems in heterogeneous materials.

36 MATERIALS SCIENCE↗

Megadrought: A Series of Unfortunate La Niña Events?

Megadroughts are multidecadal periods of aridity more persistent than most droughts during the instrumental period. Paleoclimate evidence suggests that megadroughts occur in many parts of the world, including North America, Central America, western Europe, eastern Asia, and northern Africa. It remains unclear to what extent such megadroughts require external forcing or whether they can arise from internal climate variability alone. A novel statistical–dynamical approach is used to evaluate the possibility that such events arise solely as a function of interannual tropical sea surface temperature (SST) variations. A statistical emulator of tropical SST variations is constructed by using an empirical moving-blocks bootstrap approach that randomly samples multiyear sequences of the observational SST record. This approach preserves the power spectrum, seasonal cycle, and spatial pattern of El Niño-Southern Oscillation (ENSO) but removes longer timescale fluctuations embedded in the observational record. These resampled SST anomalies are then used to force an atmospheric model (the Community Atmosphere Model Version 5). As megadroughts emerge in this run, they should, therefore, be solely a consequence of La Niña sequences combined with internal atmospheric variability and persistence driven by soil moisture storage and other land-surface processes. We indeed find that megadroughts in this simulation have an amplitude-duration rate that is generally indistinguishable from the rate documented in paleoclimate records of the western United States. Our findings support the idea that megadroughts may occur randomly when the unforced climate system evolves freely over a sufficiently long period of time, implying that an unforced unusual but statistically plausible series of La Niña events may be sufficient to generate megadrought.

54 ENVIRONMENTAL SCIENCES↗

Scale-dependent spatial variabilities of hydrological exchange flows and transit time in a large regulated river

Hydrological exchange flows (HEF) across the river-aquifer interface and the associated residence time of river water in the aquifer have important implications for contaminant plume migration and biogeochemical processes in the river corridor. HEFs and residence time are influenced by both subsurface physical features and hydrologic forcing related to the transport process, which can exhibit complex spatial and temporal variations. In this study, we used a massively parallel subsurface flow model and a particle-tracking model to study the influences of different control factors on spatial variability of HEFs and residence time distributions (RTD) in the Hanford Reach of the Columbia River in Washington State. A total number of 100M particles were randomly injected in time and space and then tracked in a model domain that covers a 51-km 2 area (15.1M model cells). We used hourly river stages and groundwater levels to drive the model to provide dynamic velocity fields for the particle tracking in the simulation period that was longer than 2 years. The groundwater flow simulation and particle-tracking results provide the first comprehensive assessment of the spatial distribution of HEFs and residence time in large complex river corridors. Overall, our results show that the aquifer hydrogeological structure has the strongest correlation with the extent and magnitude of exchange flux. The residence time exhibits complex patterns that are impacted by all the river geomorphologic, hydrodynamic, and hydrogeologic factors and are strongly correlated with the downwelling ratio of exchange flux. The new insights gained through this study can be used to support the development of reduced-order models of HEFs and RTDs for large complex river systems.

54 ENVIRONMENTAL SCIENCES↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Understanding the Seismic Ground Motion Spatial Variability Using Network Analysis Community Detection

This project is to explore ground motion spatial distribution using a new approach graph-based network analysis. In this study, we combine a large-N seismic array and graph analytics to explore spatial variability and correlation at a local scale using small local and regional earthquakes. In this method, each seismic station is modeled as a node and the similarities of the waveforms that represent ground motions between two stations are modeled as edges. By analyzing this graph network using the similarity matrices and community detection algorithm, we can group the stations spatially with similar patterns. A random forest algorithm is used to reveal the important features that affect the spatial grouping. The result suggests site conditions, and how they interact with the incident seismic wavefield, strongly condition the spatial correlation of ground motion. Future progress in characterizing ground motion spatial variability will require dense wavefield measurements, either through nodal deployments, or perhaps distributed acoustic sensing measurements of seismic wavefields.

58 GEOSCIENCES↗