Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “variability prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Predictability of the early summer surface air temperature over Western South Asia

Variability of the Surface Air Temperature (SAT) over the Western South Asia (WSA) region leads to frequent heatwaves during the early summer (May-June) season. The present study uses the European Centre for Medium-Range Weather Forecast’s fifth-generation seasonal prediction system, SEAS5, from 1981 to 2022 based on April initial conditions (1-month lead) to assess the SAT predictability during early summer season. The goal is to evaluate the SEAS5’s ability to predict the El Niño-Southern Oscillation (ENSO) related interannual variability and predictability of the SAT over WSA, which is mediated through upper-level (200-hPa) geopotential height anomalies. This teleconnection leads to anomalously warm surface conditions over the region during the negative ENSO phase, as observed in the reanalysis and SEAS5. We evaluate SEAS5 prediction skill against two observations and three reanalyses datasets. The SEAS5 SAT prediction skill is higher with high spatial resolution observations and reanalysis datasets compared to the ones with low-resolution. Overall, SEAS5 shows reasonable skill in predicting SAT and its variability over the WSA region. Moreover, the predictability of SAT during La Niña is comparable to El Niño years over the WSA region.

54 ENVIRONMENTAL SCIENCES↗

A HPC Theory-Guided Machine Learning Cyberinfrastructure for Communicating Hydrometeorological Data Across Scales

High-resolution predictions of hydrometeorological variables are critical for supporting hydropower generation decisions and flood control at hydroelectric power plants. Traditional climate and hydrologic models rely on the numerical simulation of detailed physical processes. Therefore, running these simulations is time-, labor-, and computation-intensive. Improving the spatial and temporal resolution in these modeling outputs could lead to cubic increases in both the simulation time and computational demands, rendering high-resolution hydrometeorological predictions expensive and impractical. Many past studies apply the super resolution (SR) technique to downscale climate models using deep learners. However, deep learners are deemed “black-boxes,” as their derivation processes from low-resolution outputs to high-resolution outputs are often hidden. Their results are difficult for domain scientists to interpret and validate. Thus, there is a need for an exploratory machine learning approach that can partially integrate domain-specific theory and knowledge into the data-driven mapping process between simulation outputs of different spatial scales. The domain-specific theory and knowledge can be incorporated into the data model through an inductive approach in which process-related environmental variables are used and analyzed as key drivers (i.e., environmental surrogates) to reflect the complex physical processes. Many of these variables, such as land use land cover, soil types, topography, digital elevation, air temperature, and various watershed characteristics, can be directly measured through sensors or remote sensing techniques. Additionally, SR applications that can downscale hydrological and hydrodynamics models to efficiently produce high-resolution (1 m) flood depth grids are still rare. Since the flood depth grid can be used to support critical decisions for flood control operation at hydroelectric power plants, it is crucial to enable an SR-based capability for interpolating high-resolution flood inundation maps.

13 HYDRO ENERGY↗

DNA Viral Diversity, Abundance, and Functional Potential Vary across Grassland Soils with a Range of Historical Moisture Regimes

Soil viruses are abundant, but the influence of the environment and climate on soil viruses remains poorly understood. Here, we addressed this gap by comparing the diversity, abundance, lifestyle, and metabolic potential of DNA viruses in three grassland soils with historical differences in average annual precipitation, low in eastern Washington (WA), high in Iowa (IA), and intermediate in Kansas (KS). Bioinformatics analyses were applied to identify a total of 2,631 viral contigs, including 14 complete viral genomes from three deep metagenomes (1 terabase [Tb] each) that were sequenced from bulk soil DNA. An additional three replicate metagenomes (~0.5 Tb each) were obtained from each location for statistical comparisons. Identified viruses were primarily bacteriophages targeting dominant bacterial taxa. Both viral and host diversity were higher in soil with lower precipitation. Viral abundance was also significantly higher in the arid WA location than in IA and KS. More lysogenic markers and fewer clustered regularly interspaced short palindromic repeats (CRISPR) spacer hits were found in WA, reflecting more lysogeny in historically drier soil. More putative auxiliary metabolic genes (AMGs) were also detected in WA than in the historically wetter locations. The AMGs occurring in 18 pathways could potentially contribute to carbon metabolism and energy acquisition in their hosts. Structural equation modeling (SEM) suggested that historical precipitation influenced viral life cycle and selection of AMGs. The observed and predicted relationships between soil viruses and various biotic and abiotic variables have value for predicting viral responses to environmental change.

59 BASIC BIOLOGICAL SCIENCES↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

Development of a coupled experimental–computational approach for engineering optimization of spout-fluidized bed particle coating systems

The design of spout-fluidized bed (SFB) coating systems for nuclear particle fuels typically relies on trial-and-error processes, comprising iterative and time-consuming coating deposition experiments and post-deposition characterization. At an engineering scale, this approach to guided SFB system design is inefficient, highlighting the need for streamlined experimental methodologies which can correlate fluidization conditions to downstream coating outcomes. In this study, we combine time-resolved particle image velocimetry (PIV) with CFD–DEM simulations to benchmark hydrodynamic behavior in a 3D spout-fluidized bed. By exploiting easily accessible optical measurements of particle motion at the bed wall and within the spouting region, we obtain quantitative velocity fields that can be directly compared with model predictions of the occluded bed region, without resorting to complex imaging and characterization techniques such as X-ray or magnetic resonance tomography. Experimental benchmarking reveals strong agreement between CFD–DEM and PIV in the spout and annulus regions, while discrepancies near the wall highlight areas for future model development. Here, the proposed integrated experimental–numerical framework will enable a direct connection between measured variables and numerically predicted fluidization performance of dense, surrogate nuclear particle fuel feedstock such that experimental SFB component design can be rapidly evaluated, informing design decisions for nozzle geometry and operating conditions. Future work will extend this framework by correlating quantified fluidization metrics across nozzle geometries and operating conditions with the resulting coating morphology, microstructure, and uniformity. Establishing these correlations will enable predictive links between hydrodynamic performance and coating quality, providing a rational, scalable basis for optimizing SFB design prior to coating deposition.

CFD/DEM↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Performance analysis and comparison of data-driven models for predicting indoor temperature in multi-zone commercial buildings

Building thermal models, which characterize the properties of a building’s envelope and thermal mass, are essential for accurate indoor temperature and cooling/heating demand prediction. Because of their flexibility and ease of use, data-driven models are increasingly used. Here, this study compared and analyzed the performance of gray-box (resistance-capacitance) and black-box (recurrent neural network) models for predicting indoor air temperature in a real multi-zone commercial building. The developed resistance-capacitance model served as a benchmark model for which full sets of temporal data and building information were used as inputs. The recurrent neural network models were trained and tested assuming various available types and amounts of temporal data and known building physical information to investigate the effects of data and information availability. Feature importance analysis was conducted to select the key variables for different prediction targets under different scenarios. This research provides guidance in selecting an appropriate building thermal response modeling method based on the measured data availability, building physical information, and application.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Controls and variability of soil respiration temperature sensitivity across China

Understanding the temperature sensitivity (Q 10 ) of soil respiration is critical for benchmarking the potential intensity of regional and global terrestrial soil carbon fluxes-climate feedbacks. Although field observations have demonstrated the strong spatial heterogeneity of Q 10 , a significant knowledge gap still exists regarding to the factors driving spatial and temporal variabilities of Q 10 at regional scales. Here, therefore, we used a machine learning approach to predict Q 10 from 1994 to 2016 with a spatial resolution of 1 km across China from 515 field observations at 5 cm soil depth using climate, soil and vegetation variables. Predicted Q 10 varied from 1.54 to 4.17, with an area-weighted average of 2.52. There was no significant temporal trend for Q 10 (p = 0.32), but annual vegetation production (indicated by normalized difference vegetation index, NDVI) was positively correlated to it (p < 0.01). Spatially, soil organic carbon (SOC) was the most important driving factor in 62 % of the land area across China, and varied greatly, demonstrating soil controls on the spatial pattern of Q 10 . These findings highlighted different environmental controls on the spatial and temporal pattern of soil respiration Q 10 , which should be considered to improve global biogeochemical models used to predict the spatial and temporal patterns of soil carbon fluxes to ongoing climate change.

54 ENVIRONMENTAL SCIENCES↗

An ecological niche model to predict the geographic distribution of Haemagogus janthinomys, Dyar, 1921 a yellow fever and Mayaro virus vector, in South America

Yellow fever virus (YFV) has a long history of impacting human health in South America. Mayaro virus (MAYV) is an emerging arbovirus of public health concern in the Neotropics and its full impact is yet unknown. Both YFV and MAYV are primarily maintained via a sylvatic transmission cycle but can be opportunistically transmitted to humans by the bites of infected forest dwelling Haemagogus janthinomys Dyar, 1921. To better understand the potential risk of YFV and MAYV transmission to humans, a more detailed understanding of this vector species’ distribution is critical. This study compiled a comprehensive database of 177 unique Hg . janthinomys collection sites retrieved from the published literature, digitized museum specimens and publicly accessible mosquito surveillance data. Covariate analysis was performed to optimize a selection of environmental (topographic and bioclimatic) variables associated with predicting habitat suitability, and species distributions modelled across South America using a maximum entropy (MaxEnt) approach. Our results indicate that suitable habitat for Hg . janthinomys can be found across forested regions of South America including the Atlantic forests and interior Amazon.

60 APPLIED LIFE SCIENCES↗

Deep Learning-based Surrogate Model for Efficient Reservoir Simulation in Large-scale Geological Carbon Storage: Application in IBDP Dataset

This project introduces an advanced deep learning (DL)-based surrogate modeling approach to enhance the efficiency and accuracy of large-scale geological carbon storage (GCS) simulations. Using the Illinois Basin Decatur Project (IBDP) dataset as training data, the study employs a residual U-Net architecture to predict critical state variables such as pressure and CO₂ saturation, as well as CO₂ plume migration. By incorporating key geological parameters (e.g., porosity, permeability, and rock facies) and physics-informed inputs like the diffusive time of flight and time step, the DL model effectively reduces computational complexity while maintaining robust physical constraints. Compared to traditional simulators like Eclipse, the DL model achieves remarkable accuracy, with a root mean square error (RMSE) of 1.57 psi for pressure and 0.007 for saturation, and dramatically reduces computational time from hours to just 69.9 seconds for 50-step simulations. These results demonstrate the potential of innovative DL methodologies to improve the predictivity and operational efficiency of GCS simulations, providing a reliable foundation for decision-making in CCS operations. Supported by the SMART initiative, this project underscores the success of leveraging computational innovations to advance CCS technologies.

advanced deep learning↗

Modeling and simulation of transitional Rayleigh–Taylor flow with partially averaged Navier–Stokes equations

In this work, the partially averaged Navier–Stokes (PANS) equations are used to predict the variable-density Rayleigh–Taylor (RT) flow at Atwood number 0.5 and maximum Reynolds number 500. This is a prototypical problem of material mixing, featuring laminar, transitional, and turbulent flow, instabilities and coherent structures, density fluctuations, and production of turbulence kinetic energy by both shear and buoyancy mechanisms. These features pose numerous challenges to modeling and simulation, making the RT flow ideal to develop the validation space of the recently proposed PANS Besnard–Harlow–Rauenzahn-linear eddy viscosity model closure. The numerical simulations are conducted at different levels of physical resolution and test three approaches to set the parameters $f_\phi$ defining the range of physically resolved scales. The computations demonstrate the efficiency (accuracy vs cost) of the PANS model predicting the spatiotemporal development of the RT flow. Results comparable to large-eddy simulations and direct numerical simulations are obtained at significantly lower physical resolution without the limitations of the Reynolds-averaged Navier–Stokes equations in these transitional flows. The data also illustrate the importance of appropriate selection of the physical resolution and the resolved fraction of each dependent quantity $\phi$ of the turbulent closure, $f_\phi$. These two aspects determine the ability of the model to resolve the flow phenomena not amenable to modeling by the closure and, as such, the computations’ fidelity.

Navier Stokes equations↗

An adaptive adversarial domain adaptation approach for corn yield prediction

Recently, statistical machine learning and deep learning methods have been widely explored for corn yield prediction. Though successful, machine learning models generated within a specific spatial domain often lose their validity when directly applied to new regions. To address this issue, we designed an unsupervised adaptive domain adversarial neural network (ADANN). Specifically, through domain adversarial training, the ADANN model reduced the impact of domain shift by projecting data from different domains into the same subspace. Also, the ADANN model was designed to be trained in an adaptive way, which guaranteed the model can learn the domain-invariant features and perform accurate yield prediction simultaneously. Informative variables including time-series vegetation indices and sequential weather observations were first collected from multiple data sources and aggregated to the county level. Then, we trained the ADANN model with the extracted features and corresponding reported county-level corn yield from the U.S. Department of Agriculture (USDA). Finally, the trained model was evaluated in four testing years 2016–2019. The U.S. corn belt was used as the study area and counties under study were grouped into two diverse ecological regions. Overall, the experimental results showed that the developed ADANN model had better performance than three other state-of-the-art machine learning models in both local experiments (train and test in the same region) and transfer experiments (train and test in different regions). As the first study using adversarial learning for crop yield prediction, this research demonstrates a novel solution for improving model transferability on crop yield prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Atikokan Digital Twin: Machine learning in a biomass energy system

The Atikokan Generating Station, operated by Ontario Power Generation, has a 200 MW, biomass-fired tower boiler that operates on a dispatch schedule with a five-minute cycle. The boiler is generally operated in the range of 40–100 MW using two of five burner levels. In order to optimize boiler performance, we propose the implementation of a unique digital twin. Our digital twin abstraction couples Bayesian inference from science-based models and from observations (machine learning) with decision theory to predict operating-variable set points that optimize the physical asset (the boiler) in the presence of uncertainty (artificial intelligence). We focus this paper on the continuous Bayesian machine learning part of the Atikokan Digital Twin; we discuss decision theory in a companion paper. We identify and learn about 12 operational, model, and measured-output parameters and their uncertainties from high-fidelity, science-based simulations of the Atikokan boiler and from the observed measurements at the power plant. Since the goal of the Atikokan Digital Twin is to implement it online in real time, we require fast function evaluations for the quantities of interest extracted from the simulations in the Bayesian analysis. We use Gaussian process regression/interpolation to create accurate, robust surrogate models. We define the Bayesian priors and likelihood function and solve for the posterior distributions of the 12 parameters. Here we then propagate these distributions (i.e., parameters with uncertainty) into the predicted distributions of 790 quantities of interest to learn about the relative importance of various sources of error including experimental, model, and operating-parameter errors.

09 BIOMASS FUELS↗

Multilevel Monte Carlo methods for the Grad-Shafranov free boundary problem

The equilibrium configuration of a plasma in an axially symmetric reactor is described mathematically by a free boundary problem associated with the celebrated Grad-Shafranov equation. The presence of uncertainty in the model parameters introduces the need to quantify the variability in the predictions. This is often done by computing a large number of model solutions on a computational grid for an ensemble of parameter values and then obtaining estimates for the statistical properties of solutions. In this study, we explore the savings that can be obtained using multilevel Monte Carlo methods, which reduce costs by performing the bulk of the computations on a sequence of spatial grids that are coarser than the one that would typically be used for a simple Monte Carlo simulation. We examine this approach using both a set of uniformly refined grids and a set of adaptively refined grids guided by a discrete error estimator. Numerical experiments show that multilevel methods dramatically reduce the cost of simulation, with cost reductions typically on the order of 60 or more and possibly as large as 200. Furthermore, adaptive griding results in more accurate computation of geometric quantities such as x-points associated with the model.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Using probabilistic solar power forecasts to inform flexible ramp product procurement for the California ISO

How can independent system operators (ISOs) take advantage of probabilistic solar forecasts to lower generation costs and improve reliability of power systems? We discuss one three-step approach for doing so, focusing on how such forecasts might help the California Independent System Operator (CAISO) prepare unexpected net load ramps, where net load equals gross demand minus wind and solar production. First, we enhance an existing solar forecasting system to provide well-calibrated hours-ahead probabilistic forecasts. We then relate the degree of uncertainty reflected in the forecasted prediction intervals (independent variables) to error distributions for net load ramp forecasts for the CAISO real-time market (dependent variable) using machine learning and quantile regression. Projected ramp forecast errors conditioned on solar uncertainty are translated into flexible ramp requirements that therefore reflect real-time meteorological and solar conditions, improving on typical ISO procedures. Detailed descriptions are provided on the quantile regression and kth-nearest neighbor categorization methods for accomplishing that translation. Finally, a multiple time-scale look-ahead market simulation model is applied to a 118-bus IEEE Reliability Test System, modified to represent the CAISO generation mix and demand distributions. The model runs quantify how solar-conditioned ramp requirements can, first, decrease operating costs by reducing requirements compared to often conservative unconditional methods and, second, decrease generation scarcity events and consequently improve reliability by increasing flexibility requirements at times when unconditional forecast-based requirements understate actual ramp uncertainty. Solar-conditioned ramp requirements are found to reduce generation operating costs by about 2% for the test system (which would be equivalent to over $\$100$ million per year for a CAISO-size system).

14 SOLAR ENERGY↗

Cloud Properties and Boundary Layer Stability Above Southern Ocean Sea Ice and Coastal Antarctica

Significant variability in climate predictions originates from the simulated cloud cover over the Southern Ocean. Historically, Southern Ocean cloud and aerosol properties have been less studied than their northern hemisphere counterparts, and cloud-sea-ice interactions over the Southern Ocean also remain largely unexamined. We used data from combined radar, lidar, radiometer, radiosonde, and ERA5 reanalysis profiles to investigate cloud property relationships to cloud temperature, sea-ice concentration, and boundary layer stability. Our findings show correlations between both cloud macrophysical properties and radiative effects and sea-ice concentration, and that the marine atmospheric boundary layer is more stable over higher sea-ice concentrations. Mixed-phase cloud frequency of occurrence was highest over the sea-ice zone at 15%, three times higher than over cold water south of the Antarctic Polar Front. For temperatures greater than –15°C, low-level, single-layer clouds were more likely to precipitate ice if they were coupled to cold-water or sea-ice surfaces than if they were decoupled from these surfaces, with the highest percentage of clouds precipitating ice observed over sea ice. These findings suggest a surface source of ice-nucleating particles at high southern latitudes that increases cloud glaciation probability. We discuss the implications of our results for future studies into the relationship between cloud properties, aerosols, sea ice, and boundary layer stability at high latitudes over the Southern Ocean.

54 ENVIRONMENTAL SCIENCES↗

Probabilistic and maximum entropy modeling of chemical reaction systems: Characteristics and comparisons to mass action kinetic models

We demonstrate and characterize a first-principles approach to modeling the mass action dynamics of metabolism. Starting from a basic definition of entropy expressed as a multinomial probability density using Boltzmann probabilities with standard chemical potentials, we derive and compare the free energy dissipation and the entropy production rates. We express the relation between entropy production and the chemical master equation for modeling metabolism, which unifies chemical kinetics and chemical thermodynamics. Because prediction uncertainty with respect to parameter variability is frequently a concern with mass action models utilizing rate constants, we compare and contrast the maximum entropy model, which has its own set of rate parameters, to a population of standard mass action models in which the rate constants are randomly chosen. We show that a maximum entropy model is characterized by a high probability of free energy dissipation rate and likewise entropy production rate, relative to other models. We then characterize the variability of the maximum entropy model predictions with respect to uncertainties in parameters (standard free energies of formation) and with respect to ionic strengths typically found in a cell.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗