Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble data assimilation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

CO 2 storage site characterization using ensemble-based approaches with deep generative models

Estimating spatially distributed properties such as permeability from available sparse measurements is a great challenge in efficient subsurface CO 2 storage operations. In this paper, a deep generative model that can accurately capture complex subsurface structure is tested with an ensemble-based inversion method for accurate and accelerated characterization of CO 2 storage sites. We chose Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP) for its realistic reservoir property representation and Ensemble Smoother with Multiple Data Assimilation (ES-MDA) for its robust data fitting and uncertainty quantification capability. WGAN-GP are trained to generate high-dimensional permeability fields from a low-dimensional latent space and ES-MDA then updates the latent variables by assimilating available measurements. Several subsurface site characterization examples including Gaussian, channelized, and fractured reservoirs are used to evaluate the accuracy and computational efficiency of the proposed method and the main features of the unknown permeability fields are characterized accurately with reliable uncertainty quantification. Furthermore, the estimation performance is compared with a widely-used variational, i.e., optimization-based, inversion approach, and the proposed approach outperforms the variational inversion method in several benchmark cases. We explain such superior performance by visualizing the objective function in the latent space: because of nonlinear and aggressive dimension reduction via generative modeling, the objective function surface becomes extremely complex while the ensemble approximation can smooth out the multi-modal surface during the minimization. This suggests that the ensemble-based approach works well over the variational approach when combined with deep generative models at the cost of forward model runs unless convergence-ensuring modifications are implemented in the variational inversion.

42 ENGINEERING↗

Deep Learning for Simultaneous Inference of Hydraulic and Transport Properties

Abstract Identification of a heterogeneous conductivity field and reconstruction of a contaminant release history are key aspects of subsurface remediation. These two goals are achieved by combining model predictions with sparse and noisy hydraulic head and concentration measurements. Solution of this inverse problem is notoriously difficult due to, in part, high dimensionality of the parameter space and high computational cost of repeated forward solves. We use a convolutional adversarial autoencoder (CAAE) to parameterize a heterogeneous non‐Gaussian conductivity field via a low‐dimensional latent representation. A three‐dimensional dense convolutional encoder‐decoder (DenseED) network serves as a forward surrogate of the flow and transport model. The CAAE‐DenseED surrogate is fed into the ensemble smoother with multiple data assimilation (ESMDA) algorithm to sample from the Bayesian posterior distribution of the unknown parameters, forming a CAAE‐DenseED‐ESMDA inversion framework. The resulting CAAE‐DenseED‐ESMDA inversion strategy is used to identify a three‐dimensional contaminant source and conductivity field. A comparison of the inversion results from CAAE‐ESMDA with physical flow and transport simulator and from CAAE‐DenseED‐ESMDA shows that the latter yields accurate reconstruction results at the fraction of the computational cost of the former.

Zhou, Zitong↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS↗

Nonlinear Ensemble Filtering with Diffusion Models: Application to the Surface Quasigeostrophic Dynamics

The intersection between classical data assimilation methods and novel machine learning techniques has attracted significant interest in recent years. Here, we explore another promising solution in which diffusion models are used to formulate a robust nonlinear ensemble filter for sequential data assimilation. Unlike standard machine learning methods, the proposed ensemble score filter (EnSF) is completely training free and can efficiently generate a set of analysis ensemble members. Here, in this study, we apply the EnSF to a surface quasigeostrophic model and compare its performance against the popular local ensemble transform Kalman filter (LETKF), which makes Gaussian assumptions in the analysis step. Numerical tests demonstrate that EnSF maintains stable performance in the absence of localization and for a variety of experimental settings. We find that while LETKF maintains optimal performance in the case of linear observations of the entire state and a perfect model, EnSF shows improvements over LETKF when nonlinear observations are assimilated and the system is subject to unexpected model errors. A spectral decomposition of the analysis results in this nonlinear observation regime shows that the largest improvements over LETKF occur at large scales (small wavenumbers), where LETKF lacks sufficient ensemble spread. Overall, this initial application of EnSF to a geophysical model of intermediate complexity motivates further development of the algorithm for more realistic problems.

Artificial intelligence↗

Transfer learning of neural surrogates on multifidelity groundwater simulations

Multifidelity data used in the paper published in Advances in Water Resources 206 (2025) 105140, https://doi.org/10.1016/j.advwatres.2025.105140 The code used to process the data is openly available on GitHub at https://github.com/Model-Reduction-and-UQ-Group/Transfer_Learning_K_reconstruction Computationally inexpensive surrogates of process-based models, such as deep neural networks, enable ensemble-based computations used in risk assessment, data assimilation, etc. However, generation of large datasets required to train a neural network can be as expensive as the ensemble simulations themselves. We ameliorate this challenge by using data from multifidelity (MF) groundwater simulations and transfer learning (TL) to reduce data generation costs while maintaining model accuracy. As a computational example, we train a deep convolutional neural network (CNN) to reconstruct permeability fields from saturation maps derived from a multiphase flow model. Starting with very low- and low-fidelity data generated on increasingly coarse meshes, we pretrain the CNN, followed by output-layer training and fine-tuning using only a limited number of high-fidelity samples. We demonstrate the surrogate’s robustness when interpreting low-quality inputs—such as interpolated maps or data affected by noise—which has strong implications for the applicability in practical hydrogeological scenarios. This multilevel MF-TL strategy achieves a favorable trade-off between computational efficiency and predictive accuracy, significantly outperforming high-fidelity-only approaches under the same computational budget.

Chiofalo, Alessia [University of Bologna] (ORCID:0↗

The Impact of Constrained Data Assimilation on the Forecasts of Three Convection Systems During the ARM MC3E Field Campaign

A constrained data assimilation (CDA) system based on the ensemble variational (EnVar) method and physical constraints of mass and water conservations is evaluated through three convective cases during the Midlatitude Continental Convective Clouds Experiment (MC3E) of the Atmospheric Radiation Measurement (ARM) program. Compared to the original data assimilation (ODA), the CDA is shown to perform better in the forecasted state variables and simulated precipitation. The CDA is also shown to greatly mitigate the loss of forecast skills in observation denial experiments when radar radial winds are withheld in the assimilation. Modifications to the algorithm and sensitivities of the CDA to the calculation of the time tendencies in the constraints are described.

54 ENVIRONMENTAL SCIENCES↗

Multifidelity Ensemble Kalman Filtering Using Surrogate Models Defined by Theory-Guided Autoencoders

Data assimilation is a Bayesian inference process that obtains an enhanced understanding of a physical system of interest by fusing information from an inexact physics-based model, and from noisy sparse observations of reality. The multifidelity ensemble Kalman filter (MFEnKF) recently developed by the authors combines a full-order physical model and a hierarchy of reduced order surrogate models in order to increase the computational efficiency of data assimilation. The standard MFEnKF uses linear couplings between models, and is statistically optimal in case of Gaussian probability densities. This work extends the MFEnKF into to make use of a broader class of surrogate model such as those based on machine learning methods such as autoencoders non-linear couplings in between the model hierarchies. We identify the right-invertibility property for autoencoders as being a key predictor of success in the forecasting power of autoencoder-based reduced order models. We propose a methodology that allows us to construct reduced order surrogate models that are more accurate than the ones obtained via conventional linear methods. Numerical experiments with the canonical Lorenz'96 model illustrate that nonlinear surrogates perform better than linear projection-based ones in the context of multifidelity ensemble Kalman filtering. We additionality show a large-scale proof-of-concept result with the quasi-geostrophic equations, showing the competitiveness of the method with a traditional reduced order model-based MFEnKF.

97 MATHEMATICS AND COMPUTING↗

The 4DEnVar-based weakly coupled land data assimilation system for E3SM version 2

Abstract. A new weakly coupled land data assimilation (WCLDA) system based on the four-dimensional ensemble variational (4DEnVar) method is developed and applied to the fully coupled Energy Exascale Earth System Model version 2 (E3SMv2). The dimension-reduced projection four-dimensional variational (DRP-4DVar) method is employed to implement 4DVar using the ensemble technique instead of the adjoint technique. With an interest in providing initial conditions for decadal climate predictions, monthly mean anomalies of soil moisture and temperature from the Global Land Data Assimilation System (GLDAS) reanalysis from 1980 to 2016 are assimilated into the land component of E3SMv2 within the coupled modeling framework with a 1-month assimilation window. The coupled assimilation experiment is evaluated using multiple metrics, including the cost function, assimilation efficiency index, correlation, root-mean-square error (RMSE), and bias, and compared with a control simulation without land data assimilation. The WCLDA system yields improved simulation of soil moisture and temperature compared with the control simulation, with improvements found throughout the soil layers and in many regions of the global land. In terms of both soil moisture and temperature, the assimilation experiment outperforms the control simulation with reduced RMSE and higher temporal correlation in many regions, especially in South America, central Africa, Australia, and large parts of Eurasia. Furthermore, significant improvements are also found in reproducing the time evolution of the 2012 US Midwest drought, highlighting the crucial role of land surface in drought lifecycle. The WCLDA system is intended to be a foundational resource for research to investigate land-derived climate predictability.

58 GEOSCIENCES↗

Data assimilation for combustion ignition delay time simulation using schlieren image velocimetry

This study sought to improve the accuracy of simulating spray penetration and combustion ignition delay by means of data assimilation (DA). The simulations were conducted using the Reynolds-averaged Navier–Stokes (RANS) equations and assimilating the schlieren image data. In DA, an ensemble square root filter (EnSRF) was used to build the statistical model, making the simulation results more accurate without any change in the governing equations. Recognizing that the spray-cone injection angle has a large effect on penetration, we created ensemble members with different injection angles. And we applied the two-component velocity distribution calculated via SIV and updated both velocity and temperature by using a DA statistical model derived from RANS ensemble simulations. The ignition delay time is generally known to vary even under the same experimental conditions because it is influenced by many factors. In this study, we attempted the transient DA-assisted RANS simulation to predict the ignition delay time even when the temporal resolution and accuracy of the observation data ware insufficient. Our trials offer an example of how a combination of techniques can be effectively used to assimilate experimental data obtained under restricted conditions.

Combustion simulation↗

Robustness of the Ensemble Score Filter to the Type of Assimilated Observation Networks

Recent advances in data assimilation (DA) have focused on developing more flexible approaches that can better accommodate nonlinearities in models and observations. However, it remains unclear how the performance of these advanced methods depends on the observation network characteristics. In this study, we present initial experiments with the surface quasi‐geostrophic model, in which we compare a recently developed ensemble filter using score‐based diffusion models with the standard Local Ensemble Transform Kalman Filter (LETKF). Our results show that the analysis solutions respond differently to the number, spatial distribution, and nonlinear fraction of assimilated observations. We also find notable changes in the multiscale characteristics of the analysis errors. Given that standard DA techniques will eventually be replaced by more advanced methods, we hope this study sets the ground for future efforts to reassess the value of Earth observing systems in the context of newly emerging algorithms.

97 MATHEMATICS AND COMPUTING↗

A Scalable Real-Time Data Assimilation Framework for Predicting Turbulent Atmosphere Dynamics

AI-based foundation models like FourCastNet, GraphCast are revolutionizing weather and climate predictions but are not yet ready for operational use. Their limitation lies in the absence of a data assimilation system to incorporate real-time Earth system observations, crucial for accurately forecasting events like tropical cyclones. To overcome these obstacles, we introduce a generic real-time data assimilation framework and demonstrate its end-to-end performance on the Frontier supercomputer. This framework comprises two primary modules: an ensemble score filter (EnSF), which significantly outperforms the state-of-the-art data assimilation method, and a vision transformer-based surrogate capable of real-time adaptation through the integration of observational data. We demonstrate both the strong and weak scaling of our framework up to 1024 GPUs on the Exascale supercomputer, Frontier. Our results not only illustrate the framework's exceptional scalability on high-performance computing systems, but also demonstrate the importance of supercomputers in real-time data assimilation for weather and climate predictions.

Lu, Dan↗

Simulated wind speed and initial conditions over the WFIP2 region: Cold-front (D01)

The purpose of this work is to assess the sensitivity of the forecast for turbine height wind speed to initial condition (IC) uncertainties over the Columbia River Gorge and Columbia River Basin for two typical weather phenomena: a local thermal gradient induced by a marine air intrusion and passage of a cold front. The Weather Research and Forecasting (WRF) model data assimilation system (WRFDA) was used to generate ensemble ICs from the North American Regional Analysis (NARR) for the WRF model initialization. The simulated turbine-height wind speeds were categorized into four types using the self-organizing map (SOM) technique. This work advances understanding of IC uncertainties impacts on wind speed forecasts and locates the high-impact regions.

17 WIND ENERGY↗

Simulated wind speed and initial conditions over the WFIP2 region: Cold-front (D02)

The purpose of this work is to assess the sensitivity of the forecast for turbine height wind speed to initial condition (IC) uncertainties over the Columbia River Gorge and Columbia River Basin for two typical weather phenomena: a local thermal gradient induced by a marine air intrusion and passage of a cold front. The Weather Research and Forecasting (WRF) model data assimilation system (WRFDA) was used to generate ensemble ICs from the North American Regional Analysis (NARR) for the WRF model initialization. The simulated turbine-height wind speeds were categorized into four types using the self-organizing map (SOM) technique. This work advances understanding of IC uncertainties impacts on wind speed forecasts and locates the high-impact regions.

17 WIND ENERGY↗

Simulated wind speed and initial conditions over the WFIP2 region: Sea-breeze (D01)

The purpose of this work is to assess the sensitivity of the forecast for turbine height wind speed to initial condition (IC) uncertainties over the Columbia River Gorge and Columbia River Basin for two typical weather phenomena: a local thermal gradient induced by a marine air intrusion and passage of a cold front. The Weather Research and Forecasting (WRF) model data assimilation system (WRFDA) was used to generate ensemble ICs from the North American Regional Analysis (NARR) for the WRF model initialization. The simulated turbine-height wind speeds were categorized into four types using the self-organizing map (SOM) technique. This work advances understanding of IC uncertainties impacts on wind speed forecasts and locates the high-impact regions.

17 WIND ENERGY↗

Simulated wind speed and initial conditions over the WFIP2 region: Sea-breeze (D02)

The purpose of this work is to assess the sensitivity of the forecast for turbine height wind speed to initial condition (IC) uncertainties over the Columbia River Gorge and Columbia River Basin for two typical weather phenomena: a local thermal gradient induced by a marine air intrusion and passage of a cold front. The Weather Research and Forecasting (WRF) model data assimilation system (WRFDA) was used to generate ensemble ICs from the North American Regional Analysis (NARR) for the WRF model initialization. The simulated turbine-height wind speeds were categorized into four types using the self-organizing map (SOM) technique. This work advances understanding of IC uncertainties impacts on wind speed forecasts and locates the high-impact regions.

17 WIND ENERGY↗

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat↗

An ensemble of 48 physically perturbed model estimates of the 1/8° terrestrial water budget over the conterminous United States, 1980–2015

Terrestrial water budget (TWB) data over large domains are of high interest for various hydrological applications. Spatiotemporally continuous and physically consistent estimations of TWB rely on land surface models (LSMs). As an augmentation of the operational North American Land Data Assimilation System Phase 2 (NLDAS-2) four-LSM ensemble, this paper describes a dataset simulated from an ensemble of 48 physics configurations of the Noah LSM with multi-physics options (Noah-MP). The 48 Noah-MP physics configurations are selected to give a representative cross-section of commonly used LSMs for parameterizing runoff, atmospheric surface layer turbulence, soil moisture limitation on photosynthesis, and stomatal conductance. The dataset spans from 1980 to 2015 over the conterminous United States (CONUS) at a monthly temporal resolution and a 1/8° spatial resolution. The dataset variables include total evapotranspiration and its constituents (canopy evaporation, soil evaporation, and transpiration), runoff (the surface and subsurface components), as well as terrestrial water storage (snow water equivalent, four-layer soil water content from the surface down to 2 m, and the groundwater storage anomaly). The dataset is available at https://doi.org/10.5281/zenodo.7109816. Evaluations carried out in this study and previous investigations show that the ensemble per forms well in reproducing the observed terrestrial water storage, snow water equivalent, soil moisture, and runoff. Noah-MP complements the NLDAS models well, and adding Noah-MP consistently improves the NLDAS es timations of the above variables in most areas of CONUS. Besides, the perturbed-physics ensemble facilitates the identification of model deficiencies. The parameterizations of shallow snow, spatially varying groundwater dynamics, and near-surface atmospheric turbulence should be improved in future model versions.

54 ENVIRONMENTAL SCIENCES↗

A Novel Modeling Framework for Computationally Efficient and Accurate Real-Time Ensemble Flood Forecasting With Uncertainty Quantification

A novel modeling framework that simultaneously improves accuracy, predictability, and computational efficiency is presented. It embraces the benefits of three modeling techniques integrated together for the first time: surrogate modeling, parameter inference, and data assimilation. The use of polynomial chaos expansion (PCE) surrogates significantly decreases computational time. Parameter inference allows for model faster convergence, reduced uncertainty, and superior accuracy of simulated results. Ensemble Kalman filters assimilate errors that occur during forecasting. To examine the applicability and effectiveness of the integrated framework, we developed 18 approaches according to how surrogate models are constructed, what type of parameter distributions are used as model inputs, and whether model parameters are updated during the data assimilation procedure. We conclude that (1) PCE must be built over various forcing and flow conditions, and in contrast to previous studies, it does not need to be rebuilt at each time step; (2) model parameter specification that relies on constrained, posterior information of parameters (so-called Selected specification) can significantly improve forecasting performance and reduce uncertainty bounds compared to Random specification using prior information of parameters; and (3) no substantial differences in results exist between single and dual ensemble Kalman filters, but the latter better simulates flood peaks. The use of PCE effectively compensates for the computational load added by the parameter inference and data assimilation (up to ~80 times faster). Therefore, the presented approach contributes to a shift in modeling paradigm arguing that complex, high-fidelity hydrologic and hydraulic models should be increasingly adopted for real-time and ensemble flood forecasting.

54 ENVIRONMENTAL SCIENCES↗