Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Ensemble-based methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

27 records · Page 2

A Generalized, Compactly-Supported Correlation Function for Data Assimilation Applications

Correlation functions play an essential role in modern data assimilation, where they are used to model covariances given a set of tunable parameters or applied as tapering functions to localize covariances in ensemble-based schemes. One of the most widely-used correlation functions in data assimilation is the Gaspari and Cohn (1999) piecewise-rational, compactly-supported parametric correlation function (hereafter referred to as GC99). The GC99 correlation function is useful due to its tunable cut-off parameter c and Gaussian-like shape achieved when the parameter a is set to one-half. These properties are attractive for tapering functions in data assimilation applications. However, the GC99 correlation function is homogeneous over Euclidean 3-space and isotropic when restricted to the sphere, properties that may be less than ideal for some geophysical applications. GC99 is also compactly-supported on a sphere of fixed radius, which requires tuning of the cut-off parameter c that can depend on the specific application. This work presents a generalization of the GC99 correlation function that allows the cut-off parameter c and shape parameter a to vary over space to gain more flexibility in shape while maintaining its compact support property. The function, which we call the Generalized Gaspari Cohn (GenGC) correlation function, introduces inhomogeneity in Euclidean 3-space and anisotropy when restricted to the sphere by allowing both parameters c and a to vary, as functions, over the spatial domain. The GC99 correlation function is a special case of GenGC where the functions c and a are held constant, as fixed parameters rather than functions. The GenGC correlation function also generalizes the follow-on to the work of Gaspari and Cohn (1999) presented in Gaspari et al. (2006), which allowed a to vary while keeping c fixed. We illustrate through simple one- and two-dimensional examples the variety of inhomogeneous and anisotropic correlation functions GenGC can produce by varying c and a over space, and suggest applications where they may be useful in data assimilation, such as covariance modeling or localization. In particular, we describe how the GenGC correlation function can be used to construct covariances using correlation length and variance fields derived from dynamics. For example, the correlation length field for advective dynamics is governed by a partial differential equation (PDE) in N spatial dimensions, where N is the number of space dimensions of the state. Correlation length fields can be determined from this PDE and used with GenGC to construct the corresponding correlations. We can then approximate the full covariance by rescaling by the variance, which also satisfies a PDE in N spatial dimensions for advective dynamics. Thus we can approximate the full covariance without solving the covariance PDE, which is in 2N spatial dimensions, by solving just two PDEs each in only N spatial dimensions. This approach to evolving the correlation length and variance fields, then reconstructing the correlations using GenGC, is suggested as an alternative to current methods of covariance modeling in data assimilation algorithms.

GC99↗

DART-PFLOTRAN: An ensemble-based data assimilation system for estimating subsurface flow and transport model parameters

Ensemble-based Data Assimilation (EDA), based on the Monte Carlo approach, has been effectively applied to estimate model parameters through inverse modeling in subsurface flow and transport problems. However, implementation of EDA approach involves a complicated workflow that include setting up and executing ensemble forward model simulations, processing observations and model simulation results for parameter updates, and repeat for sequential or iterative EDA. To facilitate the management of such workflow and lower the barriers for adopting EDA-based parameter estimation in subsurface science, we develop a generic software frame-work linking the Data Assimilation Research Testbed (DART) with a massively parallel subsurface FLOw and TRANsport code PFLOTRAN. The new DART-PFLOTRAN leverages both the core data assimilation engines in DART and the computational power afforded by PFLOTRAN. In addition to the standard smoother and filtering options, DART-PFLOTRAN enables an iterative EDA workflow based on the Ensemble Smoother for Multiple Data Assimilation method (ES-MDA) to improve estimation accuracy for nonlinear forward problems. Here, we verify the implementation of ES-MDA in DART-PFLOTRAN using two synthetic cases designed to estimate static permeability and dynamic exchange fluxes across the riverbed, respectively, from continuous temperature measurements made across a depth profile. One-dimensional hydro-thermal simulations are performed in both cases to relate temperature responses with the parameters of interest. In the case of estimating dynamic parameters, we demonstrate the flexibility of DART-PFLOTRAN in automating sequential ES-MDA workflow, which will significantly reduce the time researchers spend on managing complex workflows in similar applications. Both studies yield accurate estimations of the parameters compared to their synthetic truth, while ES-MDA leads to more accurate estimation when a high level of nonlinearity exist between observed responses and unknown parameters. With a code base in Python and Fortran, DART-PFLOTRAN paves the way for applications in large-scale subsurface inverse modeling by automating the complex workflow of sequential ES-MDA that can be executed on various computing platforms.

97 MATHEMATICS AND COMPUTING↗

QUIC-URB and QUIC-fire extension to complex terrain: Development of a terrain-following coordinate system

Ensemble-based approaches to prescribed fire planning cannot be supported by CFD-based models like FIRETEC and WFDS because they are too computationally expensive and cannot leverage LES approaches like CAWFE and WRF-SFIRE because too coarse of resolution. QUIC-Fire was developed to fill this gap but it cannot currently address complex terrain, typical for instance of the Western United States. In this paper, we describe the extension of the diagnostic wind model QUIC-URB, the wind engine of QUIC-Fire, to a terrain-following coordinate system. In particular, the paper presents the mathematical derivation of the wind solver leading to a linear system of equations that are solved through the successive over-relaxation method. The model is validated against a standard test used in previous works (the Askervein Hill) and against a new dataset from measurements in the Socorro Mountains, New Mexico. The terrain-following implementation captured the correct phenomenology for the isolated Askervein Hill, with a wind speed up at the top of the hill. We report the model agreed well with measurements on the upwind side of the peak, but overestimated speed-up on the downwind side of the hill. This is due to the inability of the model to generate flow separation and wake-eddy dynamics. On a common laptop, the divergence-free wind field was obtained in 6 s, making the solver appealing for coupled fire–atmosphere simulations. The Socorro Mountain was highly complex, with many cliff faces, peaks, and valleys. Although the model captures the magnitude and direction of inlet and outlet areas of the domain, it performs rather poorly in the valley region and in the regions near the steep cliffs. Hence, the model shows good agreement with data in areas of open sloped terrain but lacks in areas where flow separation and thermally driven effects may be present (neither effect was addressed in this work). Results highlight that future work should focus on the implementation of parameterizations of wake-eddies, similar to QUIC-URB’s building parameterizations, and on thermodynamic-driven flow.

54 ENVIRONMENTAL SCIENCES↗

High-Performance Semiempirical Excited-State Molecular Dynamics Powered by Graphics Processing Units

Here, this Letter introduces excited-state molecular dynamics in PYSEQM, a GPU-accelerated semiempirical quantum chemistry engine implemented in PyTorch. The new module enables Born–Oppenheimer molecular dynamics (BOMD) using configuration-interaction singles and random phase approximation for excited states, allowing long trajectories and large statistical ensembles to be simulated efficiently on a single GPU. We also implement an extended Lagrangian excited-state BOMD (XL-ESMD) scheme that propagates auxiliary electronic variables, enabling relaxed ground and excited-state convergence thresholds without compromising energy conservation. The excited-state BOMD implementation scales smoothly from small chromophores to a nearly 900-atom dendrimer (taking 6.5 s per MD step). PYSEQM also supports batched execution, allowing many geometries or trajectories to be evaluated in a single GPU launch, substantially increasing throughput and making ensemble-based protocols routine. As a demonstration, we compute absorption, emission, and infrared spectra from trajectories propagated on the ground and first excited states. The XL-ESMD scheme yields identical spectra at significantly lower computational cost, establishing the role of extended Lagrangian based dynamics for efficient excited-state BOMD simulations. Beyond raw performance, PYSEQM’s PyTorch foundation provides automatic differentiation for forces, efficient GPU batching, and seamless interfacing with machine learning models. These capabilities position PYSEQM as a practical platform for machine learning-augmented excited-state dynamics and lay the foundation for future data-driven nonadiabatic excited-state dynamics modeling of ultrafast spectroscopic probes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Earth System Reanalysis in Support of Climate Model Improvements

Recent climate model developments, established through increased model resolution, have led to substantial improvements in model simulations of the time-evolving, coupled Earth system and its subcomponents. However, regardless of resolution, climate models will always produce climate features and variability that differ from the real world and will be prone to biases. This is due to many remaining uncertainties, such as in parametric and structural model uncertainty, in the initial conditions prescribed, and in the prescribed (scenario) forcing which varies on decadal to centennial timescales. Further model improvements are expected to arise specifically from improved representation of physical processes realized through model-data fusion. This will create an unprecedented opportunity to better exploit a large array of Earth observations, from in situ measurements to weather radars and satellite observations, as the resolved scales of the models approach those of the observations. For this, climate DA will be the central tool to bring models and observations into consistency, by improving initial conditions, inferring uncertain model parameters and structure, and quantifying uncertainty. Generally, there will be advantages and complementarities of adjoint-based smoother approaches, ensemble-based filter approaches, or new ML-inspired approaches. Yet, the ever-increasing model resolution will present growing challenges arising from computational cost, calling for new ways of performing data assimilation and model optimization. Using the complementarity in a hybrid approach, blending tools and concepts from variational, ensemble and ML methods might be what is required in the future. In this context ML could be important to handle non-linear responses, and to better approximate non-Gaussian distributions.

54 ENVIRONMENTAL SCIENCES↗

Terrain-Influenced Winds and Fire-Fire Interactions in Wildland Fire Simulations [Dissertation]

Ensemble-based approaches to prescribed fire planning cannot be supported by computational fluid dynamics based models like FIRETEC and the Wildland-Urban Interface Fire Dynamics Simulator (WFDS) because they are too computationally expensive and cannot leverage large eddy simulation approaches like CAWFE and WRF-SFIRE because they have too coarse of resolution. QUIC-Fire was developed to fill this gap but it cannot currently address complex terrain, that is typical for instance in the Western United States. This dissertation describes a variety of improvements made to QUIC-Fire and its various incorporated algorithms in an effort to make it a viable tool in simulating wildland and prescribed fires on terrain. The modifications made to QUIC-Fire are described in three chapters. The first chapter describes the extension of the diagnostic wind model QUIC-URB, the wind engine of QUIC-Fire, to a terrain-following coordinate system. The terraininfluenced winds it generates are analyzed and compared. In particular, this chapter presents the mathematical derivation of the wind solver leading to a linear system of equations that are solved through the successive over-relaxation method. The model is validated against a standard test used in previous works (the Askervein Hill) and against a new dataset from measurements in the Socorro Mountains, New Mexico. The terrain-following implementation captures the correct phenomenology for the isolated Askervein Hill, with a wind speed up at the top of the hill. The model agrees well with measurements on the upwind side of the peak, but overestimates speed-up on the downwind side of the hill. This is due to the inability of the model to generate flow separation and wake-eddy dynamics. On a common laptop, the divergence-free wind field is obtained in 6 s, making the solver appealing for coupled fire-atmosphere simulations. The Socorro Mountain is highly complex, with many cliff faces, peaks, and valleys. Although the model captures the magnitude and direction of inlet and outlet areas of the domain, it performs rather poorly in the valley region and in the regions near the steep cliffs. Hence, the model shows good agreement with data in areas of open sloped terrain but lacks in areas where flow separation and thermally driven effects may be present (neither effect is addressed in this work). In the second chapter the implementation of the terrain-following version of QUIC-URB into QUIC-Fire, and the necessary changes needed to include terrain are described. No changes to the underlying fire spread algorithm are made other than what is required to correctly account for the inclusion of terrain. Previously published FIRETEC results that use five different topographies that share the same centerline profile are compared to simulation results from the modified QUIC-Fire that use the same topographies and fuels. QUIC-Fire results show overall similar behaviors in terms of how the topographies affect fire shapes and trends in spread rates. Due to the terrain-following version of QUIC-URB being unable to generate flow separations at the crest of hills, fire spread rates in these regions across all topographies are over-predicted when compared to FIRETEC. Lateral fire growth shows similar trends with FIRETEC between topographies but does not capture the increase in spread due to a diagonal interface between grassland and forested fuel region of the domain. These results suggest that there are three algorithms within QUIC-Fire that could use improvement: how flame tilt angle is accounted for, the incorporation of non-local drag effects, and the inclusion of the wake-eddy parameterizations that are used in QUIC-URB. Lastly, the third chapter describes a modification to the initial guess used for the calculation of the QUIC-URB mass-conserved wind solution during fire simulations. The modification is aimed at improving fire-fire interactions in QUIC-Fire simulations. The modification consists of using the solution from the previous timestep as the starting point for the calculation of the solution for the next timestep. Fire-fire interactions is greatly improved by the change but a new source of error is introduced. Due to how plumes are modelled in QUIC-Fire the new solution contains errors where gaps in the plume structure are present. However, these errors are mostly limited to the upper atmosphere, where they do not affect fire behavior at the surface, and their magnitude isn’t significant enough to discount the amount of new fire phenomenology now captured in QUIC-Fire with the change.

58 GEOSCIENCES↗

Accurate and Rapid Forecasts for Geologic Carbon Storage via Learning-Based Inversion-Free Prediction

Carbon capture and storage (CCS) is one approach being studied by the U.S. Department of Energy to help mitigate global warming. The process involves capturing CO 2 emissions from industrial sources and permanently storing them in deep geologic formations (storage reservoirs). However, CCS projects generally target “green field sites,” where there is often little characterization data and therefore large uncertainty about the petrophysical properties and other geologic attributes of the storage reservoir. Consequently, ensemble-based approaches are often used to forecast multiple realizations prior to CO 2 injection to visualize a range of potential outcomes. In addition, monitoring data during injection operations are used to update the pre-injection forecasts and thereby improve agreement between forecasted and observed behavior. Thus, a system for generating accurate, timely forecasts of pressure buildup and CO 2 movement and distribution within the storage reservoir and for updating those forecasts via monitoring measurements becomes crucial. This study proposes a learning-based prediction method that can accurately and rapidly forecast spatial distribution of CO 2 concentration and pressure with uncertainty quantification without relying on traditional inverse modeling. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO 2 storage site operators with an effective tool for timely and informative decision making based on limited simulation and monitoring data.

58 GEOSCIENCES↗

Spaceborne Lidar Retrievals of PM2.5 for Air Quality Studies and Applications

Fine particulate matter (PM2.5) substantially contributes to air pollution and negatively affects human health. While many studies have investigated the use of passive column-integrated aerosol optical depth to infer surface PM2.5, the use of lidar observations for air quality characterization is not nearly as extensive. Lidar measurements are critical, however, due to the vertical aerosol information they provide, including near the surface. In this presentation, we first provide an overview of various lidar-based approaches for estimating PM2.5 concentrations and then discuss how lidar measurements can assist other air quality applications. For example, estimates of PM2.5 have been obtained in a physics-based approach through CALIOP near-surface aerosol extinction retrievals, assumptions on the mass extinction efficiency, and incorporating other parameters (an aerosol hygroscopic growth factor and PM2.5/PM10 ratio). Application of this algorithm over the contiguous United States (CONUS) from 2006 to 2018 yielded larger PM2.5 values over the eastern and western CONUS (~10-15 μg/m³) and lower PM2.5 levels in the central CONUS (~5 μg/m³). These spatial patterns were similar to those from gridded PM2.5 concentrations obtained through in situ measurements at ground stations operated by the US Environmental Protection Agency. In another approach, the Cloud Aerosol Transport System (CATS) lidar was used with the Goddard Earth Observing System (GEOS) model in a 1D ensemble-based variational technique to obtain PM2.5 over the US and Europe, and the spatial patterns of the CATS/GEOS based PM2.5 concentrations generally captured those from surface stations (with corresponding hourly EPA PM2.5 vs CATS PM2.5 statistics of R=0.4 and bias=1.5 μg/m³). In our recent work, as part of the Models, In situ, and Remote sensing of Aerosols (MIRA) Working Group, we have applied both the CALIOP and CATS/GEOS based approaches over the highly polluted country of India during the post-monsoon season (September-October 2016). We derived elevated levels of two-month mean PM2.5 (~100 μg/m³) in northern India, especially near New Delhi. These high PM2.5 concentrations in the Indo-Gangetic plain are driven in large part from the seasonal burning of crop residue and meteorological conditions typical at this time of the year, such as low wind speeds and a shallow boundary layer. While the satellite-derived PM2.5 moderately replicates (R = ~0.7-0.9) the spatial variability in the two-month mean of surface in situ PM2.5 from monitoring sites operated by the Central and State Pollution Control Boards, we show results from specific scenes for which there are large deviations between the satellite-derived PM2.5 and in situ measurements. Other current work on this topic focuses on developing PM2.5 estimates using airborne high spectral resolution lidar measurements through machine learning regression algorithms and involves several parameters (e.g., aerosol extinction, color ratio, lidar ratio). Application of this method over major metropolitan areas in the US and Asia have resulted in high correlations (R = 0.93) with surface measurements. This airborne lidar approach can be adapted to spaceborne lidar measurements, and all three of these approaches can be applied to ESA’s EarthCARE Atmospheric Lidar instrument, setting the stage for the future Cloud Aerosol Lidar for Global Scale Observations of the Ocean-Land Atmosphere System (CALIGOLA) mission. Ultimately, beyond estimates of PM2.5, the aerosol vertical distribution from lidars can benefit studies involving passive sensor approaches for PM2.5 proxies (including from geostationary satellites), wildfire smoke plume injection heights, volcanic emissions (e.g., ash height retrievals), and aerosol/air quality model assimilation, evaluation, and forecasts.

Travis D Toth↗

Visual HPC Workflows for the Analysis of System Dynamics Models

Visual analytics supported by high performance computing (HPC) accelerates and enhances the discovery, exploration, and analysis of causal patterns in complex system dynamics (SD) models. We present a suite of visualization-assisted ensemble-based techniques for hypothesis generation and testing, and for sensitivity analysis. By employing HPC to provide parallel, on-demand simulation of SD models, one can “steer” an ensemble of simulated scenarios in real time as one first formulates and then informally tests those hypotheses: this provides rapid feedback for analysts to refine their understanding of the causal relationships emergent from a model. Such understandings can be followed and augmented by rigorous application of statistical methods, namely global variance-based sensitivity analysis, Monte-Carlo filtering, adaptive regional sensitivity analysis, and self-organized maps: here timely computation relies on HPC, while effective presentation emphasizes high-dimensional multivariate data visualization. Immersive visualization in virtual 3D environments provides an excellent adjunct to the traditional 2D graphics typically used for SD models, as it generates an embodied understanding of model behavior and facilitates an active, collaborative critique of model structure and output. Finally, we summarize prospects for HPC-enabled visual analytics applied to SD modeling.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗