Scalable, Trustworthy, Agile, Goal-oriented, Efficient, and Data-assimilated Data-Driven Closure Models (STAGED DDCMs)
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Presentation video for ML/DL workshop
The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. When comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that GLLS predictions fail to replicate computed response distributions for nonlinear applications, while MOCABA shows near agreement, and IUQ uses computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.
The dynamics of shattered pellet injection (SPI) shutdowns are simulated using a time-dependent global energy balance model, based on a modification of the KPRAD framework. The new SPI particle source in the model calculates the ablation of individual pellet fragments that enter the plasma as a temporally resolved plume, thus capturing the effects of earlier fragments on the ablation of those that follow, which has a significant impact on the overall assimilation. Despite the reduced physics and the global averaging of all quantities, results from a large number of DIII-D, KSTAR, and JET experiments are well reproduced, including the plasma cooling timescales, particle assimilation, and current quench (CQ) rates. Cooling timescales and CQ rates are in good agreement for pellets containing as little as ∼1% neon, while particle assimilations are most accurate for neon fractions above ∼15% by number of atoms. Below this, the assimilation tends to be overestimated due to the lack of radial transport in the particle balance, which becomes important in the low-Z limit. Predictive simulations of mixed-composition dual-SPI shutdowns in ITER are compared against those with the 3D non-linear magnetohydrodynamic code JOREK and are found to reproduce overall trends observed in the higher-fidelity modeling across a range of injection scenarios. The general success of the model points to the critical role of energy balance in determining SPI particle assimilation and the subsequent disruption dynamics and highlights the value of these simulations for experimental interpretation and for optimizing the deployment of computationally expensive, higher-fidelity models.
Lignin is one of the most common biopolymers on Earth. In nature, lignin is primarily deconstructed by fungi into mixtures of aromatic compounds that are then assimilated by bacteria and fungi. Industrially, lignin is primarily generated as a byproduct of pulp and paper production and burned for process heat. However, if the appropriate assimilatory pathways were identified, deconstructed lignin could be funneled into value-added products using engineered bacteria. Foundational work has described pathways for assimilation of diverse monomeric aromatic compounds such as protocatechuate, ferulate, and syringate, as well as select dimers including those with β-O-4 and 5-5 interunit linkages. Recent advances have elucidated additional pathways for dimer assimilation, including pathways for new substrates as well as parallel pathways for previously characterized substrates. Comparing these dimer assimilation pathways can illuminate the underlying biochemical logic of assimilation for lignin-associated aromatic dimers and provide opportunities for metabolic engineering to enhance lignin valorization.
This dataset presents a high-resolution historical streamflow reanalysis for NHDPlusV2 stream reaches across Puerto Rico (PR) spanning 1950 - 2019. The reanalysis is generated using the calibrated VIC-RAPID hydrologic modeling framework at the Hydrologic Unit Code Sub-basin (HUC08) scale, forced with sub-daily and daily meteorological forcings from Daymet. Runoff is simulated on 1- and 6-km grids, and the resulting total runoff is routed through the NHDPlusV2 river network using the RAPID routing model to produce Naturalized Streamflow Reanalysis. Where complete observational records are available over 1980 - 2019, streamflows are assimilated (substituted) and subsequently routed downstream through the river network to produce Assimilated Streamflow Reanalysis. The dataset includes streamflow outputs from eight distinct hydrologic modeling configurations along with key performanc evaluation metrics at daily and monthly scales, supporting a wide range of water resource applications. This dataset is derived to support the Non-Powered Dam Assessment, as well as 9505 Secure Water Assessment projects for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). For further details, refer to Ghimire et al. (2023), Kao et al. (2024), and Ghimire et al. (2025).
This manuscript proposes a novel information-theoretic approach to the quantification of experimental relevance, i.e., coverage, to achieve optimal data assimilation results for nuclear engineering applications. Specifically, this work posits the need for a new metric, called coverage (q C ) of an application’s quantity of interest, i.e., eigenvalue or power peaking for an advanced reactor concept, defined herein as the theoretically maximum achievable reduction in the quantity’s uncertainty given measurements from a pool of experiments in a manner that is independent of the data assimilation procedure employed. Currently, reduction in a quantity’s uncertainty is strongly biased by the underlying assumptions of the assimilation procedure to account for the under-determined nature of such problems and the similarity criterion employed to identify relevant experiments. To address this challenge, this work has developed a coverage metric, q C , based on mutual information, which establishes a new conceptual framework for assessing coverage, one that is independent of the model parameters and responses degree of variations in both the experimental and application domains, i.e., linear vs non-linear, and their prior uncertainty distributions, i.e., Gaussian vs. non-Gaussian. The q C is an entropic measure capable of addressing coverage for general nonlinear problems with non-Gaussian uncertainties and inclusive of the measurement uncertainties from multiple experiments. Numerical experiments from manufactured analytical problems as well as a set of benchmarks from the ICSBEP handbook are employed to demonstrate its theoretical and practical performance as compared to the c k -based experiment selection methodology, commonly employed in the neutronic community. The manuscript then employs other well-known adaptations to existing data assimilation methodologies for nonlinear and non-Gaussian problems capable of achieving the coverage posited by q C .
Yarrowia lipolytica , an oleaginous yeast, shows promise for industrial fermentation due to its robust acetyl-CoA flux and well-developed genetic engineering tools. However, its lack of an active xylose metabolism restricts the conversion of cellulosic sugars to valuable products. To address this, metabolic engineering, and adaptive laboratory evolution (ALE) were applied to the Y. lipolytica PO1f strain, resulting in an efficient xylose-assimilating strain (XEV). Whole-genome sequencing (WGS) of the XEV followed by reverse engineering revealed that the amplification of the heterologous oxidoreductase pathway and a mutation in the GTPase-activating protein gene (YALI0B12100g) might be the primary reasons for improved xylose assimilation in the XEV strain. When a sorghum hydrolysate was used, the XEV strain showed superior xylose consumption and lipid production compared to its parental strain (X123). This study advances our understanding of xylose metabolism in Y. lipolytica and proposes effective metabolic engineering strategies for optimizing lignocellulosic hydrolysates.
Abstract. This article presents a validation study of the popular aeroservoelastic code suite OpenFAST leveraging weeks of measurements obtained during normal operation of a 2.8 MW land-based wind turbine. Measured wind conditions were used to generate one-to-one turbulent flow fields (i.e., comparing simulation to measurement in 10 min increments, or bins) through unconstrained and constrained assimilation methods using the kinematic turbulence generators TurbSim and PyConTurb. A total of 253 bins of 10 min of normal turbine operation were selected for analysis, and a statistical comparison in terms of performance and loads is presented. We show that successful validation of the model was not strongly dependent on the type of inflow assimilation method used for mean quantities of interest, which had median modeling errors per wind-speed interval generally within 5 %–10 % of the measurement. The type of inflow assimilation method did have a larger effect on the fatigue predictions for blade-root flapwise and tower-base fore–aft quantities, which surprisingly saw larger errors from the assumed higher-fidelity assimilation methods. Avenues for further work are discussed and include possible improvements to the aerodynamic, structural, and controller modeling that may offer insight on the origin of the up to ∼ 40 % median overprediction of fatigue for these quantities.
Multiphysics problems that are characterized by complex interactions among fluid dynamics, heat transfer, structural mechanics, and electromagnetics, are inherently challenging due to their coupled nature. While experimental data on certain state variables may be available, integrating these data with numerical solvers remains a significant challenge. Physics-informed neural networks (PINNs) have shown promising results in various engineering disciplines, particularly in handling noisy data and solving inverse problems in partial differential equations (PDEs). However, their effectiveness in forecasting nonlinear phenomena in multiphysics regimes, particularly involving turbulence, is yet to be fully established. Here, this study introduces NeuroSEM, a hybrid framework integrating PINNs with the highfidelity Spectral Element Method (SEM) solver, Nektar++. NeuroSEM leverages the strengths of both PINNs and SEM, providing robust solutions for multiphysics problems. PINNs are trained to assimilate data and model physical phenomena in specific subdomains, which are then integrated into the Nektar++ solver. We demonstrate the efficiency and accuracy of NeuroSEM for thermal convection in cavity flow and flow past a cylinder. The framework effectively handles data assimilation by addressing those subdomains and state variables where the data is available. We applied NeuroSEM to the Rayleigh-B´enard convection system, including cases with missing thermal boundary conditions and noisy datasets. Finally, we applied the proposed NeuroSEM framework to real particle image velocimetry (PIV) data to capture flow patterns characterized by horseshoe vortical structures. Our results indicate that NeuroSEM accurately models the physical phenomena and assimilates the data within the specified subdomains. The framework’s plug-and-play nature facilitates its extension to other multiphysics or multiscale problems. Furthermore, NeuroSEM is optimized for efficient execution on emerging integrated GPU-CPU architectures. This hybrid approach enhances the accuracy and efficiency of simulations, making it a powerful tool for tackling complex engineering challenges in various scientific domains.
This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.
The rapid increase in the volume and variety of terrestrial biosphere observations (i.e., remote sensing data and in situ measurements) offers a unique opportunity to derive ecological insights, refine process‐based models, and improve forecasting for decision support. However, despite their potential, ecological observations have primarily been used to benchmark process‐based models, as many past and current models lack the capability to directly integrate observations and their associated uncertainties for parameterization. In contrast, data assimilation frameworks such as the CARbon DAta MOdel fraMework (CARDAMOM) and its suite of process‐based models, known as the Data Assimilation Linked Ecosystem Carbon Model (DALEC), are specifically designed for model‐data fusion. This review, motivated by a recent CARDAMOM community workshop, examines the development and applications of CARDAMOM, with an emphasis on its role in advancing ecosystem process understanding. CARDAMOM employs a Bayesian approach, using a Markov Chain Monte Carlo algorithm to enable data‐driven calibration of DALEC parameters and initial states (i.e., carbon pool sizes) through observation operators. CARDAMOM's unique ability to retrieve localized model process parameters from diverse datasets—ranging from in situ measurements to global satellite observations—makes it a highly flexible tool for analyzing spatially variable ecosystem responses to environmental change. However, assimilating these data also presents challenges, including data quality issues that propagate into model skill, as well as trade‐offs between model complexity, parameter equifinality, and predictive performance. We discuss potential solutions to these challenges, such as reducing parameter equifinality by incorporating new observations. This review also offers community recommendations for incorporating emerging datasets, integrating machine learning techniques, strengthening collaboration with remote sensing, field, and modeling communities, and expanding CARDAMOM's relevance for localized ecosystem monitoring and decision‐making. CARDAMOM enables a deep, mechanistic understanding of terrestrial ecosystem dynamics that cannot be achieved through empirical analyses of observational datasets or weakly constrained models alone.
Providing robust real time flood warnings is of paramount importance to coastal communities. Although state-of-the-art hydrodynamic models are capable of robustly predicting Coastal Water Levels (CWL), unresolved drivers affecting level fluctuations are often not represented by the model governing equations. This work evaluates a novel method to improve the performance of the ADvanced CIRCulation (ADCIRC) hydrodynamic model by assimilating observations from four nadir-only satellite altimetry missions against a set of National Oceanic and Atmospheric Administration (NOAA) gauge stations located across the entire U.S. East Coast. Two different types of simulations were performed – Open Loop (OL) and Data Assimilation (DA). Five different simulations were performed where four different satellite altimetry observations were assimilated individually and combined with two different scenarios – with and without considering the data quality flags. Results indicate that, despite their limited spatial coverage, merging nadir-only observations into ADCIRC from the newly launched Surface Water and Ocean Topography (SWOT)’s nadir altimeter can improve the model performance at 76% of the gauge locations, whereas Sentinel-6 improves 73% of the total locations, Jason-3 74%, and SARAL 21%. Furthermore, combining observations from SWOT-nadir, Jason-3, and Sentinel-6 can improve the ADCIRC performance at more than 80% of the gauge locations for 107-day simulation. Nadir-only satellite altimetry observations can be useful for improving the model performance even if flagged as “poor quality” near the coast. When the flagged data are disregarded, SWOT can improve ADCIRC at 78%, Sentinel-6 at 73%, Jason-3 at 53%, and SARAL at 21% of the gauge locations. The ability to improve the model simulations largely depends on the availability of a satellite overpass nearby. Therefore, model performance can be further enhanced if satellite observations are available during a storm surge event, stressing the importance of frequent satellite overpasses.