A Statistical Perspective on Climate Informatics
No abstract available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
No abstract available
Accurate simulation of the quasi-biennial oscillation (QBO) is challenging due to uncertainties in representing convectively generated gravity waves. We develop an end-to-end uncertainty quantification workflow that calibrates these gravity wave processes in E3SM for a realistic QBO. Central to our approach is a domain knowledge-informed, compressed representation of high-dimensional spatio-temporal wind fields. By employing a parsimonious statistical model that learns the fundamental frequency from complex observations, we extract interpretable and physically meaningful quantities capturing key attributes. Building on this, we train a probabilistic surrogate model that approximates the fundamental characteristics of the QBO as functions of critical physics parameters governing gravity wave generation. Leveraging the Karhunen–Loève decomposition, our surrogate efficiently represents these characteristics as a set of orthogonal features, capturing cross-correlations among multiple physics quantities evaluated at different pressure levels and enabling rapid surrogate-based inference at a fraction of the computational cost of full-scale simulations. Finally, we analyze the inverse problem using a multi-objective approach. Our study reveals a tension between amplitude and period that constrains the QBO representation, precluding a single optimal solution. To navigate this, we quantify the bi-criteria trade-off and generate a set of Pareto optimal parameter values that balance the conflicting objectives. This integrated workflow improves the fidelity of QBO simulations and offers a versatile template for uncertainty quantification in complex geophysical models.
This paper presents the development of a new entropy-based feature selection method for identifying and quantifying impacts. Here, impacts are defined as statistically significant differences in spatio-temporal fields when comparing datasets with and without an external forcing in an Earth system model. Temporal feature selection is performed by first computing the cross-fuzzy entropy to quantify similarity of patterns between two datasets and then applying changepoint detection to identify regions of statistically constant entropy. The method is used to capture temperate north surface cooling from a 9-member simulation ensemble of the Mt. Pinatubo volcanic eruption, which injected 10 Tg of SO 2 into the stratosphere. The results estimate a mean difference decrease in near surface air temperature of -0.560 K with a 99% confidence interval between -0.864 K and -0.257 K between April and November of 1992, one year following the eruption. A sensitivity analysis with decreasing SO 2 injection revealed that the impact is statistically significant at 5 Tg but not at 3 Tg. Using identified features, a dependency graph model based on a 9-day lag had significantly fewer nodes than a graph based on monthly means. Furthermore, this demonstrates our method’s ability to perform dimension reduction while still uncovering source-to-impact pathways.
Recent years have seen a growing concern about climate change and its impacts. While Earth System Models (ESMs) can be invaluable tools for studying the impacts of climate change, the complex coupling processes encoded in ESMs and the large amounts of data produced by these models, together with the high internal variability of the Earth system, can obscure important source-to-impact relationships. Here, this paper presents a novel and efficient unsupervised data-driven approach for detecting statistically-significant impacts and tracing spatio-temporal source-impact pathways in the climate through a unique combination of ideas from anomaly detection, clustering and Natural Language Processing (NLP). Using as an exemplar the 1991 eruption of Mount Pinatubo in the Philippines, we demonstrate that the proposed approach is capable of detecting known post-eruption impacts/events. We additionally describe a methodology for extracting meaningful sequences of post-eruption impacts/events by using NLP to efficiently mine frequent multivariate cluster evolutions, which can be used to confirm or discover the chain of physical processes between a climate source and its impact(s).
This work demonstrates the development and implementation of a Fully Constrained Least Squares (FCLS) unmixing model developed in C++ programming language with OpenCV package and boost C++ libraries in the NASA Earth Exchange (NEX). Visualization of the results is supported by GRASS GIS and statistical analysis is carried in R in a Linux system environment. FCLS was first tested on computer simulated data with Gaussian noise of various signal-to-noise ratio, and Landsat data of an agricultural scenario and an urban environment using a set of global end members of substrate (soils, sediments, rocks, and non-photosynthetic vegetation), vegetation that includes green photosynthetic plants and dark objects which encompasses absorptive substrate materials, clear water, deep shadows, etc. For the agricultural scenario, a spectrally diverse collection of 11 scenes of Level 1 terrain corrected, cloud free Landsat-5 TM data of Fresno, California, USA were unmixed and the results were validated with the corresponding ground data. To study an urbanized landscape, a clear sky Landsat-5 TM data were unmixed and validated with coincident World View-2 abundance maps (of 2 m spatial resolution) for an area of San Francisco, California, USA. The results were evaluated using descriptive statistics, correlation coefficient, RMSE, probability of success, boxplot and bivariate distribution function. Finally, FCLS was used for sub-pixel land cover analysis of the monthly WELD (Wen-enabled Landsat data) repository from 2008 to 2011 of North America. The abundance maps in conjunction with DMSP-OLS nighttime lights data were used to extract the urban land cover features and analyze their spatial-temporal growth.
In the standard picture of fully developed turbulence, highly intermittent hydrodynamic fields are nonlinearly coupled across scales, where local energy cascades from large scales into dissipative vortices and large density gradients. Microscopically, however, constituent fluid molecules are in constant thermal (Brownian) motion, but the role of molecular fluctuations in large-scale turbulence is largely unknown, and with rare exceptions, it has historically been considered irrelevant at scales larger than the molecular mean free path. Recent theoretical and computational investigations have shown that molecular fluctuations can impact energy cascade at Kolmogorov length scales. Here, we show that molecular fluctuations not only modify energy spectrum at wavelengths larger than the Kolmogorov length in compressible turbulence, but also significantly inhibit spatio-temporal intermittency across the entire dissipation range. Using large-scale direct numerical simulations of computational fluctuating hydrodynamics, we demonstrate that the extreme intermittency characteristic of turbulence models is replaced by nearly Gaussian statistics in the dissipation range. These results demonstrate that the compressible Navier–Stokes equations should be augmented with molecular fluctuations to accurately predict turbulence statistics across the dissipation range. Our findings have significant consequences for turbulence modelling in applications such as astrophysics, reactive flows and hypersonic aerodynamics, where dissipation-range turbulence is approximated by closure models.
NASA has been collecting massive amounts of remote sensing data about Earth's systems for more than a decade. Missions are selected to be complementary in quantities measured, retrieval techniques, and sampling characteristics, so these datasets are highly synergistic. To fully exploit this, a rigorous methodology for combining data with heterogeneous sampling characteristics is required. For scientific purposes, the methodology must also provide quantitative measures of uncertainty that propagate input-data uncertainty appropriately. We view this as a statistical inference problem. The true but notdirectly- observed quantities form a vector-valued field continuous in space and time. Our goal is to infer those true values or some function of them, and provide to uncertainty quantification for those inferences. We use a spatiotemporal statistical model that relates the unobserved quantities of interest at point-level to the spatially aggregated, observed data. We describe and illustrate our method using CO2 data from two NASA data sets.
As part of JPL's ocean data assimilation effort to study ocean circulation and seasonal-interannual climate variability, sea level anomaly observed by TOPEX altimeter, together with sea surface temperature and wind stress data, are assimilated into a simple coupled ocean atmosphere model of the tropical Pacific. Model-data consistency is examined. Impact of the assimilation (as initialization) on El Nino Southern Oscillation (ENSO) forecasts is evaluated. The coupled model consists of a shallow water component with two baroclinic modes, an Ekman shear layer, a simplified mixed-layer temperature equation, and a statistical atmosphere based on dominant correlations between historical surface temperature and wind stress anomaly data. The adjoins method is used to fit the coupled model to the data over various six-month periods from late 1996 to early 1998 by optimally adjusting the initial state, model parameters, and basis functions of the statistical atmosphere. On average, the coupled model can be fitted to the data to approximately within the data and representation errors (5 cm, 0.5 C, and 10 sq m/sq m for sea level, surface temperature, and pseudo wind stress anomalies, respectively). The estimated fields resemble observed spatio-temporal structure reasonably well. Hindcasts/forecasts of the 1997/1998 El Nino initialized from forced estimated ocean states and parameters are much more realistic than those simply initialized from ocean states (see figure below). In particular, the ability of the model to produce significant warming beyond the initial state is dramatically improved. Parameter estimation, which compensates for some model errors, is found to be important to obtaining better fits of the model to data and to improving forecasts.
Earth system models (ESMs) are essential for understanding the interaction between human activities and the Earth's climate. However, the computational demands of ESMs often limit the number of simulations that can be run, hindering the robust analysis of risks associated with extreme weather events. While low-cost climate emulators have emerged as an alternative to emulate ESMs and enable rapid analysis of future climate, many of these emulators only provide output on at most a monthly frequency. This temporal resolution is insufficient for analyzing events that require daily characterization, such as heat waves or heavy precipitation. We propose using diffusion models, a class of generative deep learning models, to effectively downscale ESM output from a monthly to a daily frequency. Trained on a handful of ESM realizations, reflecting a wide range of radiative forcings, our DiffESM model takes monthly mean precipitation or temperature as input, and is capable of producing daily values with statistical characteristics close to ESM output. Combined with a low-cost emulator providing monthly means, this approach requires only a small fraction of the computational resources needed to run a large ensemble. We evaluate model behavior using a number of extreme metrics, showing that DiffESM closely matches the spatio-temporal behavior of the ESM output it emulates in terms of the frequency and spatial characteristics of phenomena such as heat waves, dry spells, or rainfall intensity.
pyTCR is a climatology software package developed in the Python programming language. It integrates the capabilities of several legacy physical models and increases computational efficiency to allow rapid estimation of tropical cyclone (TC) rainfall consistent with the large-scale environment. Specifically, pyTCR implements a horizontally distributed and vertically integrated model [Zhu et al., 2013] for simulating rainfall driven by TCs. Along storm tracks, rainfall is estimated by computing the cross-boundary-layer, upward water vapor transport caused by different mechanisms including frictional convergence, vortex stretching, large-scale baroclinic effect (i.e., wind shear), topographic forcing, and radiative cooling [Lu et al., 2018]. The package provides essential functionalities for modeling and interpreting spatio-temporal TC rainfall data. pyTCR requires a limited number of model input parameters, making it a convenient and useful tool for analyzing rainfall mechanisms driven by TCs. To sample rare (most intense) rainfall events that are often of great societal interest, pyTCR adapts and leverages outputs from a statistical-dynamical TC downscaling model [Lin et al., 2023] capable of rapidly generating a large number of synthetic TCs given a certain climate. As a result, pyTCR significantly reduces computational effort and improves the efficiency in capturing extreme TC rainfall events at the tail of the distributions from limited datasets. Furthermore, the TC downscaling model is forced entirely by large-scale environmental conditions from reanalysis data or coupled General Circulation Models (GCMs), simplifying the projection of TC-induced rainfall and wind speed under future climate using pyTCR. Finally, pyTCR can be coupled with hydrological and wind models to assess risks associated with independent and compound events (e.g., storm surges and freshwater flooding).
Crop phenology regulates seasonal agroecosystem carbon, water, and energy exchanges, and is a key component in empirical and process-based crop models for simulating biogeochemical cycles of farmlands, assessing gross and net primary production, and forecasting the crop yield. The advances in phenology matching models provide a feasible means to monitor crop phenological progress using remote sensing observations, with a priori information of reference shapes and reference phenological transition dates. Yet the underlying geometrical scaling assumption of models, together with the challenge in defining phenological references, hinders the applicability of phenology matching in crop phenological studies. The objective of this study is to develop a novel hybrid phenology matching model to robustly retrieve a diverse spectrum of crop phenological stages using satellite time series. The devised hybrid model leverages the complementary strengths of phenometric extraction methods and phenology matching models. It relaxes the geometrical scaling assumption and can characterize key phenological stages of crop cycles, ranging from farming practice-relevant stages (e.g., planted and harvested) to crop development stages (e.g., emerged and mature). To systematically evaluate the influence of phenological references on phenology matching, four representative phenological reference scenarios under varying levels of phenological calibrations in terms of time and space are further designed with publicly accessible phenological information. The results indicate that the hybrid phenology matching model can achieve high accuracies for estimating corn and soybean phenological growth stages in Illinois, particularly with the year- and region-adjusted phenological reference (R-squared higher than 0.9 and RMSE less than 5 days for most phenological stages). The inter-annual and regional phenological patterns characterized by the hybrid model correspond well with those in the crop progress reports (CPRs) from the USDA National Agricultural Statistics Service (NASS). Compared to the benchmark phenology matching model, the hybrid model is more robust to the decreasing levels of phenological reference calibrations, and is particularly advantageous in retrieving crop early phenological stages (e.g., planted and emerged stages) when the phenological reference information is limited. This innovative hybrid phenology matching model, together with CPR-enabled phenological reference calibrations, holds 3 considerable promise in revealing spatio-temporal patterns of crop phenology over extended geographical regions.
Clouds cover approximately 60% of the earth's surface. When obscuring the satellite's field of view (FOV), clouds complicate the retrieval of ozone, trace gases and aerosols from data collected by earth observing satellites. Cloud properties associated with optical thickness, cloud pressure, water phase, drop size distribution (DSD), cloud fraction, vertical and areal extent can also change significantly over short spatio-temporal scales. The radiative transfer models used to retrieve column estimates of atmospheric constituents typically do not account for all these properties and their variations. The OMI science team is preparing to release a new data product, OMMYDCLD, which combines the cloud information from sensors on board two earth observing satellites in the NASA A-Train: Aura/OMI and Aqua/MODIS. OMMYDCLD co-locates high resolution cloud and radiance information from MODIS onto the much larger OMI pixel and combines it with parameters derived from the two other OMI cloud products: OMCLDRR and OMCLDO2. The product includes histograms for MODIS scientific data sets (SDS) provided at 1 km resolution. The statistics of key data fields - such as effective particle radius, cloud optical thickness and cloud water path - are further separated into liquid and ice categories using the optical and IR phase information. OMMYDCLD offers users of OMI data cloud information that will be useful for carrying out OMI calibration work, multi-year studies of cloud vertical structure and in the identification and classification of multi-layer clouds.
We present a hybrid machine learning framework that combines physics-informed neural operators (PINOs) with score-based generative diffusion models to simulate the full spatio-temporal evolution of two-dimensional, incompressible, resistive magnetohydrodynamic turbulence across a broad range of Reynolds numbers (Re). The framework leverages the equation-constrained generalization capabilities of PINOs to predict coherent, low-frequency dynamics, while a conditional diffusion model stochastically corrects high-frequency residuals, enabling accurate modeling of fully developed turbulence. Trained on a comprehensive ensemble of high-fidelity simulations with Re ϵ {100, 250, 500, 750, 1000, 3000, 10000}, the approach achieves state-of-the-art accuracy in regimes previously inaccessible to deterministic surrogates. At Re = 1000 and 3000, the model faithfully reconstructs the full spectral energy distributions of both velocity and magnetic fields late into the simulation, capturing non-Gaussian statistics, intermittent structures, and cross-field correlations with high fidelity. At extreme turbulence levels (Re = 10 000), it remains the first surrogate capable of recovering the high-wavenumber evolution of the magnetic field, preserving large-scale morphology and enabling statistically meaningful predictions.
Neural operators are promising surrogates for dynamical systems but when trained with standard L 2 losses they tend to oversmooth fine-scale turbulent structures. Here, we show that combining operator learning with generative modeling overcomes this limitation. We consider three practical turbulent-flow challenges where conventional neural operators fail: spatio-temporal super-resolution, forecasting, and sparse flow reconstruction. For Schlieren jet super-resolution, an adversarially trained neural operator (adv-NO) reduces the energy-spectrum error by 15 × while preserving sharp gradients at neural operator-like inference cost. For 3D homogeneous isotropic turbulence, adv-NO trained on only 160 timesteps from a single trajectory forecasts accurately for five eddy-turnover times and offers 114 × wall-clock speed-up at inference than the baseline diffusion-based forecasters, enabling near-real-time rollouts. For reconstructing cylinder wake flows from highly sparse Particle Tracking Velocimetry-like inputs, a conditional generative model infers full 3D velocity and pressure fields with correct phase alignment and statistics. These advances enable accurate reconstruction and forecasting at low compute cost, bringing near-real-time analysis and control within reach in experimental and computational fluid mechanics.
Human mobility trajectories provide valuable information for developing mobility applications, as they contain diverse and rich information about the users. User mobility data is valuable for various applications such as intelligent transportation systems (ITS), commercial business models, and disease-spread models. However, such spatio-temporal traces may pose a threat to user privacy. GPS trajectories in their raw form are not suitable for transportation studies, as they require matching locations with nearest road links — a process called map-matching. This software implements a differential privacy (DP)-based map-matching algorithm, called DPMM, that generates link-level location trajectories in a privacy-preserving manner to protect users' origin destinations (OD) and travel paths. OD privacy is achieved by injecting Planar Laplace noise to the user OD GPS points. Travel-path privacy is provided with randomized travel path construction using exponential DP mechanism. The injected noise level is selected adaptively, by considering the link density of the location and the functional category of the localized links. For path privacy, our mechanism samples waypoints and selects candidate paths between waypoints. DPMM provides privacy effectively with respect to link density instead of other trajectory samples in the database compared to other privacy mechanisms. Compared to the different baseline models our DP-based privacy model offers closer query responses to the raw data in terms of individual and aggregate trajectory-level statistics with an average at absolute deviation from the baseline for individual statistics on ϵ = 1.0. Beyond individual trajectory statistics, the DPMM outperforms the other benchmark DP-based mechanisms on different aggregate statistics with up to 8x improvement in utility.
Satellite precipitation products, as all quantitative estimates, come with some inherent degree of uncertainty. To associate a quantitative value of the uncertainty to each individual estimate, error modeling is necessary. Most of the error models proposed so far compute the uncertainty as a function of precipitation intensity only, and only at one specific spatio-temporal scale. We propose a spectral error model which accounts for the neighboring space-time dynamics of precipitation into the uncertainty quantification. Systematic distortions of the precipitation signal and random errors are characterized distinctively in every frequency-wavenumber band in the Fourier domain, to accurately characterize error across scales. The systematic distortions are represented as a deterministic space-time linear filtering term. The random errors are represented as a non-stationary additive noise. The spectral error model is applied to the IMERG multi satellite precipitation product and its parameters are estimated empirically through a system identification approach using the GV-MRMS gauge-radar measurements as reference (“truth”) over the eastern United States. The filtering term is found to be essentially low-pass. While traditional error models attribute most of the error variance to random errors, it is found here that the systematic filtering term explains 48% of the error variance at the native resolution of IMERG. This fact confirms that, at high resolution, filtering effects in satellite precipitation products cannot be ignored, and that the error cannot be represented as a purely random additive or multiplicative term. An important consequence is that precipitation estimates derived from totally different sources shall not be expected to automatically have statistically independent errors.
Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system’s internal variability.
Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.