Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistics of extremes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Quantifying the impacts of compound extremes on agriculture

Agricultural production and food prices are affected by hydroclimatic extremes. There has been a growing amount of literature measuring the impacts of individual extreme events (heat stress or water stress) on agricultural and human systems. Yet, we lack a comprehensive understanding of the significance and the magnitude of the impacts of compound extremes. This study combines a fine-scale weather product with outputs of a hydrological model to construct functional metrics of individual and compound hydroclimatic extremes for agriculture. Then, a yield response function is estimated with individual and compound metrics, focusing on corn in the United States during the 1981–2015 period. Supported by statistical evidence, the findings suggest that metrics of compound hydroclimatic extremes are better predictors of corn yield variations than metrics of individual extremes. The results also confirm that wet heat is more damaging than dry heat for corn. This study shows the average yield damage from heat stress has been up to four times more severe when combined with water stress.

54 ENVIRONMENTAL SCIENCES↗

Emergence of robust anthropogenic increase of heat stress-related variables projected from CORDEX-CORE climate simulations

The information of when and where region-specific patterns in both mean and extreme temperatures leading to heat stress will emerge from the present-day climate variability is important to plan adaptation options, but to date studies on this issue still remain limited and fragmented. Here, we estimate the time of emergence (TOE) of temperature and wet-bulb temperature (Tw), a better indication of heat stress, using fine-scale, long-term regional climate model projections under the RCP2.6 and RCP8.5 scenarios across six different domains. Differently from previous studies, the TOE is determined using three methods applied on impact-relevant variables: two different signal-to-noise frameworks based on summer mean temperature and Tw and a statistical test to identify significant differences in daily extreme distributions. The TOE response to RCP2.6 and RCP8.5 with respect to the end of 20th century variability differs significantly regardless of which TOE metric is applied. For summer mean temperature, the land fraction reaching TOE is expected to exceed 90% by the 2050s under the RCP8.5, whereas the increase rate of land exposure to TOE tends to stagnate over time under the RCP2.6 so that more than 40% of land will not experience TOE by the end of the 21st century. Compared to temperature, the TOE of Tw is reached earlier in most of the wet tropics but is delayed in hot and dry regions because of the nonlinear response of Tw to humidity. For both temperature and Tw, the TOE appears earlier in regions with low baseline variability, such as in the tropics. Despite the uncertainties arising from the choice of TOE metrics, the vast majority of regions in Africa and southeast Asia experience TOE in the early 21st century under both the RCP2.6 and RCP8.5 scenarios, which stresses the urgent need for developing adequate adaptation strategies in these regions.

54 ENVIRONMENTAL SCIENCES↗

Accelerating quantum optics experiments with statistical learning

Quantum optics experiments, involving the measurement of low-probability photon events, are known to be extremely time-consuming. We present a methodology for accelerating such experiments using physically motivated ansatzes together with simple statistical learning techniques such as Bayesian maximum a posteriori estimation based on few-shot data. We show that it is possible to reconstruct time-dependent data using a small number of detected photons, allowing for fast estimates in under a minute and providing a one-to-two order of magnitude speed-up in data acquisition time. We test our approach using real experimental data to retrieve the second order intensity correlation function, G (2) ($τ$), as a function of time delay τ between detector counts, for thermal light as well as anti-bunched light emitted by a quantum dot driven by periodic laser pulses. The proposed methodology has a wide range of applicability and has the potential to impact the scientific discovery process across a multitude of domains.

36 MATERIALS SCIENCE↗

How Frequent Will the Rarest Daily Rainfall Records of Hurricane Ida’s Remnants Be in the Future?

Abstract Gaining continued insights into the impact of global warming on the occurrence of hurricane-associated intense record downpours is essential for building climate resilient communities. This study investigates projected future changes in extreme rainfall over the Northeast United States, as represented by extreme daily amounts during Hurricane Ida in 2021. We used historical control simulations of Weather Research and Forecasting (WRF) Model generated from 40 years of weather events (1980–2014, 12 km) forced by the fifth generation European Centre for Medium-Range Weather Forecasts atmospheric reanalysis. These simulations are thermodynamically modified (2060–2100) via an imposed warming for the high-emission scenario of shared socioeconomic pathway (SSP585) from a range of general circulation models. Ground observations from the Global Historical Climatology Network (1950–2014) and WRF simulations (historical, 1980–2014, and future, 2060–2100) are integrated into a nonstationary generalized extreme value (GEV) framework to assess the frequency of Ida’s heaviest daily rain rates under the SSP585 scenario. Results show that Ida’s daily maximum rainfall recorded at different observation locations was higher than the single highest September daily maximum observed (1950–2014) for 5 out of 17 stations (∼30% of the stations). Ida-like extreme daily rain rates are projected to be, on average, more than 2 times more likely to occur at the end of the century in the simulations (with some regions as high as 5 times). This work demonstrates that integrating a high-resolution atmospheric model’s present-day and thermodynamically modified future simulations along with ground observations, within a nonstationary statistical framework, is crucial for understanding changing characteristics of extreme weather events. Significance Statement Daily scale extreme precipitation is expected to become more frequent and severe, as evidenced by observations and model simulations. While it is important to investigate how these intensifying heavy rainfall events affect current engineering standards, fewer studies have contextualized how warming impacts the most extreme rainfall from a single storm event relative to historical heavy downpours. In this study, we focused on the daily extreme rainfall associated with the extratropical transition of Hurricane Ida (2021), particularly over the northeastern United States—some of which exceeded the commonly used hydrologic design criteria for a 100-yr storm. Using a high-resolution atmospheric model simulation, we investigated how continued warming may influence the frequency of such daily rain rates. Under a high-emission scenario, these events are projected to become up to 5 times more likely at the end of the twenty-first century.

Dollan, Ishrat J↗

Subseasonal Clustering of Atmospheric Rivers Over the Western United States

Abstract The serial occurrence of atmospheric rivers (ARs) along the US West Coast can lead to prolonged and exacerbated hydrologic impacts, threatening flood‐control and water‐supply infrastructure due to soil saturation and diminished recovery time between storms. Here a statistical approach for quantifying subseasonal temporal clustering among extreme events is applied to a 41‐year (1979–2019) wintertime AR catalog across the western United States (US). Observed AR occurrence, compared against a randomly distributed AR timeseries with the same average event density, reveals temporal clustering at a greater‐than‐random rate across the western US with a distinct geographical pattern. Compared to the Pacific Northwest, significant AR clusters over the northern Coastal Range of California and Sierra Nevada are more frequent and occur over longer time periods. Clusters along the California Coastal Range typically persist for 2 weeks, are composed of 4–5 ARs per cluster, and account for over 85% of total AR occurrence. Across the northwest Coast‐Cascade Ranges, clusters account for ∼50% of total AR occurrence, typically last 8–10 days, and contain 3–4 individual AR events. Based on precipitation data from a high‐resolution dynamical downscaling of reanalysis, the fractions of total and extreme hourly precipitation attributable to AR clusters are largest along the northern California coast and in the Sierra Nevada. Interannual variability among clusters highlights their importance for determining whether a particular water year is anomalously wet or dry. The mechanisms behind this unusual clustering are unclear and require further research.

Meteorology & Atmospheric Sciences↗

Going Off Grid: A Comparative Study of the Lagrangian and Eulerian Perspectives of New Particle Formation Events

New particle formation and growth (NPF&G) is the process by which ultrafine particles are formed from gas-phase precursors. NPF&G is the dominant source of global aerosol number with important influences on climate. Most observations of NPF&G events are conducted at stationary sites; however, NPF&G observed from stationary sites is influenced by gradual or rapid changes in the air masses passing over the site, complicating NPF&G analysis. In this work, we use observations and a 3D aerosol model to compare aerosol size distributions at a stationary site (Southern Great Plains [SGP] observatory, Oklahoma, USA) and along Lagrangian trajectories crossing the site. The model simulates the NPF&G events reasonably well at SGP. Using the model to compare the Lagrangian and stationary perspectives, we can explain previously unanalyzable days with some evidence of NPF&G as either non-event or analyzable NPF&G days. We find most of the unanalyzable NPF&G days are due to isolated and inhomogeneous NPF&G occurring upwind of the stationary site, often in the outflow of urban regions. Finally, we compare formation rates of 3 nm particles, growth rates, and the survival probability of 3 nm particles growing to 25 nm between the stationary and Lagrangian perspectives. Because of the much larger number of analyzable days along the Lagrangian trajectories, this perspective potentially provides more robust statistics and better characterization of NPF&G event extremes. Our method for extracting chemical/physical properties along Lagrangian trajectories from 3D models can be applied to a wide range of science questions.

O’Donnell, Samuel E. [Colorado State Univ., Fort C↗

Climate Nowcasting

The climate is changing so rapidly that climatologies based on historical statistics cannot reliably capture the current risk of extreme weather events hazardous to society. Decision relevant projections of weather extreme probability over the next 10–15 years are needed to enable adaptation and resilience in the face of this evolving risk. Current weather forecasts/predictions and long-term climate projections for decades into the future are inadequate for providing this information to stakeholders that need it, targeting forecast horizons either too short or too far into the future. We argue that a new approach is needed: climate nowcasting. Climate nowcasting would focus on user-inspired extreme metrics over the next 10–15 year time frame, targeting specific impacts and locations down to a local scale, by engaging with stakeholders to understand their needs and provide information in a format relevant for decision making. Importantly, climate nowcasting will not consist of a single approach or data set, involving rather data fusion from different sources of information: simulations, observations and data driven methods, likely through different weights depending on the metric of interest. Predictions must be accompanied by serious engagement with stakeholders facing climate risks, and clearly present uncertainties and limitations of any prediction. Such a vision is very different from how typical climate or weather forecasts are applied today.

Gettelman, Andrew↗

Inference of the Mass Composition of Cosmic Rays with Energies from 10 18.5 to 10 20 eV Using the Pierre Auger Observatory and Deep Learning

We present measurements of the atmospheric depth of the shower maximum X max , inferred for the first time on an event-by-event level using the surface detector of the Pierre Auger Observatory. Using deep learning, we were able to extend measurements of the X max distributions up to energies of 100 EeV ( 10 20 eV ), not yet revealed by current measurements, providing new insights into the mass composition of cosmic rays at extreme energies. Gaining a 10-fold increase in statistics compared to the fluorescence detector data, we find evidence that the rate of change of the average X max with the logarithm of energy features three breaks at 6.5 ± 0.6 ( stat ) ± 1 ( syst ) EeV , 11 ± 2 ( stat ) ± 1 ( syst ) EeV , and 31 ± 5 ( stat ) ± 3 ( syst ) EeV , in the vicinity to the three prominent features (ankle, instep, suppression) of the cosmic-ray flux. The energy evolution of the mean and standard deviation of the measured X max distributions indicates that the mass composition becomes increasingly heavier and purer, thus being incompatible with a large fraction of light nuclei between 50 and 100 EeV. Published by the American Physical Society 2025

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

(U) Correlated Sampling Using Batch Statistics to Reduce the Uncertainty of Combinations of KSEN Outputs with MCNP6

The relative sensitivity of k eff to the densities of nuclides in a material are combined to compute relative sensitivities to other inputs. When computed in a single Monte Carlo run, the nuclide density sensitivities are correlated, and the statistical uncertainties propagated to other inputs will be incorrect unless those correlations are accounted for. Equations are presented to apply correlated sampling using batch statistics for the sum of an arbitrary number of random tallies and the difference of two random tallies when each is multiplied by a different constant. When correlated sampling is used on a recent benchmark evaluation, the correct statistical uncertainties for certain combinations of sensitivities are dramatically smaller than the incorrect uncertainties. A user-controlled, regular output of MCNP6’s KSEN sensitivities that allows batch statistics to be applied to combinations would be extremely valuable.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Reference Site Conditions for Floating Wind Arrays in the United States

Floating offshore wind farm design is highly site-specific, requiring detailed information about the specific conditions of a project area for realistic design studies. Unfortunately, publicly available site condition data for potential floating offshore wind project sites in the United States is scarce. To support U.S. offshore wind research, we developed reference site condition datasets, including metocean and seabed information, for four potential floating wind project areas in the U.S.: Humboldt Bay, Morro Bay, the Gulf of Maine, and the Gulf of Mexico. These datasets were compiled using publicly available data. Our metocean analysis, covering wind, waves, and surface currents, utilized measurement data from 2000 to 2020. Sources included the National Renewable Energy Laboratory’s National Offshore Wind Dataset for wind data, National Data Buoy Center buoys for wave data, and the High Frequency Radar Network for surface currents. These data were integrated into hourly time series used to compute extreme return periods up to 500 years, monthly statistics, and joint probability clusters for fatigue analysis. Soil conditions were evaluated using the usSEABED database and bathymetry grids were interpolated from the NCEI Digital Elevation Model Global Mosaic. In addition to providing curated reference site condition datasets for four U.S. areas, our assessment highlights the need for more publicly available metocean and soil condition data.

17 WIND ENERGY↗

Can Simple Metrics Identify the Process(es) Driving Extreme Precipitation?

This work seeks an automatic algorithm to determine the primary meteorological cause(s) of individual extreme precipitation events. Such determinations have been made before, but required a by-hand analysis of each separate event. This is very time-consuming and the field would benefit from an automatic process. This is especially relevant when comparing different datasets to determine which ones most closely hew towards reality. This paper tests three simple metrics over the continental United States using the European Center for Medium-Range Weather Forecasting’s (ECMWF) atmospheric reanalysis (ERA5). The metrics tested measure and compare the strength of three meteorological processes associated with extreme precipitation: fronts, convection, and cyclones. A multivariate statistical technique as well as individual case studies show evidence that the three meteorological processes of interest cannot be isolated from one another using these simple physical metrics. This shows the difficulty in finding “pure” cases of these precipitation-generating processes and suggests approaching these processes with an eye toward mixed-type events.

Swenson, Leif M. (ORCID:0000000199708735)↗

A Detailed Investigation into the Use of Planetary Nebulae as Standard Candles

The program's goal was to understand the physics underlying the [O III] (lambda)5007 planetary nebula luminosity function (PNLF) and evaluate its accuracy as an extragalactic distance indicator. Work under the grant concentrated in two areas. The first major goal was to extensively test the PNLF method to find its limits. We did this performing yet another internal test of the method in the core galaxies of the Fornax Cluster, performing external comparisons of PNLF distances with distances derived from Cepheids and the Surface Brightness Fluctuation method (SBF), and, in general, examining the PNLF in as many different galactic environments as possible, including the disks of late-type spirals. Because of the difficulty distinguishing planetary nebulae (PNe) from H II regions, and because spiral galaxies have uneven internal extinction, the process of identifying "statistical" samples of PNe in these objects is extremely complicated. Nevertheless, by using the ratio of [O III] (lambda)5007 to H(alpha) as a diagnostic, we were able to effectively discriminate PNe from most H II regions, and apply the method to systems such as NGC 300, M101, M51, and M96. The second goal of this research was to determine theoretically, why the PNLF is such an excellent standard candle.

Ciardullo, Robin↗

Data-Driven Transition Path Analysis Yields a Statistical Understanding of Sudden Stratospheric Warming Events in an Idealized Model

Abstract Atmospheric regime transitions are highly impactful as drivers of extreme weather events, but pose two formidable modeling challenges: predicting the next event (weather forecasting) and characterizing the statistics of events of a given severity (the risk climatology). Each event has a different duration and spatial structure, making it hard to define an objective “average event.” We argue here that transition path theory (TPT), a stochastic process framework, is an appropriate tool for the task. We demonstrate TPT’s capacities on a wave–mean flow model of sudden stratospheric warmings (SSWs) developed by Holton and Mass, which is idealized enough for transparent TPT analysis but complex enough to demonstrate computational scalability. Whereas a recent article (Finkel et al. 2021) studied near-term SSW predictability, the present article uses TPT to link predictability to long-term SSW frequency. This requires not only forecasting forward in time from an initial condition, but also backward in time to assess the probability of the initial conditions themselves. TPT enables one to condition the dynamics on the regime transition occurring, and thus visualize its physical drivers with a vector field called the reactive current . The reactive current shows that before an SSW, dissipation and stochastic forcing drive a slow decay of vortex strength at lower altitudes. The response of upper-level winds is late and sudden, occurring only after the transition is almost complete from a probabilistic point of view. This case study demonstrates that TPT quantities, visualized in a space of physically meaningful variables, can help one understand the dynamics of regime transitions.

Meteorology & Atmospheric Sciences↗

A Comparison of Infectious Disease Forecasting Methods across Locations, Diseases, and Time

Accurate infectious disease forecasting can inform efforts to prevent outbreaks and mitigate adverse impacts. This study compares the performance of statistical, machine learning (ML), and deep learning (DL) approaches in forecasting infectious disease incidences across different countries and time intervals. We forecasted three diverse diseases: campylobacteriosis, typhoid, and Q-fever, using a wide variety of features (n = 46) from public datasets, e.g., landscape, climate, and socioeconomic factors. We compared autoregressive statistical models to two tree-based ML models (extreme gradient boosted trees [XGB] and random forest [RF]) and two DL models (multi-layer perceptron and encoder–decoder model). The disease models were trained on data from seven different countries at the region-level between 2009–2017. Forecasting performance of all models was assessed using mean absolute error, root mean square error, and Poisson deviance across Australia, Israel, and the United States for the months of January through August of 2018. The overall model results were compared across diseases as well as various data splits, including country, regions with highest and lowest cases, and the forecasted months out (i.e., nowcasting, short-term, and long-term forecasting). Overall, the XGB models performed the best for all diseases and, in general, tree-based ML models performed the best when looking at data splits. There were a few instances where the statistical or DL models had minutely smaller error metrics for specific subsets of typhoid, which is a disease with very low case counts. Feature importance per disease was measured by using four tree-based ML models (i.e., XGB and RF with and without region name as a feature). The most important feature groups included previous case counts, region name, population counts and density, mortality causes of neonatal to under 5 years of age, sanitation factors, and elevation. This study demonstrates the power of ML approaches to incorporate a wide range of factors to forecast various diseases, regardless of location, more accurately than traditional statistical approaches.

59 BASIC BIOLOGICAL SCIENCES↗

Trust Your Gut: Comparing Human and Machine Inference from Noisy Visualizations

People commonly utilize visualizations not only to examine a given dataset, but also to draw generalizable conclusions about the underlying models or phenomena. Prior research has compared human visual inference to that of an optimal Bayesian agent, with deviations from rational analysis viewed as problematic. However, human reliance on non-normative heuristics may prove advantageous in certain circumstances. We investigate scenarios where human intuition might surpass idealized statistical rationality. In two experiments, we examine individuals’ accuracy in characterizing the parameters of known data-generating models from bivariate visualizations. Our findings indicate that, although participants generally exhibited lower accuracy compared to statistical models, they frequently outperformed Bayesian agents, particularly when faced with extreme samples. Participants appeared to rely on their internal models to filter out noisy visualizations, thus improving their resilience against spurious data. However, participants displayed overconfidence and struggled with uncertainty estimation. They also exhibited higher variance than statistical machines. Our findings suggest that analyst gut reactions to visualizations may provide an advantage, even when departing from rationality. These results carry implications for designing visual analytics tools, offering new perspectives on how to integrate statistical models and analyst intuition for improved inference and decision-making. The data and materials for this paper are available at https://osf.io/qmfv6

human-machine collaboration↗

Resolving turbulent magnetohydrodynamics: a hybrid operator-diffusion framework

We present a hybrid machine learning framework that combines physics-informed neural operators (PINOs) with score-based generative diffusion models to simulate the full spatio-temporal evolution of two-dimensional, incompressible, resistive magnetohydrodynamic turbulence across a broad range of Reynolds numbers (Re). The framework leverages the equation-constrained generalization capabilities of PINOs to predict coherent, low-frequency dynamics, while a conditional diffusion model stochastically corrects high-frequency residuals, enabling accurate modeling of fully developed turbulence. Trained on a comprehensive ensemble of high-fidelity simulations with Re ϵ {100, 250, 500, 750, 1000, 3000, 10000}, the approach achieves state-of-the-art accuracy in regimes previously inaccessible to deterministic surrogates. At Re = 1000 and 3000, the model faithfully reconstructs the full spectral energy distributions of both velocity and magnetic fields late into the simulation, capturing non-Gaussian statistics, intermittent structures, and cross-field correlations with high fidelity. At extreme turbulence levels (Re = 10 000), it remains the first surrogate capable of recovering the high-wavenumber evolution of the magnetic field, preserving large-scale morphology and enabling statistically meaningful predictions.

Diffusion-Integrated Neural Operators↗

Quantifying Graph Uncertainty from Communication Data

Graphs are a widely used abstraction for representing a variety of important real-world problems including emulating cyber networks for situational awareness, or studying social networks to understand human interactions or pandemic spread. Communication data is often converted into graphs to help understand social and technical patterns in the underlying communication data. However, prior to this project, little work had been performed analyzing how best to develop graphs from such data. Thus, many critical, national security problems were being performed against graph representations of questionable quality. Herein, we describe our analyses that were precursors to our final statistically grounded technique for creating static graph snapshots from a stream of communication events. The first analyzes the statistical distribution properties of a variety of real-world communication datasets generally fit best by Pareto, log normal, and extreme value distributions. The second derives graph properties that can be estimated given the expected statistical distribution for communication events and the communication interval to be viewed node observability, edge observability, and expected accuracy of node degree. Unfortunately, as that final process is under review for publication, we can't publish it here at this time.

97 MATHEMATICS AND COMPUTING↗