Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Impact of survey spatial variability on galaxy redshift distributions and the cosmological 3 × 2-point statistics for the Rubin Legacy Survey of Space and Time (LSST)

We investigate the impact of spatial survey non-uniformity on the galaxy redshift distributions for forthcoming data releases of the Rubin Observatory Legacy Survey of Space and Time (LSST). Specifically, we construct a mock photometry data set degraded by the Rubin OpSim observing conditions, and estimate photometric redshifts of the sample using a template-fitting photo-z estimator, BPZ, and a machine learning method, FlexZBoost. We select the Gold sample, defined as $i\lt 25.3$ for 10 yr LSST data, with an adjusted magnitude cut for each year and divide it into five tomographic redshift bins for the weak lensing lens and source samples. We quantify the change in the number of objects, mean redshift, and width of each tomographic bin as a function of the coadd i-band depth for 1-yr (Y1), 3-yr (Y3), and 5-yr (Y5) data. In particular, Y3 and Y5 have large non-uniformity due to the rolling cadence of LSST, hence provide a worst-case scenario of the impact from non-uniformity. We find that these quantities typically increase with depth, and the variation can be $10\!-\!40~{{\rm per\ cent}}$ at extreme depth values. Using Y3 as an example, we propagate the variable depth effect to the weak lensing $3\times 2$ pt analysis, and assess the impact on cosmological parameters via a Fisher forecast. We find that galaxy clustering is most susceptible to variable depth, and non-uniformity needs to be mitigated below 3 per cent to recover unbiased cosmological constraints. There is little impact on galaxy–shear and shear–shear power spectra, given the expected LSST Y3 noise.

cosmology↗

High-Resolution South American Wind Resource Data Downscaled with Generative Machine Learning Conditioned on Near-Surface Observations

High-resolution historical wind data was developed for the entirety of South America using the innovative Super-Resolution for Renewable Resource Data (sup3r) machine learning framework. The publicly available Sup3rWind South America dataset represents a significant advancement in wind resource data generation, leveraging generative machine learning conditioned on near-surface observations from the Meteorological Assimilation Data Ingest System (MADIS) to efficiently and accurately downscale coarse reanalysis data from the European Centre for Medium-Range Weather Forecasts (ERA5). This approach produces fine-scale, spatially and temporally coherent wind and meteorological fields hundreds of times more computationally efficient than traditional numerical weather modeling methods, enabling access to high-fidelity wind information across both continental and offshore regions. Sup3rWind South America builds on the earlier Sup3rWind Ukraine dataset through improvements in model architecture and outputs conditioned on near-surface observation inputs. As with the Ukraine data release, this dataset includes wind speed, wind direction, temperature, relative humidity, and pressure at a horizontal resolution of ~2 km, representing a 15x spatial enhancement relative to the 31 km ERA5 grid. Wind speed and direction are provided at 5-minute resolution, a 12x temporal refinement compared to the hourly ERA5 data, while temperature, relative humidity, and pressure remain at hourly resolution. The data covers all years from 2005 to 2024. Before downscaling, ERA5 inputs were bias-corrected using long-term monthly means and a limited number of quality-controlled observations to align large-scale statistics with regional conditions. The resulting dataset is the first publicly available high-resolution timeseries wind record that provides full spatial coverage of South America. Model validation demonstrates strong agreement with observations across several statistical metrics, consistent with other state-of-the-art high-resolution wind resource datasets. The potential applications of Sup3rWind South America span renewable energy resource assessment, energy system modeling, and grid resilience analysis. The 20-year record and high spatial and temporal resolution support accurate estimation of long-term energy yield and the economic feasibility of potential wind development sites. Continuous coverage across both continental and offshore regions enables comprehensive site prospecting within exclusive economic zones. The 2 km, 5-minute resolution data provide the spatial and temporal variability required for power system simulation, operational planning, and regional risk assessments.

17 WIND ENERGY↗

Coupling localized Noah-MP-Crop model with the WRF model improved dynamic crop growth simulation across Northeast China

Croplands play a critical role in regulating the energy and moisture exchanges between the land surface and atmosphere. However, the interactions between cropland and climate are usually poorly represented due to a lack of detailed representation in crop types and field management. Here, we coupled the Noah-MP-Crop model with the state-of-the-art Weather Research and Forecasting (WRF) model to explore and evaluate the crop growth dynamics in response to climate variations across Northeast China. The default parameters of the crop model were not exactly suitable for the agricultural ecosystems in Northeast China. The detailed cropland distribution, and crop phenology parameters including growing degree days (GDD) and planting (harvesting) date were first created using multi-source remote sensing products and reanalysis data, and was then successfully used to simulate the growth and yield for corn and soybean and associated energy exchanges. We also optimized and calibrated other crop parameters using the time-series of the Moderate Resolution Imaging Spectroradiometer (MODIS) land surface products. The modified crop model substantially improved the simulation of crop growth, plant physiology, and biomass accumulation for both corn and soybean. Coupling the localized dynamic crop model into the WRF led to considerable decreases in the simulated mean-absolute-errors (MAEs) and biases of the leaf area index, evapotranspiration, and gross primary production compared with the MODIS observed values. Compared with the statistical yield from each province, the modified crop model underestimated the corn yield from 11.1% to 48.6%, whereas overestimated the soybean yield from 16.5% to 162.6%.

54 ENVIRONMENTAL SCIENCES↗

How Frequent Will the Rarest Daily Rainfall Records of Hurricane Ida’s Remnants Be in the Future?

Abstract Gaining continued insights into the impact of global warming on the occurrence of hurricane-associated intense record downpours is essential for building climate resilient communities. This study investigates projected future changes in extreme rainfall over the Northeast United States, as represented by extreme daily amounts during Hurricane Ida in 2021. We used historical control simulations of Weather Research and Forecasting (WRF) Model generated from 40 years of weather events (1980–2014, 12 km) forced by the fifth generation European Centre for Medium-Range Weather Forecasts atmospheric reanalysis. These simulations are thermodynamically modified (2060–2100) via an imposed warming for the high-emission scenario of shared socioeconomic pathway (SSP585) from a range of general circulation models. Ground observations from the Global Historical Climatology Network (1950–2014) and WRF simulations (historical, 1980–2014, and future, 2060–2100) are integrated into a nonstationary generalized extreme value (GEV) framework to assess the frequency of Ida’s heaviest daily rain rates under the SSP585 scenario. Results show that Ida’s daily maximum rainfall recorded at different observation locations was higher than the single highest September daily maximum observed (1950–2014) for 5 out of 17 stations (∼30% of the stations). Ida-like extreme daily rain rates are projected to be, on average, more than 2 times more likely to occur at the end of the century in the simulations (with some regions as high as 5 times). This work demonstrates that integrating a high-resolution atmospheric model’s present-day and thermodynamically modified future simulations along with ground observations, within a nonstationary statistical framework, is crucial for understanding changing characteristics of extreme weather events. Significance Statement Daily scale extreme precipitation is expected to become more frequent and severe, as evidenced by observations and model simulations. While it is important to investigate how these intensifying heavy rainfall events affect current engineering standards, fewer studies have contextualized how warming impacts the most extreme rainfall from a single storm event relative to historical heavy downpours. In this study, we focused on the daily extreme rainfall associated with the extratropical transition of Hurricane Ida (2021), particularly over the northeastern United States—some of which exceeded the commonly used hydrologic design criteria for a 100-yr storm. Using a high-resolution atmospheric model simulation, we investigated how continued warming may influence the frequency of such daily rain rates. Under a high-emission scenario, these events are projected to become up to 5 times more likely at the end of the twenty-first century.

Dollan, Ishrat J↗

Decadal predictability of North Atlantic blocking and the NAO

Can multi-annual variations in the frequency of North Atlantic atmospheric blocking and mid-latitude circulation regimes be skilfully predicted? Recent advances in seasonal forecasting have shown that mid-latitude climate variability does exhibit significant predictability. However, atmospheric predictability has generally been found to be quite limited on multi-annual timescales. New decadal prediction experiments from NCAR are found to exhibit remarkable skill in reproducing the observed multi-annual variations of wintertime blocking frequency over the North Atlantic and of the North Atlantic Oscillation (NAO) itself. This is partly due to the large ensemble size that allows the predictable component of the atmospheric variability to emerge from the background chaotic component. The predictable atmospheric anomalies represent a forced response to oceanic low-frequency variability that strongly resembles the Atlantic Multi-decadal Variability (AMV), correctly reproduced in the decadal hindcasts thanks to realistic ocean initialization and ocean dynamics. The occurrence of blocking in certain areas of the Euro-Atlantic domain determines the concurrent circulation regime and the phase of known teleconnections, such as the NAO, consequently affecting the stormtrack and the frequency and intensity of extreme weather events. Therefore, skilfully predicting the decadal fluctuations of blocking frequency and the NAO may be used in statistical predictions of near-term climate anomalies, and it provides a strong indication that impactful climate anomalies may also be predictable with improved dynamical models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Contrasting Trends in Colorado Fire Weather Index from Reanalysis and Observations

Recent wildfires in Colorado raise the question of whether rising global temperatures have increased fire weather occurrences in Colorado. The U.S. National Weather Service defines fire weather as when “forecast weather conditions will result in a significant threat for the ignition and/or spread of wildfires.” We use two datasets to address the question: “How has the occurrence of fire weather changed in Colorado?” Using 22 years of observed weather conditions from a meteorological tower at the National Renewable Energy Laboratory and 67 years of ERA5 reanalysis data, we assess changing trends in Colorado fire weather as defined by hot, dry, and windy conditions. Additionally, we explore if the difference in recorded wind speeds between observational data and reanalysis data can be explained by differences in spatial and temporal resolution and what are the implications in the context of quantifying fire weather occurrences. The observational data are limited in temporal extent and spatial representativeness, but they capture exact real-world conditions at a location in complex terrain. The reanalysis data are available for an extended period of time and for the entire state, but the data are of relatively coarse spatial and temporal resolution and may fail to capture extremes. To quantify fire risk, we calculate the hot–dry–windy index (HDWI), which relies on wind speed and vapor pressure deficit. No statistically significant trend in the HDWI appears in the observational dataset. However, according to the reanalysis data, strong increasing trends in HDWI values emerge across all of Colorado. This apparent conflict between observational and reanalysis data suggests that reanalysis data may not be representative. Further, more long-term observational datasets are required to assess fire risk.

17 WIND ENERGY↗

Climate Dynamics Preceding Summer Forest Fires in California and the Extreme Case of 2018

Recent record-breaking wildfire seasons in California prompt an investigation into the climate patterns that typically precede anomalous summer burned forest area. Using burned-area data from the U.S. Forest Service’s Monitoring Trends in Burn Severity (MTBS) product and climate data from the fifth major global reanalysis produced by the European Centre for Medium-Range Weather Forecasts (ERA5) over 1984–2018, relationships between the interannual variability of antecedent climate anomalies and July California burned area are spatially and temporally characterized. Lag correlations show that antecedent high vapor pressure deficit (VPD), high temperatures, frequent extreme high temperature days, low precipitation, high subsidence, high geopotential height, low soil moisture, and low snowpack and snowmelt anomalies all correlate significantly with July California burned area as far back as the January before the fire season. Seasonal regression maps indicate that a global midlatitude atmospheric wave train in late winter is associated with anomalous July California burned area. July 2018, a year with especially high burned area, was to some extent consistent with the general patterns revealed by the regressions: low winter precipitation and high spring VPD preceded the extreme burned area. However, geopotential height anomaly patterns were distinct from those in the regressions. Extreme July heat likely contributed to the extent of the fires ignited that month, even though extreme July temperatures do not historically significantly correlate with July burned area. While the 2018 antecedent climate conditions were typical of a high-burned-area year, they were not extreme, demonstrating the likely limits of statistical prediction of extreme fire seasons and the need for individual case studies of extreme years.

54 ENVIRONMENTAL SCIENCES↗

Advanced Laboratory and Field Arrays (ALFA)/Lab Collaboration Project (LCP) for Marine Energy (Final Scientific/Technical Report)

The objective of the Advanced Laboratory and Field Arrays (ALFA) project was to reduce the Levelized Cost of Energy (LCOE) of Marine and Hydrokinetic (MHK) energy by leveraging research, development, and testing capabilities at Oregon State University, University of Washington, and the University of Alaska, Fairbanks. ALFA is a project within the Pacific Marine Energy Center (PMEC; formerly NNMREC), a multi-institution entity with a diverse funding base that focuses on research and development for marine renewables. The ALFA project aimed to accelerate the development of next-generation arrays of wave energy conversion (WEC) and tidal energy conversion (TEC) devices through a suite of field-focused R&D activities spanning a broad range of strategic opportunity areas identified in the Funding Opportunity Announcement: • Device and/or array operation and maintenance (O&M) logistics development; • High-fidelity resource characterization and/or modeling technique development and validation; • Array-specific component technology development (e.g. moorings and foundations, transmission, and other offshore grid components); • Array performance testing and evaluation; and • Novel cost-effective environmental monitoring techniques and instrumentation testing and evaluation. The objective of the Lab Collaboration Project (LCP) was to accelerate the development of next-generation marine energy conversion systems. The LCP aimed to achieve these project objectives in collaboration with the national laboratories by: • Developing concept generation and assessment tools; • Improving access to existing testing resources; • Validating collision risk models between fish and turbines; and • Advancing analysis and simulation capabilities for wave-WEC interactions and PTO analysis in nonlinear ocean waves. The ALFA portion of the project was comprised of six overarching technical tasks: • Task 1: Debris Modeling, Detection and Mitigation; • Task 2: Autonomous Monitoring & Intervention; • Task 3: Resource Characterization for Extreme Conditions; • Task 4: Robust Models for Design of Offshore Anchoring and Mooring Systems; • Task 5: Performance Enhancement for Marine Energy Converter (MEC) Arrays; and • Task 6: Evaluating Sampling Techniques for MHK Biological Monitoring. The LCP was divided into four overarching technical tasks: • Task 7: Project Management and Reporting • Task 8: Novel Design and Assessment Methodologies for Wave Energy Converter Design (Wave- SPARC) • Task 9: Testing Access for Commercial Marine Renewable Energy Technology Developers • Task 10: Quantifying Collision Risk for Fish and Turbines • Task 11: Nonlinear Ocean Waves and PTO Control Strategy Each ALFA/LCP task listed above functioned as a separate and discreet project. A final Technical Report was written for each individual task and these reports were uploaded to OSTI, after receiving DOE approval. The following document is a compilation of each of these final, approved reports arranged as individual chapters.

13 HYDRO ENERGY↗

Confronting the Challenge of Modeling Cloud and Precipitation Microphysics

In the atmosphere, microphysics refers to the microscale processes that affect cloud and precipitation particles and is a key linkage among the various components of Earth’s atmospheric water and energy cycles. The representation of microphysical processes in models continues to pose a major challenge leading to uncertainty in numerical weather forecasts and climate simulations. In this paper, the problem of treating microphysics in models is divided into two parts: i) how to represent the population of cloud and precipitation particles, given the impossibility of simulating all particles individually within a cloud, and ii) uncertainties in the microphysical process rates owing to fundamental gaps in knowledge of cloud physics. The recently-developed Lagrangian particle-based method is advocated as a way to address several conceptual and practical challenges of representing particle populations using traditional bulk and bin microphysics parameterization schemes. For addressing critical gaps in cloud physics knowledge, sustained investment for observational advances from laboratory experiments, new probe development, and next-generation instruments in space is needed. Greater emphasis on laboratory work, which has apparently declined over the past several decades relative to other areas of cloud physics research, is argued to be an essential ingredient for improving process-level understanding. More systematic use of natural cloud and precipitation observations to constrain microphysics schemes is also advocated. Because it is generally difficult to quantify individual microphysical process rates from these observations directly, this presents an inverse problem that can be viewed from the standpoint of Bayesian statistics. Following this idea, a probabilistic framework is proposed that combines elements from statistical and physical modeling. Besides providing rigorous constraint of schemes, there is an added benefit of quantifying uncertainty systematically. Finally, a broader hierarchical approach is proposed to accelerate improvements in microphysics schemes, leveraging the advances described in this paper related to process modeling (using Lagrangian particle-based schemes), laboratory experimentation, cloud and precipitation observations, and statistical methods.

54 ENVIRONMENTAL SCIENCES↗

Stochastic Price Generation for Evaluating Wholesale Electricity Market Bidding Strategies

This work presents a novel method for generating electricity price scenarios from statistical properties of past electricity prices using a hybrid statistical and reduced-form stochastic model. Previous work in applying stochastic differential equations (SDE) to model electricity prices has focused on daily average prices. To extend stochastic price generation methods to hourly or sub-hourly pricing, we address several weaknesses in the state-of-the-art: (1) we replace the mean-reversion component of the SDE with an ARIMA process that is better able to characterize the daily and weekly trends; (2) we extend the price-spike, or jump process to account for conditional probabilities of price spikes occurring in consecutive time steps by replacing the traditional Poisson process for modeling jumps with a generalized point process model inspired by brain neuron models; and (3) we replace the traditional method of estimating spike intensity with empirical variance with a Markov process based on observed price spike intensity transitions. The method is demonstrated with electricity prices from the US ERCOT market and a use-case example is provided for bidding an energy storage unit into the day-ahead and real-time energy markets of ERCOT using stochastic optimization methods. Results show that the the synthetic price model out performs a (naive) persistence forecast model by resulting in 24% to 47% more in profits over 168 simulated days.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Atmospheric condition identification in multivariate data through a metric for total variation

Identification of atmospheric conditions within a multivariable atmospheric data set is a necessary step in the validation of emerging and existing high-fidelity models used to simulate wind plant flows and operation.Atmospheric conditions relevant for wind energy research include stationary conditions, given the need for well-converged statistics for model validation, as well as conditions observed less frequently, such as extreme atmospheric events, which are used in wind turbine and wind plant design.Aggregation of observations without regard to covariance between time series discounts the dynamical nature of the atmosphere and is not sufficiently representative of atmospheric conditions.Identification and characterization of continuous time periods with atmospheric conditions that have a high value for analysis or simulation set the stage for more advanced model validation and the development of real-time control and operational strategies.The current work explores a single metric for variation in a multivariate data sample that quantifies variability within each channel as well as covariance between channels.The total variation is used to identify conditions of interest that conform to desired objective functions, such as stationary conditions, ramps or waves of wind speed, and changes in wind direction.Total variation is somewhat sensitive to the presence of outliers in the input data, and the method is best complemented by quality-control procedures to ensure reliable results.The direct detection and classification of events or conditions of interest within atmospheric data sets is vital to developing our understanding of wind plant response and to the formulation of forecasting and control models.

17 WIND ENERGY↗

Hemispherical variance anomaly and reionization optical depth

ABSTRACT Cosmic microwave background (CMB) full-sky temperature data show a hemispherical asymmetry in power nearly aligned with the Ecliptic, with the Northern hemisphere displaying an anomalously low variance, while the Southern hemisphere appears consistent with expectations from the best-fitting theory, Lambda Cold Dark Matter (ΛCDM). The low signal-to-noise ratio in current polarization data prevents a similar comparison. Polarization realizations constrained by temperature data show that in ΛCDM the lack of variance is not expected to be present in polarization data. Therefore, a natural way of testing whether the temperature result is a fluke is to measure the variance of CMB polarization components. In anticipation of future CMB experiments that will allow for high-precision large-scale polarization measurements, we study how the variance of polarization depends on ΛCDM-parameter uncertainties by forecasting polarization maps with Planck’s Markov chain Monte Carlo chains. We show that polarization variance is sensitive to present uncertainties in cosmological parameters, mainly due to current poor constraints on the reionization optical depth τ, which drives variance at low multipoles. We demonstrate how the improvement in the τ measurement seen between Planck’s two latest data releases results in a tighter constraint on polarization variance expectations. Finally, we consider even smaller uncertainties on τ and how more precise measurements of τ can drive the expectation for polarization variance in a hemisphere close to that of the cosmic-variance-limited distribution.

79 ASTRONOMY AND ASTROPHYSICS↗

Dynamically Downscaled Hourly Future Weather Data with 12-km Resolution Covering Most of North America

This is an hourly future weather dataset for energy modeling applications. The dataset is primarily based on the output of a regional climate model (RCM), i.e., the Weather Research and Forecasting (WRF) model version 3.3.1. The WRF simulations are driven by the output of a general circulation model (GCM), i.e., the Community Climate System Model version 4 (CCSM4). This dataset is in the EPW format, which can be read or translated by more than 25 building energy modeling programs (e.g., EnergyPlus, ESP-r, and IESVE), energy system modeling programs (e.g., System Advisor Model (SAM)), indoor air quality analysis programs (e.g., CONTAM), and hygrothermal analysis programs (e.g., WUFI). It contains 13 weather variables, which are the Dry-Bulb Temperature, Dew Point Temperature, Relative Humidity, Atmospheric Pressure, Horizontal Infrared Radiation Intensity from Sky, Global Horizontal Irradiation, Direct Normal Irradiation, Diffuse Horizontal Irradiation, Wind Speed, Wind Direction, Sky Cover, Albedo, and Liquid Precipitation Depth. The weather data is created for two emissions scenarios: RCP4.5 and RCP8.5 and spans two 10-year time slices in the future: 2045 - 2054 and 2085 - 2094. It offers a spatial resolution of 12 km by 12 km with extensive coverage across most of North America. Due to the enormous size of the entire dataset, in the first stage of its distribution, we provide 20 years of future weather data for the centroid of each Public Use Microdata Area (PUMA), excluding Hawaii. PUMAs are non-overlapping, statistical geographic areas that partition each state or equivalent entity into geographic areas containing no fewer than 100,000 people each. The 2,378 PUMAs as a whole cover the entirety of the U.S. The weather data can be utilized alongside the large-scale energy analysis tools, ResStock and ComStock, developed by National Renewable Energy Laboratory, whose smallest resolution is at the PUMA scale. The data for RCP4.5 is still being processed and will be published soon.

Array↗

A new metrics framework for quantifying and intercomparing atmospheric rivers in observations, reanalyses, and climate models

We present a new atmospheric river (AR) analysis and benchmarking tool, namely Atmospheric River Metrics Package (ARMP). It includes a suite of new AR metrics that are designed for quick analysis of AR characteristics via statistics in gridded climate datasets such as model output and reanalysis. This package can be used for climate model evaluation in comparison with reanalysis and observational products. Integrated metrics such as mean bias and spatial pattern correlation are efficient for diagnosing systematic AR biases in climate models. For example, the package identifies the fact that, in CMIP5 and CMIP6 (Coupled Model Intercomparison Project Phases 5 and 6) models, AR tracks in the South Atlantic are positioned farther poleward compared to ERA5 reanalysis, while in the South Pacific, tracks are generally biased towards the Equator. For the landfalling AR peak season, we find that most climate models simulate a completely opposite seasonal cycle over western Africa. This tool can also be used for identifying and characterizing structural differences among different AR detectors (ARDTs). For example, ARs detected with the Mundhenk algorithm exhibit systematically larger size, width, and length compared to the TempestExtremes (TE) method. The AR metrics developed from this work can be routinely applied for model benchmarking and during the development cycle to trace performance evolution across model versions or generations and set objective targets for the improvement of models. They can also be used by operational centers to perform near-real-time climate and extreme event impact assessments as part of their forecast cycle.

58 GEOSCIENCES↗

Quantifying microbial control of soil organic matter dynamics at macrosystem scales

Soil organic matter (SOM) stocks, decomposition and persistence are largely the product of controls that act locally. Yet the controls are shaped and interact at multiple spatiotemporal scales, from which macrosystem patterns in SOM emerge. Theory on SOM turnover recognizes the resulting spatial and temporal conditionality in the effect sizes of controls that play out across macrosystems, and couples them through evolutionary and community assembly processes. For example, climate history shapes plant functional traits, which in turn interact with contemporary climate to influence SOM dynamics. Selection and assembly also shape the functional traits of soil decomposer communities, but it is less clear how in turn these traits influence temporal macrosystem patterns in SOM turnover. Here, we review evidence that establishes the expectation that selection and assembly should generate decomposer communities across macrosystems that have distinct functional effects on SOM dynamics. Representation of this knowledge in soil biogeochemical models affects the magnitude and direction of projected SOM responses under global change. Yet there is high uncertainty and low confidence in these projections. To address these issues, we make the case that a coordinated set of empirical practices are required which necessitate (1) greater use of statistical approaches in biogeochemistry that are suited to causative inference; (2) long-term, macrosystem-scale, observational and experimental networks to reveal conditionality in effect sizes, and embedded correlation, in controls on SOM turnover; and (3) use of multiple measurement grains to capture local- and macroscale variation in controls and outcomes, to avoid obscuring causative understanding through data aggregation. Here, when employed together, along with process-based models to synthesize knowledge and guide further empirical work, we believe these practices will rapidly advance understanding of microbial controls on SOM and improve carbon cycle projections that guide policies on climate adaptation and mitigation.

59 BASIC BIOLOGICAL SCIENCES↗

Divide and conquer: Learning chaotic dynamical systems with multistep penalty neural ordinary differential equations

Forecasting high-dimensional dynamical systems is a fundamental challenge in various fields, such as geosciences and engineering. Neural Ordinary Differential Equations (NODEs), which combine the power of neural networks and numerical solvers, have emerged as a promising algorithm for forecasting complex nonlinear dynamical systems. However, classical techniques used for NODE training are ineffective for learning chaotic dynamical systems. In this work, we propose a novel NODE-training approach that allows for robust learning of chaotic dynamical systems. Here, our method addresses the challenges of non-convexity and exploding gradients associated with underlying chaotic dynamics. Training data trajectories from such systems are split into multiple, non-overlapping time windows. In addition to the deviation from the training data, the optimization loss term further penalizes the discontinuities of the predicted trajectory between the time windows. The window size is selected based on the fastest Lyapunov time scale of the system. Multi-step penalty(MP) method is first demonstrated on Lorenz equation, to illustrate how it improves the loss landscape and thereby accelerates the optimization convergence. MP method can optimize chaotic systems in a manner similar to least-squares shadowing with significantly lower computational costs. Our proposed algorithm, denoted the Multistep Penalty NODE, is applied to chaotic systems such as the Kuramoto-Sivashinsky equation, the two-dimensional Kolmogorov flow, and ERA5 reanalysis data for the atmosphere. It is observed that MP-NODE provide viable performance for such chaotic systems, not only for short-term trajectory predictions but also for invariant statistics that are hallmarks of the chaotic nature of these dynamics.

Chaotic dynamical systems↗

A Machine‐Learning‐Assisted Stochastic Cloud Population Model as a Parameterization of Cumulus Convection

Abstract A machine‐learning‐assisted stochastic cloud population model is coupled with the Advanced Research Weather Research and Forecasting (WRF) model to represent fluctuations in the cloud‐base mass flux associated with the life cycles and interactions among cumulus convection cells. In this cloud population model, the size distribution and the associated cloud‐base mass flux of the convective cells are related to their previous state and to the change in the total convective area via a transition function. The convective area tendency in turn is assumed to depend on the cloud‐base mass flux that is resolved by the host WRF model. The transition function is represented by a single hidden‐layer neural network trained by the evolution of convective cell size distributions in a 1‐km grid‐spacing WRF simulation run over the Australian Monsoon region. At every grid point of the host model, the cloud population model predicts the cell size and cloud‐base mass flux distributions from which a random sample of cells is fed to an entraining parcel model that calculates precipitation as well as the associated liquid water potential temperature and total moisture tendencies. These tendencies are averaged over the cells and provided to the host model. Several regional simulations are performed over tropical and midlatitude domains to test this as a potential approach to scale‐aware parameterization. It is shown that such an approach could be a new promising path to simulating realistic precipitation statistics and propagation of precipitation associated with the Madden‐Julian Oscillation while maintaining realistic depictions of the diurnal cycle over both land and ocean.

54 ENVIRONMENTAL SCIENCES↗