Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model skill”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Hierarchical Testing of a Hybrid Machine Learning‐Physics Global Atmosphere Model

Machine learning (ML)-based models have demonstrated high skill and computational efficiency, often outperforming conventional physics-based models in weather and subseasonal predictions. While prior studies have assessed their fidelity in capturing synoptic-scale atmospheric dynamics, their performance across timescales and under out-of-distribution forcing, such as +3K or +4K uniform-warming forcings, and the sources of biases remain elusive, to establish the model's reliability for Earth science. Here, we design three sets of experiments targeting synoptic-scale phenomena, interannual variability, and out-of-distribution uniform-warming forcings. We evaluate the Neural General Circulation Model (NeuralGCM), a hybrid model integrating a dynamical core with ML-based component, against observations and physics-based Earth system models (ESMs). At the synoptic scale, NeuralGCM captures the evolution and propagation of extratropical cyclones with performance comparable to ESMs. At the interannual scale, when forced by El Niño-Southern Oscillation sea surface temperature (SST) anomalies, NeuralGCM successfully reproduces associated teleconnection patterns but exhibits deficiencies in capturing nonlinear response. Under out-of-distribution uniform-warming forcings, NeuralGCM simulates similar responses in global-average temperature and precipitation and reproduces large-scale tropospheric circulation features similar to those in ESMs. Notable weaknesses include overestimating the tracks and spatial extent of extratropical cyclones, biases in the teleconnected wave train triggered by tropical SST anomalies, and differences in upper-level warming and stratospheric circulation responses to SST warming compared to physics-based ESMs. The causes of these weaknesses were explored. Despite the noted weaknesses, NeuralGCM reproduces responses across experiments reasonably and performs comparably to ESMs. By integrating a dynamical core with ML, NeuralGCM shows potential for developing ML-based ESMs.

global warming↗

“Godzilla,” the Extreme African Dust Event of June 2020: Origins, Transport, and Impact on Air Quality in the Greater Caribbean Basin

In June 2020, the tropical Atlantic and the Caribbean Basin were affected by a series of African dust outbreaks unprecedented in size and intensity. These events, informally named “Godzilla,” coincided with CALIMA, a large field campaign, offering a rare opportunity to assess the impact of African dust on air quality in the Greater Caribbean Basin. Network measurements of respirable particles (i.e., PM 10 and PM 2.5 ) showed that dust significantly degraded regional air quality and increased the risk to public health in the Caribbean, the southern United States, northern South America, and Central America. CALIMA examined the meteorological context of Godzilla dust events over North Africa and how these conditions might relate to the greatly increased dust emissions and enhanced transport to the Americas. Godzilla was linked to strong pressure anomalies over West Africa, resulting in a large-scale geostrophic wind anomaly at 700 hPa over North Africa. We used surface-based and columnar measurements to test the performance of two frequently used aerosol forecast models: the NASA Goddard Earth Observing System (GEOS) and Weather Research and Forecasting Model coupled with Chemistry (WRF-Chem) models. The models showed some skills but differed substantially between their forecasts, suggesting large uncertainties in these forecasts that are critical for issuing early warnings of health-threatening dust events. Our results demonstrate the value of an integrated approach in characterizing the spatial and temporal variability of African dust transport and assessing its impact on regional air quality. Future studies are needed to improve models and to track the long-term changes in dust transport from Africa under a changing climate.

Aerosols/particulates↗

Relating flow resistance to equivalent roughness

Describing flow resistance using the physical properties of an underlying surface is a recalcitrant problem in overland flow models. If discharge measurements are available, an equivalent roughness (e.g., Manning’s n) can be calibrated to represent the effects of surface properties within the domain with a single numerical value. Alternatively, the flow resistance can be estimated from discharge and velocity measured at a point, typically a runoff plot outlet. However, such experimental estimates are often inconsistent with the equivalent roughness determined from calibration to discharge, even if both derive from the same dataset. For example, if Manning’s equation is used to parameterize flow resistance, the Manning’s n obtained by calibrating a model to discharge differs from the value of n calculated from measured flow and velocity at the hillslope outlet. Here, this discrepancy is resolved by deriving a correction factor relating experimentally-determined flow resistance to the equivalent roughness. The derived correction factor is tested for four commonly-used resistance formulations using 129 rainfall simulator experiments. The correction factor is necessary to reproduce measured velocities, and yields minor improvements in discharge prediction. Plain Language Summary: Accurate runoff prediction is needed for land and water management in dryland regions, where sporadic and limited rainfall necessitate efficient water use and drought mitigation strategies. The skill of runoff models is known to be hindered by out ability to estimate flow resistance, which is the quantity that describes how energy is lost from flowing water to the underlying surface. Typically, models represent flow resistance with an equivalent roughness, e.g., Manning’s n, that is adjusted until the model can reproduce available discharge observations at watershed scale. However, the flow resistance measured in plot-scale experiments (1–10 m) often exceeds equivalent roughness coefficients by a factor of 10. This means that the direct use of plot-scale experimental data to parameterize runoff models could cause errors in discharge and runoff velocity predictions. Here, we resolve these differences by deriving an analytic correction factor that relates flow resistance to the equivalent roughness required for models to reproduce experimental velocity and discharge data. This correction factor is tested using rainfall simulator data from 129 experiments performed in the US Southwest covering a wide range of precipitation intensities, soil textures and vegetation types. Use of the correction factor substantially improves model prediction of flow velocity, which is needed for reproducing the timing of flood events and the estimation of erosion.

54 ENVIRONMENTAL SCIENCES↗

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗

Complementing Dynamical Downscaling With Super‐Resolution Convolutional Neural Networks

Despite advancements in Artificial Intelligence (AI) methods for climate downscaling, significant challenges remain for their practicality in climate research. Current AI-methods exhibit notable limitations, such as limited application in downscaling Global Climate Models (GCMs), and accurately representing extremes. To address these challenges, we implement an AI-based methodology using super-resolution convolutional neural networks (SRCNN), trained and evaluated on 40 years of daily precipitation data from a reanalysis and a high-resolution dynamically downscaled counterpart. The dynamical downscaled simulations, constrained using spectral nudging, enable the replication of historical events at a higher resolution. This allows the SRCNN to emulate dynamical downscaling effectively. Modifications, such as incorporating elevation data and data pre-processing enhances overall model performance, while using exponential and quantile loss functions improve the simulation of extremes. Our findings show SRCNN models efficiently and skillfully downscale precipitation from GCMs. Future work will expand this methodology to downscale additional variables for future climate projections.

54 ENVIRONMENTAL SCIENCES↗

Incorrect computation of Madden-Julian oscillation prediction skill

The Madden–Julian oscillation (MJO) is a major tropical weather system and one of the largest sources of predictability for subseasonal-to-seasonal weather forecasts. Skillful prediction of the MJO has been a highly active area of research due to its large socio-economic impacts. Silini et al., herein S21, developed a machine learning model to predict the MJO, which they claimed to have an MJO prediction skill of 26–27 days over all seasons and 45 days for December–February (DJF) winter. If true, this would make the skill of their model competitive with that of the state-of-the-art dynamical MJO prediction systems at 20–35 days. However, here we show that the MJO prediction was calculated incorrectly in S21, which spuriously increased the performance of their model. Correctly computed skill of their model was substantially lower than that reported in S21; the skill for all seasons drops to 11–12 days and the skill for forecasts initialized during DJF drops to 15 days. Our findings clarify that the S21 machine learning model is not competitive with state-of-the-art numerical weather prediction models in predicting the MJO.

54 ENVIRONMENTAL SCIENCES↗

A Practical Probabilistic Benchmark for AI Weather Models

Since the weather is chaotic, it is necessary to forecast an ensemble of future states. Recently, multiple AI weather models have emerged claiming breakthroughs in deterministic skill. Unfortunately, it is hard to fairly compare ensembles of AI forecasts because variations in ensembling methodology become confounding and the baseline data volume is immense. We address this by scoring lagged initial condition ensembles—whereby an ensemble can be constructed from a library of deterministic hindcasts. This allows the first parameter‐free intercomparison of leading AI weather models' probabilistic skill against an operational baseline. Lagged ensembles of the two leading AI weather models, GraphCast and Pangu, perform similarly even though the former outperforms the latter in deterministic scoring. These results are elaborated upon by sensitivity tests showing that commonly used multiple time‐step loss functions damage ensemble calibration.

54 ENVIRONMENTAL SCIENCES↗

Leveraging High-resolution Molecular Composition of Soil Organic Matter to Enhance Carbon Cycling Modeling

Soils store more carbon than the atmosphere and vegetation combined, yet Earth system models still struggle to predict how this vast reservoir will respond to environmental change. A central limitation is that most soil biogeochemical models represent organic matter using bulk conceptual pools or chemically homogeneous fractions, preventing direct use of rapidly expanding molecular-scale datasets. Here we develop and test a new soil decomposition framework that explicitly integrates high-resolution information on organic matter composition. First, we construct a molecularly informed litter decomposition module in which plant inputs are partitioned into five functional compound classes—carbohydrates, proteins, lignin-like aromatics, lipids, and carbonyls—using a molecular mixing model calibrated to solid-state 13 C Nuclear Magnetic Resonance (NMR) spectra. Class-specific kinetics, lignin-dependent physical protection, and substrate-driven microbial carbon use efficiency allow the module to capture metabolic tradeoffs associated with enzyme production and nutrient limitation. We then embed this litter module within a microbially explicit whole-soil model that tracks the transformation of these compound classes through particulate organic matter, dissolved organic matter, mineral-associated organic matter, and microbial biomass. High-resolution Fourier Transform Ion Cyclotron Resonance mass spectrometry (FTICR-MS) data are used to link internal pools to measurable soil organic matter fractions and to constrain key process parameters. Applications at soil-core and ecosystem scales demonstrate that the new model reproduces observed soil respiration dynamics while providing mechanistic attribution of CO 2 fluxes to specific chemical classes and pools. Compared to existing frameworks such as the Community Land Model soil biogeochemistry module and the Millennial model, our approach maintains competitive predictive skill while substantially improving interpretability and opportunities for data–model integration. This work illustrates a viable pathway for leveraging molecular-scale observations to reduce structural uncertainty in soil carbon–climate feedback projections.

54 ENVIRONMENTAL SCIENCES↗

Quantifying Uncertainties in Earth's Energy Budget by Cloud Feedback and Ocean Heat Uptake Using E3SM-Slab Ocean Configurations

In order to improve predictive skills of Earth System Model, we need to better understand processes that control Earth's energy budget via ocean, atmosphere, and cryosphere interactions. Simulated energy budget in comprehensive Earth System Models shows a wide range, leading to large uncertainties in predicting Earth system dynamic and thermodynamic variations and associated social-economic impacts. Uncertainties in cloud feedbacks have been identified as the main cause of the large inter-model spread, but oceanic adjustments, especially those associated with ocean heat uptake (OHU) and the Atlantic Meridional Overturning Circulation (AMOC), also play an important role. In this proposed work, we focus on understanding the individual and combined roles of cloud feedbacks and ocean adjustments on modulating Earth's energy balance. This research is motivated by our overarching hypothesis that oceanic adjustment is a key source of uncertainty, in addition to those associated with the cloud feedbacks; further, the ocean adjustment and associated OHU work through the cloud feedbacks to modulate Earth's energy budget and temperature variations. We test this hypothesis using numerical experiments where we systematically enable and disable cloud feedbacks in conjunction with perturbations to OHU.

58 GEOSCIENCES↗

Machine learning methods for weather forecasting

SAND2025-14466O This repository contains code for developing, training, and evaluating machine learning models for weather and climate forecasting, including forecast skill assessment, feature importance analysis, and reproducible workflows for model comparison. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Holthuijzen, Maike [Sandia National Lab. (SNL-CA),↗

Analyzing and Exploring Training Recipes for Large-Scale Transformer-Based Weather Prediction

Abstract The rapid rise of deep learning (DL) in numerical weather prediction (NWP) has led to a proliferation of models which forecast atmospheric variables with comparable or superior skill than traditional physics-based NWP. However, among these leading DL models, there is a wide variance in both the training settings and architecture used. Further, the lack of thorough ablation studies makes it hard to discern which components are most critical to success. In this work, we show that it is possible to attain high forecast skill even with relatively off-the-shelf architectures, simple training procedures, and moderate compute budgets. Specifically, we train a minimally modified Swin Transformer V2 (SwinV2) on ERA5 data and find that it attains superior skill in terms of mean-square errors of deterministic forecasts when compared against the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS). Almost all DL–NWP systems share a core set of hyperparameters and design decisions. To aid and expedite future DL–NWP research, we present an in-depth, systematic exploration of different loss functions, model sizes and depths, patch sizes, and multistep training objectives. We also examine the model performance with metrics beyond the typical accuracy (ACC) and RMSE and investigate how the performance scales with model size. Through our open-source code, scoring pipelines, and models, we share our findings on key aspects of the training pipeline. These ablations reduce the necessity for expensive hyperparameter tuning and lower the barrier to entry for future DL–NWP research. Significance Statement This study investigates the potential of using large-scale transformer-based models for weather prediction, showing that it is possible to achieve high forecast accuracy with simpler, off-the-shelf architectures. By training a minimally modified SwinV2 transformer on ERA5 data, we show that the model achieves competitive forecast skill in terms of mean-square error for key variables, outperforming the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS) at all lead times. Our findings suggest that effective training strategies, such as multistep fine-tuning and channel-weighted losses, significantly enhance the model’s performance. However, we also highlight that these improvements come with trade-offs in other areas, such as ensemble spread and high-frequency spatial detail. This work highlights the promise of deep learning in improving weather forecasts, which could lead to better preparedness and response to weather events, ultimately benefiting society by providing more reliable weather predictions.

Willard, Jared D. [Lawrence Berkeley National Labo↗

Evaluating mesoscale model predictions of diurnal speedup events in the Altamont Pass Wind Resource Area of California

Mesoscale model predictions of wind, turbulence, and wind energy capacity factors are evaluated in the Altamont Pass Wind Resource Area of California (APWRA), where the diurnal regional sea breeze and associated terrain-driven speedup flows drive wind energy production during the summer months. Results from the Weather Research and Forecasting model version 4.4 using a novel three-dimensional planetary boundary layer (3D PBL) scheme, which treats both vertical and horizontal turbulent mixing, are compared to those using a well-established one-dimensional (1D) scheme that treats only vertical turbulent mixing. Each configuration is evaluated over a nearly 3-month-long period during the Hill Flow Study, and due to the recurring nature of the observed speedup flows, diurnal composite averaging is used to capture robust trends in model performance. Both model configurations showed similar overall skill. The general timing and direction of the speedup flows is captured, but their magnitude is overestimated within a typical wind turbine rotor layer. Both also fail to capture a persistent observed near-surface jet-like flow, likely due to the limited grid resolution that is typical of mesoscale models. However, the 3D PBL configuration shows several minor improvements over the 1D PBL configuration, including improved wind speed and turbulence kinetic energy profiles during the accelerating phase of the speedup events, as well as reduced positive wind speed bias at surface stations across the APWRA region. Using a mesoscale wind farm parameterization, modeled capacity factors are also compared to monthly data reported to the US Energy Information Administration (EIA) during the study period. Although the monthly trend in the data is captured, both model configurations overestimate capacity factors by roughly 7 %–11 %. Through model evaluation, this study provides confidence in the 3D PBL scheme for wind energy applications in complex terrain and provides guidance for future testing.

17 WIND ENERGY↗

Simulating Hurricane Katrina in the Simple Cloud‐Resolving E3SM Atmosphere Model v1

Climate models are important tools for advancing understanding and prediction of tropical cyclones (TCs). Traditional global climate models, however, do not have the ability to properly simulate TC intensity due to their coarse horizontal resolution. Regional models can be run at convection‐permitting resolutions, but these models are often strongly influenced by the data used in the lateral boundary forcing, and domain choice can have a large impact on the simulation. Cloud‐resolving global climate models have demonstrated great potential for realism in TC simulations, and in this study we focus specifically on the Simple Cloud‐Resolving Energy Exascale Earth System Model (E3SM) Atmosphere Model (SCREAM) v1 configuration. We evaluate SCREAMv1 against the observational record and the Weather Research and Forecasting (WRF) model run at a convection‐permitting resolution with Hurricane Katrina as our case study. We found that both models produced realistic simulations of Hurricane Katrina. SCREAMv1 demonstrated skill in simulating TC track, size, and intensity, while the model produced an excessive amount of precipitation. In comparison, WRF more accurately simulated TC precipitation and intensity, although the TC wind extent was smaller than the observations.

54 ENVIRONMENTAL SCIENCES↗

Continental-Scale Controls on Hyporheic Respiration Revealed by Knowledge-Guided Machine Learning

Hyporheic zone sediments regulate organic matter turnover and in-stream respiration, yet controls on sediment respiration remain poorly constrained across heterogeneous river networks, limiting prediction of stream metabolism and carbon processing at continental scales. Here, we integrate observations from ~90 river corridors across the United States in the WHONDRS consortium with a knowledge-guided machine learning (KGML) framework that couples thermodynamic rate theory with machine learning to identify dominant controls on hyporheic respiration. Diagnostic analyses show that organic matter concentration and thermodynamic favorability define an upper bound on respiration potential, whereas biological catalytic capacity and physical accessibility jointly govern realized respiration rates through interaction effects. To represent unmeasurable accessibility constraints, we use the mechanistic model as a scaffold for KGML, allowing machine learning to target residual structure not explained by process theory. This hybrid framework improves predictive skill relative to both the mechanistic model alone and fully data-driven models while preserving interpretability. These results indicate that variability in hyporheic respiration is largely mechanistically structured and demonstrate how integrating process theory with explainable AI enhances predictive performance while enabling scalable synthesis of river corridor observations.

Zheng, Jianqiu↗

Educational Consortium for Energy-related Data Science & Computation in Building Engineering Programs

The project spearheaded by Pennsylvania State University aims to address the growing need for integrating energy-focused computation and data science into building engineering education. As the demand for energy-efficient building designs and operations increases, the educational sector must adapt to equip future engineers with the necessary skills. This initiative responds to this need by developing a consortium that unites multiple institutions to enhance curriculum development, dataset curation, and resource sharing, thereby ensuring students are well-prepared for the evolving energy sector. The primary goal of the project is to establish a consortium that will develop and disseminate educational materials and training programs focused on energy-related data science and computation. Key accomplishments include the creation of a beta website for resource sharing, the development of training programs and standalone modules, and the curation of datasets accessible to the public. This effort will culminate in a curriculum that incorporates advanced modeling technologies and data science skills into building engineering programs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Evaluating E3SM Global Storm‐Resolving Model Simulations of Deep Convection: Insights From DP‐SCREAM During TRACER

Global Storm-Resolving Models (GSRMs) are becoming increasingly vital for advancing climate modeling and improving the prediction of extreme weather events. Houston, a coastal region frequently affected by deep convective storms, offers an ideal setting to evaluate the ability of GSRMs to simulate deep convection. This study assesses the performance of the Doubly Periodic Simple Cloud-Resolving E3SM (Energy Exascale Earth System Model) Atmosphere Model (DP-SCREAM) using observations from the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign. DP-SCREAM effectively reproduces the diurnal cycles of clouds and precipitation, demonstrating much greater skill than the E3SM single column model. The DP-SCREAM is demonstrated to be applicable to coastal regions, partially due to the forcing data sets already capturing the influence of breezes. DP-SCREAM also replicates biases persistent in the global version of SCREAM: the underrepresentation of boundary layer shallow clouds, a lack of mid-level congestus clouds, and the popcorn convection, characterized by small and disorganized convective cells generating the strongest precipitation. To investigate these issues, two sensitivity experiments were conducted: increasing the mixing length and scaling up the buoyancy flux within the Simplified Higher Order Closure scheme. Increasing the mixing length improved mid-level congestus representation and reduced unrealistic early morning fog occurrence. Enhancing buoyancy flux only marginally improved the bias of underproduced big convective cells. In conclusion, an additional resolution sensitivity test at 0.5 km grid spacing demonstrated that a refined horizontal resolution alone is insufficient to resolve these biases.

54 ENVIRONMENTAL SCIENCES↗

Linking the subseasonal variability of the East Asia winter monsoon and the Madden-Julian Oscillation through wave disturbances along the subtropical jet

Despite an urgent demand for reliable subseasonal-to-seasonal (S2S) predictions to guide disaster preparedness, our current climate models show limited S2S prediction skill, particularly for precipitation, due to an inadequate understanding of the key processes that drive regional S2S variability. Here we demonstrate that the leading subseasonal variability mode of precipitation over the East Asian Winter Monsoon (EAWM) region is not only closely tied to the activity of the Madden-Julian Oscillation (MJO), but also linked to precipitation and temperature extremes worldwide, influenced by a circumglobal Rossby wave-train along the subtropical westerly jet. Despite a close phase-lock relationship between the MJO and subseasonal EAWM precipitation, our findings indicate that the MJO itself may only play a minor role in the subseasonal EAWM variability. Given its significant impact on the S2S variability of global weather extremes, we call for coordinated community efforts to enhance the understanding and prediction of the circumglobal Rossby wave-train.

Atmospheric science↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗