Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

The ASHRAE Great Energy Predictor III competition: Overview and results

In late 2019, ASHRAE hosted the Great Energy Predictor III (GEPIII) machine learning competition on the Kaggle platform. This launch marked the third energy prediction competition from ASHRAE and the first since the mid-1990s. In this updated version, the competitors were provided with over 20 million points of training data from 2,380 energy meters collected for 1,448 buildings from 16 sources. This competition’s overall objective was to find the most accurate modeling solutions for the prediction of over 41 million private and public test data points. Furthermore, the competition had 4,370 participants, split across 3,614 teams from 94 countries who submitted 39,403 predictions. In addition to the top five winning workflows, the competitors publicly shared 415 reproducible online machine learning workflow examples (notebooks), including over 40 additional, full solutions. This paper gives a high-level overview of the competition preparation and dataset, competitors and their discussions, machine learning workflows and models generated, winners and their submissions, discussion of lessons learned, and competition outputs and next steps. The most popular and accurate machine learning workflows used large ensembles of mostly gradient boosting tree models, such as LightGBM. Similar to the first predictor competition, preprocessing of the data sets emerged as a key differentiator.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A benchmark dataset for Hydrogen Combustion

The generation of reference data for deep learning models is challenging for reactive systems, and more so for combustion reactions due to the extreme conditions that create radical species and alternative spin states during the combustion process. Here, we extend intrinsic reaction coordinate (IRC) calculations with ab initio MD simulations and normal mode displacement calculations to more extensively cover the potential energy surface for 19 reaction channels for hydrogen combustion. A total of ~290,000 potential energies and ~1,270,000 nuclear force vectors are evaluated with a high quality range-separated hybrid density functional, ωB97X-V, to construct the reference data set, including transition state ensembles, for the deep learning models to study hydrogen combustion reaction.

08 HYDROGEN↗

Evaluating the Limits of QAOA Parameter Transfer at High-Rounds on Sparse Ising Models With Geometrically Local Cubic Terms

The emergent practical applicability of the Quantum Approximate Optimization Algorithm (QAOA) for approximate combinatorial optimization is a subject of considerable interest. One of the primary limitations of QAOA is the task of finding a set of good parameters, which is usually done using a variational optimization loop. Parameter transfer, or parameter concentration, is a phenomenon where QAOA angles trained on problem instances that are self-similar tend to perform well for other problem instances from that similar class. This suggests a potentially highly efficient and scalable non-variational learning method for QAOA angle finding. In this work, we systematically study QAOA parameter transferability from small problem sizes (16 and 27 decision variables) onto large problem instances (up to 156 qubits) for heavy-hex graph Ising models with geometrically local higher order terms using the Julia based QAOA simulation tool \texttt{JuliQAOA} to perform classical angle finding for up to $49$ QAOA layers ($p$). Parameter transfer of the fixed angles is validated using a combination of full statevector, Projected Entangled Pair States (PEPS), Matrix Product State (MPS), and LOWESA numerical simulations. We find that the QAOA parameter transfer from single instances applied to other (unseen) problem instances does not in general provide monotonically improving performance as a function of $p$ - there are many cases where the performance temporarily decreases as a function of $p$ - but despite this the transferred angles have a general trend of improved expectation value as the QAOA depth increases, in many cases converging close to the true ground-state energy of the $100+$ qubit instances. We also sample the hardware-compatible Ising models using the ensemble of transfer-learned QAOA parameters on several superconducting qubit IBM Quantum processors with 127, 133, and 156 qubits. We find continuous solution quality improvement of the hardware-compatible QAOA circuits run on the IBM NISQ processors up to $p=5$ on \texttt{ibm\_fez}, up to $p=9$ on \texttt{ibm\_torino}, and up to $p=10$ on \texttt{ibm\_pittsburgh}.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Characterization and prediction of the electromechanical wear of contact tips during wire arc additive manufacturing of 316L stainless steel

Here, this study seeks to better understand the degradation of the contact tip with respect to WAAM for a 316L wire electrode as well as explore methods of monitoring the contact tip state from process data. The contact tip, a consumable component, positions the wire and serves as the electrical contact surface between the wire electrode and the welding power supply. The wear of the contact tip was characterized in terms of material loss and material contamination for a set of tips worn to discrete levels as measured by the amount of wire fed or arc time. Geometrical characterization found a 49% increase in the bore exit area at 180 meters of wire fed. Machine learning models were developed to predict the relative bore exit area of the contact tip from arc-based process data and a random forest classifier exhibited favorable performance with a cross-validated f1-score of 0.84. The regression architecture implemented a multi-layer perceptron with the ability to predict the relative exit area with an $R^2$ score of 0.75. Key features used in the prediction include the standard deviation of the voltage and the time between shorts.

Contact tip wear↗

Rational design of heterogeneous single-site catalysts via surface organometallic chemistry

Single-site heterogeneous catalysts offer an attractive route to unite the molecular precision of homogeneous catalysis with the durability and practical advantages of solids. Surface organometallic chemistry (SOMC) provides a particularly powerful strategy for this purpose by grafting molecular precursors onto tailored surfaces and converting support functionalities into ligand environments for isolated metal centers. As a result, SOMC brings the language and logic of coordination chemistry to heterogeneous catalysis, where the support becomes an integral part of the active site coordination sphere. This Review surveys recent progress in the rational design of SOMC-derived single-site catalysts, with emphasis on synthetic routes, post synthetic transformations, and the deliberate tuning of catalytic behavior through metal-support interactions. Discussions are made on how support identity, hydroxyl topology, acidity, and redox activity shape the geometry, electronic structure, and oxidation state of supported metal sites, as well as how these factors determine activity, selectivity, and stability. We also examine a central limitation of these systems: despite their molecularly informed design, supported single sites often exist as structurally distributed ensembles rather than uniform species, particularly on amorphous supports. This site heterogeneity, along with catalyst dynamics under operating conditions, remains a major barrier to definitive structure-activity relationships. Therefore, emerging approaches that combine advanced characterization, first-principles modeling, ensemble kinetics, and machine learning to resolve active-site structure and guide catalyst development are highlighted. Together, these advances position SOMC as a versatile coordination chemistry framework for the predictive design of heterogeneous catalysts with well-defined molecularly tailored active sites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Chemical-Free Lithium Separation from High-Salinity Brines Using Model-Informed and Machine Learning-Optimized Multi-Column Zwitterionic Chromatography

Direct Lithium Extraction (DLE) technologies often struggle to produce high-purity lithium salts from high-salinity brines, as current approaches require chemical-based elution, regeneration, and precipitation steps, resulting in significant environmental footprints. A novel salt fractionation approach using carboxybetaine resin, known as zwitterionic chromatography (ZIC), has demonstrated that lithium ions can be separated from divalent cations under high-salinity conditions using only water as eluent, with no regeneration required. To enable continuous and scalable deployment of this approach, we developed a chemical-free Multi-column Zwitterionic Chromatography (MZC) process and its theoretical and process models. To predict and optimize this nontraditional separation system, we introduced a novel anti-Langmuir isotherm, and the isotherm parameters were estimated through a machine learning-driven optimization based on artificial neural network ensembles with numerical feasibility assessment. Using machine learning-driven optimization, the MZC process achieved 98.0% lithium recovery, 99.5 % Li/(Li + Mg + Ca) purity, a 31.3% productivity increase, and a 33% reduction in water use compared to batch operation. The proposed MZC process enables lithium separation at $0.6-1.2 kg-1 Li, with costs dominated by resin manufacturing, while offering lower separation costs and carbon footprint compared with conventional carbonation. Overall, these findings position the MZC process as an effective polishing step within scalable and sustainable lithium production pipelines.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Convolutional Neural Network to Emulate Seasonal Tropical Cyclone Activity

Abstract It has been widely recognized that tropical cyclone (TC) genesis requires favorable large‐scale environmental conditions. Based on these linkages, numerous efforts have been made to establish an empirical relationship between seasonal TC activities and large‐scale environmental favorability in a quantitative way, which lead to conceptual functions such as the TC genesis index. However, due to the limited amount of reliable TC observations and complexity of the climate system, a simple analytic function may not be an accurate portrait of the empirical relationship between TCs and their ambiences. In this research, we use convolution neural networks (CNNs) to disentangle this complex relationship. To circumvent the limited amount of seasonal TC observation records, we implement transfer‐learning technique to train ensemble of CNNs first on suites of high‐resolution climate model simulations with realistic seasonal TC activities and large‐scale environmental conditions, and then on a state‐of‐the‐art reanalysis from 1950 to 2019. The trained CNNs can well reproduce the historical TC records and yields significant seasonal prediction skills when the large‐scale environmental inputs are provided by operational climate forecasts. Furthermore, by inputting the ensemble CNNs with 20th century reanalysis products and Phase 6 of the Coupled Model Intercomparison Project (CMIP6) simulations, we investigated TC variability and its changes in the past and future climates. Specifically, our ensemble CNNs project a decreasing trend of global mean TC activity in the future warming scenario, which is consistent with our future projections using high‐resolution climate model.

Meteorology & Atmospheric Sciences↗

Mesoscale Modeling to Characterize Eagle Soaring Habitat

Uncovering drivers of risk is crucial to understanding interactions between wildlife and wind turbines, and identifying options for impact minimization. These drivers tie to co-variates linked to behavior and movement patterns that allow us to estimate locations and periods of risk. For volant species, atmospheric flow can have significant influence on flight patterns. For obligate soaring birds, like golden eagles, updraft velocities can inform where eagles are likely to travel, at what altitude, and where conditions are not likely sufficient to sustain soaring flight. This has been an active area of study in recent years, using relatively coarse atmospheric data generally at the 20km x 20km scale or larger. Leveraging a 20-year dataset from the Weather Research and Forecasting Model (WRF) (https://www.mmm.ucar.edu/weather-research-and-forecasting-model), we are quantifying vertical velocities across the continental United States at a 2km x 2km resolution. Wind resource data sets originally were static maps showing the mean annual wind speed over an area. However, for these data sets to be optimally used for various applications they must be high-resolution time series, seamlessly span large geographic contexts, and account for uncertainty in wind speed. The National Renewable Energy Laboratory is producing public available datasets that meet these criteria and will be bias corrected to yield the most accurate wind resource data. This effort is an augment to the current WIND Toolkit which houses a high resolution data set. In the new iteration of the WIND Toolkit, a 20-year dataset will be used to improve the accuracy and estimate uncertainty using ensemble and machine-learning techniques. The resulting product will be the most accurate dataset of its size and at a 2km x 2km spatial and 5-minute temporal resolution. Through this work, a mesoscale vertical velocity layer will be produced by calculating the likelihood of orographic updraft and thermal updraft conditions across the continental United States. Specifically, we will use WRF model output combined with digital elevation maps to predict updrafts and then determine if vertical velocities are sufficient to support Golden Eagle soaring and gliding. Ultimately this data layer will be made available as a GIS layer in the Wind Prospector (maps.nrel.gov/wind-prospector/) tool or a similar framework. Data that will be incorporated include wind speed, direction temperature, relative humidity, barometric pressure, air density, precipitation rate, solar radiation, atmospheric stability, skin temperature, and upward heat flux. These products will advance research on interactions between volant species and wind energy by providing open access to highly resolved data with uncertainty quantification not previously available at this scale.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Semi-supervised Bayesian Low-shot Learning

Deep neural networks (NNs) typically outperform traditional machine learning (ML) approaches for complicated, non-linear tasks. It is expected that deep learning (DL) should offer superior performance for the important non-proliferation task of predicting explosive device configuration based upon observed optical signature, a task which human experts struggle with. However, supervised machine learning is difficult to apply in this mission space because most recorded signatures are not associated with the corresponding device description, or “truth labels.” This is challenging for NNs, which traditionally require many samples for strong performance. Semi-supervised learning (SSL), low-shot learning (LSL), and uncertainty quantification (UQ) for NNs are emerging approaches that could bridge the mission gaps of few labels and rare samples of importance. NN explainability techniques are important in gaining insight into the inferential feature importance of such a complex model. In this work, SSL, LSL, and UQ are merged into a single framework, a significant technical hurdle not previously demonstrated. Exponential Average Adversarial Training (EAAT) and Pairwise Neural Networks (PNNs) are chosen as the SSL and LSL methods of choice. Permutation feature importance (PFI) for functional data is used to provide explainability via the Variable importance Explainable Elastic Shape Analysis (VEESA) pipeline. A variety of uncertainty quantification approaches are explored: Bayesian Neural Networks (BNNs), ensemble methods, concrete dropout, and evidential deep learning. Two final approaches, one utilizing ensemble methods and one utilizing evidential learning, are constructed and compared using a well-quantified synthetic 2D dataset along with the DIRSIG Megascene.

97 MATHEMATICS AND COMPUTING↗

Uncertainty guided online ensemble for non-stationary data streams in fusion science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.

AI↗

Parameter inference from event ensembles and the top-quark mass

One of the key tasks of any particle collider is measurement. In practice, this is often done by fitting data to a simulation, which depends on many parameters. Sometimes, when the effects of varying different parameters are highly correlated, a large ensemble of data may be needed to resolve parameter-space degeneracies. An important example is measuring the top-quark mass, where other physical and unphysical parameters in the simulation must be profiled when fitting the top-quark mass parameter. We compare four different methodologies for top-quark mass measurement: a classical histogram fit similar to one commonly used in experiment augmented by soft-drop jet grooming; a 2D profile likelihood fit with a nuisance parameter; a machine-learning method called DCTR; and a linear regression approach, either using a least-squares fit or with a dense linearly-activated neural network. Despite the fact that individual events are totally uncorrelated, we find that the linear regression methods work most effectively when we input an ensemble of events sorted by mass, rather than training them on individual events. Although all methods provide robust extraction of the top-quark mass parameter, the linear network does marginally best and is remarkably simple. For the top study, we conclude that the Monte-Carlo-based uncertainty on current extractions of the top-quark mass from LHC data can be reduced significantly (by perhaps a factor of 2) using networks trained on sorted event ensembles. More generally, machine learning from ensembles for parameter estimation has broad potential for collider physics measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Contrasting Carbon–Water–Energy Dynamics in Perennial and Annual Bioenergy Agroecosystems Using Eddy Covariance and Interpretable Machine Learning

Understanding how agroecosystems respond to environmental variability is fundamental to predicting productivity and sustainability under a changing climate. We analyzed 55 site-years of high-frequency eddy covariance observations from five agroecosystems—two perennial grasses (miscanthus and switchgrass), two annual rotation systems (maize–soybean and sorghum–soybean), and a restored native prairie—to examine ecosystem-scale carbon, water, and energy fluxes. Using an interpretable machine-learning framework with regression tree ensembles, Shapley Additive Explanations, and Accumulated Local Effects, we quantified how environmental and temporal factors regulate gross primary productivity (GPP), evapotranspiration (ET), water-use efficiency, and the Bowen ratio. Perennials exhibited stronger physiological buffering and maintained fluxes across a broader range of temperature and moisture conditions, reflecting deeper rooting and persistent canopy cover. Annuals, in contrast, showed greater short-term variability and stronger coupling to atmospheric demand, with GPP and ET declining rapidly under low humidity or soil moisture. Differences in temperature sensitivity of Bowen ratio further revealed that perennials sustained proportionally greater sensible heat flux under cool conditions, whereas annuals exhibited constrained energy exchange when evaporative demand was low. Together, these results demonstrate that crop life cycle and canopy structure are fundamental determinants of ecosystem-scale carbon–water–energy coupling. By integrating long-term flux observations with interpretable machine learning, this study identifies the environmental drivers that shape agroecosystem function and highlights how conversion from annual to perennial feedstocks can enhance climatic resilience and alter land–atmosphere energy feedbacks. These findings provide a data-driven basis for improving crop and Earth-system models and for guiding bioenergy landscape design under future climate scenarios.

Accumulated Local Effects↗

Reduced-order autodifferentiable ensemble Kalman filters

This paper introduces a computational framework to reconstruct and forecast a partially observed state that evolves according to an unknown or expensive-to-simulate dynamical system. Our reduced-order autodifferentiable ensemble Kalman filters (ROAD-EnKFs) learn a latent low-dimensional surrogate model for the dynamics and a decoder that maps from the latent space to the state space. The learned dynamics and decoder are then used within an EnKF to reconstruct and forecast the state. Numerical experiments show that if the state dynamics exhibit a hidden low-dimensional structure, ROAD-EnKFs achieve higher accuracy at lower computational cost compared to existing methods. If such structure is not expressed in the latent state dynamics, ROAD-EnKFs achieve similar accuracy at lower cost, making them a promising approach for surrogate state reconstruction and forecasting.

Mathematics↗

Learning and discovering multiple solutions using physics-informed neural networks with random initialization and deep ensemble

In this work we explore the capability of physics-informed neural networks (PINNs) to discover multiple solutions. Many real-world phenomena governed by nonlinear differential equations (DEs), such as fluid flow, exhibit multiple solutions under the same conditions, yet capturing this solution multiplicity remains a significant challenge. A key difficulty lies in providing appropriate initial conditions or guesses, as widely used time-marching schemes and Newton’s method are highly sensitive to these choices when solving complex computational problems. While machine learning models, particularly PINNs, have shown promise in solving DEs, their ability to capture multiple solutions remains underexplored. In this work, we propose a simple and practical approach using PINNs to learn and discover multiple solutions. We first demonstrate that PINNs, when combined with random initialization and deep ensemble method—originally developed for uncertainty quantification—can effectively uncover multiple solutions to nonlinear ordinary and partial DEs. Although training large ensembles of PINNs may appear computationally demanding, this can be done efficiently using vectorization techniques supported by modern deep learning frameworks, allowing many networks to be trained simultaneously. Our approach highlights the critical role of initialization in shaping solution diversity, addressing an often-overlooked aspect of machine learning for scientific computing. Furthermore, we propose utilizing PINN-generated solutions as initial conditions or initial guesses for conventional numerical solvers to enhance accuracy and efficiency in capturing multiple solutions. Extensive numerical experiments, including the Allen–Cahn equation and cavity flow, where our approach successfully identifies both stable and unstable solutions, validate the effectiveness of our method. These findings establish a general and efficient framework for addressing solution multiplicity in nonlinear DEs.

97 MATHEMATICS AND COMPUTING↗

Insight into molecular basis and dynamics of full-length CRaf kinase in cellular signaling mechanisms

Raf kinases play key roles in signal transduction in cells for regulating proliferation, differentiation, and survival. Despite decades of research into functions and dynamics of Raf kinases with respect to other cytosolic proteins, understanding Raf kinases is limited by the lack of their full-length structures at the atomic resolution. Here, we present the first model of the full-length CRaf kinase obtained from artificial intelligence/machine learning algorithms with a converging ensemble of structures simulated by large-scale temperature replica exchange simulations. Our model is validated by comparing simulated structures with the latest cryo-EM structure detailing close contacts among three key domains and regions of the CRaf. Our simulations identify potentially new epitopes of intramolecule interactions within the CRaf and reveal a dynamical nature of CRaf kinases, in which the three domains can move back and forth relative to each other for regulatory dynamics. The dynamic conformations are then used in a docking algorithm to shed insight into the paradoxical effect caused by vemurafenib in comparison with a paradox breaker PLX7904. In this study, we propose a model of Raf-heterodimer/KRas-dimer as a signalosome based on the dynamics of the full-length CRaf.

59 BASIC BIOLOGICAL SCIENCES↗

A Bayesian Multi-fidelity Neural Network to Predict Nonlinear Frequency Backbone Curves

The use of structural mechanics models during the design process often leads to the development of models of varying fidelity. Often low-fidelity models are efficient to simulate but lack accuracy, while the high-fidelity counterparts are accurate with less efficiency. Here, this paper presents a multi-fidelity surrogate modeling approach that combines the accuracy of a high-fidelity finite element model with the efficiency of a low-fidelity model to train an even faster surrogate model that parameterizes the design space of interest. The objective of these models is to predict the nonlinear frequency backbone curves of the Tribomechadynamics Research Challenge benchmark structure which exhibits simultaneous nonlinearities from frictional contact and geometric nonlinearity. The surrogate model consists of an ensemble of neural networks that learn the mapping between low and high-fidelity data through nonlinear transformations. Bayesian neural networks are used to assess the surrogate model's uncertainty. Once trained, the multi-fidelity neural network is used to perform sensitivity analysis to assess the influence of the design parameters on the predicted backbone curves. Additionally, Bayesian calibration is performed to update the input parameter distributions to correlate the model parameters to the collection of experimentally measured backbone curves.

42 ENGINEERING↗

Quantifying the Known Unknown: Including Marine Sources of Greenhouse Gases in Climate Modeling

Researchers have recently estimated that Arctic submarine permafrost currently traps 60 billion tons of methane and contains 560 billion tons of organic carbon in seafloor sediments and soil, a giant pool of carbon with potentially large feedbacks on the climate system. Unlike terrestrial permafrost, the submarine permafrost system has remained a “known unknown” because of the difficulty in acquiring samples and measurements. Consequently, this potentially large carbon stock never yet considered in global climate models or policy discussions, represents a real wildcard in our understanding of Earth’s climate. This report summarizes our group’s effort at developing a numerical modeling framework designed to produce a first-of-its-kind estimate of Arctic methane gas releases from the marine sediments to the water column, and potentially to the atmosphere, where positive climate feedback may occur. Newly developed modeling capability supported by the Laboratory Directed Research and Development (LDRD) program at Sandia National Laboratories now gives us the ability to probabilistically map gas distribution and quantity in the seabed by using a hybrid approach of geospatial machine learning, and predictive numerical thermodynamic ensemble modeling. The novelty in this approach is its ability to produce maps of useful data in regions that are only sparsely sampled, a common challenge in the Arctic, and a major obstacle to progress in the past. By applying this model to the circum-Arctic continental shelves and integrating the flux of free gas from in situ methanogenesis and dissociating gas hydrates from the sediment column under climate forcing, we can provide the most reliable estimate of a spatially and temporally varying source term for greenhouse gas flux that can be used by global oceanographic circulation and Earth system models (such as DOE’s E3SM). The result will allow us to finally tackle the wildcard of the submarine permafrost carbon system, and better inform us about the severity of future national security threats that sustained climate change poses.

54 ENVIRONMENTAL SCIENCES↗

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur↗