Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Kronecker-structured covariance models for multiway data

Many applications produce multiway data of exceedingly high dimension. Modeling such multi-way data is important in multichannel signal and video processing where sensors produce multi-indexed data, e.g. over spatial, frequency, and temporal dimensions. We will address the challenges of covariance representation of multiway data and review some of the progress in statistical modeling of multiway covariance over the past two decades, focusing on tensor-valued covariance models and their inference. We will illustrate through a space weather application: predicting the evolution of solar active regions over time.

97 MATHEMATICS AND COMPUTING↗

Simple Building Calculator

Whole building energy modeling is a powerful tool for analyzing energy use in buildings. For large buildings benefits of modeling services can easily be quantified due to the short payback period of the implemented measures. For small buildings, the upfront costs of developing robust energy models often deter building owners from investing in energy modeling services. To address this shortcoming, a ”Simple Building Calculator” (SBC) was developed. SBC uses pre-simulated results of a range of common measures over a wide range of efficiency inputs. It combines whole building simulation results with statistical modeling techniques to predict energy impact of measures very quickly. In this paper we present the modeling methodology used to develop the data supporting SBC.

Building Energy Simulation, commercial building, E↗

Nepheline Crystallization Studies and the Structural Integrity of the Residual (SIR) - FY 2022 Study Glasses

In FY22 a set of 367 HLW glasses with compositions representative of the Waste Treatment and Immobilization Plant (WTP) processing region were aggregated and analyzed with the Structural Integrity of the Residual (SIR) model developed by the Savannah River National Laboratory (SRNL). An additional 12 glasses targeting 27 – 37 wt % alumina (in glass) and at waste loadings up to ~70 wt % were fabricated and added to the database. Collectively, this data set was analyzed with regression statistics and partitioning functions to reveal correlations between boron release and the chemical composition. FY22 results indicate that the SIR theoretical alkali-silicate precipitation term, Al 2 O 3 and B 2 O 3 concentration, and SIR computed non-bridging oxygen are significant predictors of chemical durability. A statistical model using the SIR parameters was shown to have greater overall accuracy than the nepheline discriminator model while maintaining false predictions of pour durability below 1%. In FY23, we will continue this work to focus on refining the SIR with glass compositions that challenge the SIR correlation terms, namely the theoretical crystalline precipitation and non-bridging oxygen terms.

36 MATERIALS SCIENCE↗

Sheaves as a Framework for Understanding and Interpreting Model Fit

As data grows in size and complexity, finding frameworks which aid in interpretation and analysis has become critical. This is particularly true when data comes from complex systems where extensive structure is available, but must be drawn from peripheral sources. In this paper we argue that in such situations, sheaves can provide a natural framework to analyze how well a statistical model fits at the local level (that is, on subsets of related datapoints) vs the global level (on all the data). The sheaf-based approach that we propose is suitably general enough to be useful in a range of applications, from analyzing sensor networks to understanding the feature space of a deep learning model.

Kvinge, Henry J.↗

In Situ Inference for Earth System Predictability

An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

In Situ Inference for Earth System Predictability

Focal Area: Focal Area 3: Insight gleaned from complex simulated data using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

Accurate prediction of oxygen vacancy concentration with disordered A-site cations in high-entropy perovskite oxides

Abstract Entropic stabilized ABO 3 perovskite oxides promise many applications, including the two-step solar thermochemical hydrogen (STCH) production. Using binary and quaternary A-site mixed {A}FeO 3 as a model system, we reveal that as more cation types, especially above four, are mixed on the A-site, the cell lattice becomes more cubic-like but the local Fe–O octahedrons are more distorted. By comparing four different Density Functional Theory-informed statistical models with experiments, we show that the oxygen vacancy formation energies ( $${E}_{V}^{f}$$ E V f ) distribution and the vacancy interactions must be considered to predict the oxygen non-stoichiometry ( δ ) accurately. For STCH applications, the $${E}_{V}^{f}$$ E V f distribution, including both the average and the spread, can be optimized jointly to improve Δ δ (difference of δ between the two-step conditions) in some hydrogen production levels. This model can be used to predict the range of water splitting that can be thermodynamically improved by mixing cations in {A}FeO 3 perovskites.

08 HYDROGEN↗

Evaluation of obstacle modelling approaches for resource assessment and small wind turbine siting: case study in the northern Netherlands

Abstract. Growth in adoption of distributed wind turbines for energy generation is significantly impacted by challenges associated with siting and accurate estimation of the wind resource. Small turbines, at hub heights of 40 m or less, are greatly impacted by terrestrial obstacles such as built structures and vegetation that can cause complex wake effects. While some progress in high-fidelity complex fluid dynamics (CFD) models has increased the potential accuracy for modelling the impacts of obstacles on turbulent wind flow, these models are too computationally expensive for practical siting and resource assessment applications. To understand the efficacy of available models in situ, this study evaluates classic and commonly used methods alongside new state-of-the-art lower-order models derived from CFD simulations and machine learning approaches. This evaluation is conducted using a subset of an extensive original dataset of measurements from more than 300 operational wind turbines in the northern Netherlands. The results show that data-driven methods (e.g. machine learning and statistical modelling) are most effective at predicting production at real sites with an average error in annual energy production of 2.5 %. When sufficient data may not be available de novo to support these data-driven approaches, models derived from high-fidelity simulations show promise and reliably outperform classic methods. On average these models have 6.3 %–11.5 % error compared with 26 % for classic methods and 27 % baseline error for reanalysis data without obstacle correction. While more performant on average, these methods are also sensitive to the quality of obstacle descriptions and reanalysis inputs.

17 WIND ENERGY↗

Classification of Photovoltaic Failures with Hidden Markov Modeling, an Unsupervised Statistical Approach

Failure detection methods are of significant interest for photovoltaic (PV) site operators to help reduce gaps between expected and observed energy generation. Current approaches for field-based fault detection, however, rely on multiple data inputs and can suffer from interpretability issues. In contrast, this work offers an unsupervised statistical approach that leverages hidden Markov models (HMM) to identify failures occurring at PV sites. Using performance index data from 104 sites across the United States, individual PV-HMM models are trained and evaluated for failure detection and transition probabilities. This analysis indicates that the trained PV-HMM models have the highest probability of remaining in their current state (87.1% to 93.5%), whereas the transition probability from normal to failure (6.5%) is lower than the transition from failure to normal (12.9%) states. A comparison of these patterns using both threshold levels and operations and maintenance (O&M) tickets indicate high precision rates of PV-HMMs (median = 82.4%) across all of the sites. Although additional work is needed to assess sensitivities, the PV-HMM methodology demonstrates significant potential for real-time failure detection as well as extensions into predictive maintenance capabilities for PV.

classification↗

A Comparison of PV Resource Modeling for Sizing Microgrid Components

Microgrid systems are being deployed with increased frequency to meet critical building loads during unplanned power outages. Determining the appropriate type and size of components that make up a microgrid can greatly affect its ability to meet this need. This paper presents a new tool to appropriately size PV, battery, and generator capacities to meet all site loads, given a resilience goal. It uses a statistical model to simulate microgrid behavior under a large range of solar resource conditions, providing resilience planners with confidence that critical loads will be met even under extreme weather conditions. We compare this model with one that uses typical weather data for sizing microgrids, and quantify the resilience effects of using these different methods. Our method produces more conservative generator capacities and fuel requirements in an outage, predicting up to 30% more fuel required in the most extreme case. Microgrid systems that are designed using more conservative estimates are more likely to continue to serve load during emergency outages under a large range of conditions.

Newman, Sarah F.↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

Optimal experimental design: Formulations and computations

Questions of ‘how best to acquire data’ are essential to modelling and prediction in the natural and social sciences, engineering applications, and beyond. Optimal experimental design (OED) formalizes these questions and creates computational methods to answer them. This article presents a systematic survey of modern OED, from its foundations in classical design theory to current research involving OED for complex models. We begin by reviewing criteria used to formulate an OED problem and thus to encode the goal of performing an experiment. We emphasize the flexibility of the Bayesian and decision-theoretic approach, which encompasses information-based criteria that are well-suited to nonlinear and non-Gaussian statistical models. We then discuss methods for estimating or bounding the values of these design criteria; this endeavour can be quite challenging due to strong nonlinearities, high parameter dimension, large per-sample costs, or settings where the model is implicit. A complementary set of computational issues involves optimization methods used to find a design; we discuss such methods in the discrete (combinatorial) setting of observation selection and in settings where an exact design can be continuously parametrized. Finally we present emerging methods for sequential OED that build non-myopic design policies, rather than explicit designs; these methods naturally adapt to the outcomes of past experiments in proposing new experiments, while seeking coordination among all experiments to be performed. Throughout, we highlight important open questions and challenges.

97 MATHEMATICS AND COMPUTING↗

Quantitative assessment of fitting errors associated with streak camera noise in Thomson scattering data analysis

Thomson scattering measurements in high energy density experiments are often recorded using optical streak cameras. In the low-signal regime, noise introduced by the streak camera can become an important and sometimes the dominant source of measurement uncertainty. In this paper, we present a formal method of accounting for the presence of streak camera noise in our measurements. We present a phenomenological description of the noise generation mechanisms and present a statistical model that may be used to construct the covariance matrix associated with a given measurement. This model is benchmarked against simulations of streak camera images. We demonstrate how this covariance may then be used to weight fitting of the data and provide quantitative assessments of the uncertainty in the fitting parameters determined by the best fit to the data and build confidence in the ability to make statistically significant measurements in the low-signal regime, where spatial correlations in the noise become apparent. These methods will have general applicability to other measurements made using optical streak cameras.

47 OTHER INSTRUMENTATION↗

Evaluation of Turbulence and Dispersion in Multiscale Atmospheric Simulations over Complex Urban Terrain during the Joint Urban 2003 Field Campaign

Abstract This paper evaluates the representation of turbulence and its effect on transport and dispersion within multiscale and microscale-only simulations in an urban environment. These simulations, run using the Weather Research and Forecasting Model with the addition of an immersed boundary method, predict transport and mixing during a controlled tracer release from the Joint Urban 2003 field campaign in Oklahoma City, Oklahoma. This work extends the results of a recent study through analysis of turbulence kinetic energy and turbulence spectra and their role in accurately simulating wind speed, direction, and tracer concentration. The significance and role of surface heat fluxes and use of the cell perturbation method in the numerical simulation setup are also examined. Our previous study detailed the model development necessary for our multiscale simulations, examined model skill at predicting wind speeds and tracer concentrations, and demonstrated that dynamic downscaling from mesoscale to microscale through a sequence of nested simulations can improve predictions of transport and dispersion relative to a microscale-only simulation forced by idealized meteorology. Here, predictions are compared with observations to assess qualitative agreement and statistical model skill at predicting wind speed, wind direction, tracer concentration, and turbulent kinetic energy at locations throughout the city. We also investigate the scale distribution of turbulence and the associated impact on model skill, particularly for predictions of transport and dispersion. Our results show that downscaled large-scale turbulence, which is unique to the multiscale simulations, significantly improves predictions of tracer concentrations in this complex urban environment. Significance Statement Simulations of atmospheric transport and mixing in urban environments have many applications, including pollution modeling for urban planning or informing emergency response following a hazardous release. These applications include phenomena with spatial scales spanning from millimeters to kilometers. Most simulations resolve flow only within the urban area of interest, omitting larger scales of turbulence and regional influences. This study examines a method that resolves both the small and large-scale flow features. We evaluate simulation accuracy by comparing predictions with observations from an experiment involving the release of a tracer gas in Oklahoma City, Oklahoma, with emphasis on correctly modeling turbulent fluctuations. Our results demonstrate the importance of resolving large-scale flow features when predicting transport and dispersion in urban environments.

42 ENGINEERING↗

Experimental Validation of a Module Cell Cracking Model

The What's Cracking app can predict how changes in crystalline silicon photovoltaic (PV) module materials, design, and mounting affect its susceptibility for cell fracture under uniform loading. This work has experimentally validated the app. A set of commercial crystalline silicon PV modules was obtained for this study. The modules were uniformly loaded at three different mounting points, and their subsequent cell fractures were recorded. A large sample size allowed for the development of an experimental statistical model for cell fracture. Here, the comparison of the experiment to predictions from the app is in excellent agreement. Both experimental and modeling results also elucidate how moving the module mounting points toward the center of the module increases the probability of cell fracture.

14 SOLAR ENERGY↗

Predicting battery capacity from impedance at varying temperature and state of charge using machine learning

Prediction of battery health from electrochemical impedance spectroscopy (EIS) data can enable rapid measurement of battery state in real-world applications without using additional sensors or time-consuming performance measurements. However, deconvoluting the effect of capacity, state of charge, and temperature on EIS response is complicated analytically. Here, various machine-learning models, such as linear, Gaussian process, random forest, and artificial neural network regression, are utilized to predict capacity from EIS using hundreds of capacity, direct current (DC) resistance, and EIS measurements recorded under varying conditions of health, temperature, and state of charge (SOC). Several feature extraction and selection methods from traditional electrochemical analysis and statistical modeling are explored using machine-learning pipelines. EIS data from just two frequencies can accurately predict capacity, and interrogation shows that the optimal set of frequencies is not usually intuitive. Best results are achieved with an ensemble model, which predicts battery capacity with a mean absolute error of 1.9% on data from unobserved cells.

25 ENERGY STORAGE↗

Considering uncertainties expands the lower tail of maize yield projections

Crop yields are sensitive to extreme weather events. Improving the understanding of the mechanisms and the drivers of the projection uncertainties can help to improve decisions. Previous studies have provided important insights, but often sample only a small subset of potentially important uncertainties. Here we expand on a previous statistical modeling approach by refining the analyses of two uncertainty sources. Specifically, we assess the effects of uncertainties surrounding crop-yield model parameters and climate forcings on projected crop yield. We focus on maize yield projections in the eastern U.S.in this century. We quantify how considering more uncertainties expands the lower tail of yield projections. We characterized the relative importance of each uncertainty source and show that the uncertainty surrounding yield model parameters is the main driver of yield projection uncertainty.

59 BASIC BIOLOGICAL SCIENCES↗

New 26 P( p, γ ) 27 S Thermonuclear Reaction Rate and Its Astrophysical Implications in the rp -process

Accurate nuclear reaction rates for 26 P(p, γ) 27 S are pivotal for a comprehensive understanding of the rp-process nucleosynthesis path in the region of proton-rich sulfur and phosphorus isotopes. However, large uncertainties still exist in the current rate of 26 P(p, γ) 27 S because of the lack of nuclear mass and energy level structure information for 27 S. We reevaluate this reaction rate using the experimentally constrained 27 S mass, together with the shell model predicted level structure. It is found that the 26 P(p, γ) 27 S reaction rate is dominated by a direct capture reaction mechanism despite the presence of three resonances at E = 1.104, 1.597, and 1.777 MeV above the proton threshold in 27 S. The new rate is overall smaller than the other previous rates from the Hauser–Feshbach statistical model by at least 1 order of magnitude in the temperature range of X-ray burst interest. In addition, we consistently update the photodisintegration rate using the new 27 S mass. The influence of new rates of forward and reverse reaction in the abundances of isotopes produced in the rp-process is explored by postprocessing nucleosynthesis calculations. The final abundance ratio of 27 S/ 26 P obtained using the new rates is only 10% of that from the old rate. The abundance flow calculations show that the reaction path 26 P(p, γ) 27 S(β + ,ν) 27 P is not as important as previously thought for producing 27 P. The adoption of the new reaction rates for 26 P(p, γ) 27 S only reduces the final production of aluminum by 7.1% and has no discernible impact on the yield of other elements.

79 ASTRONOMY AND ASTROPHYSICS↗