Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Average Accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Evaluating Image Classification Deep Convolutional Neural Network Architectures for Remaining Useful Life Estimation of Turbofan Engines

Accurate estimation of the remaining useful life (RUL) is a key component of condition-based maintenance (CBM) and prognosis and health management (PHM). Data-based models for the estimation of RUL are of particular interest because expert knowledge of systems is not always available, and physical modeling is often not feasible. Additionally, using data-based models, which make decisions based on raw sensor data, allow features to be learned instead of manually determined. In this work, deep convolutional neural network (CNN) architectures are investigated for their ability to estimate the RUL of turbofan engines. To improve the accuracy of the models, CNN architectures, which have proven successful in image classification, are implemented and tested. Specifically, the blocks used in the Visual Geometry Group (VGG) architecture, inception modules used in the GoogLeNet architecture, and residual blocks used in the ResNet architecture are incorporated. To account for varying flight lengths, the input to the models is a window of time series data collected from the engine under test. Window locations at the climb, cruise, and descent stages are considered. To further improve the RUL estimations, multiple overlapping windows at each location are used. This increases the amount of training data available and is found to increase the accuracy of the resulting RUL estimations by averaging the estimates from all overlapping segments. The model is trained and tested using the new Commercial Modular Aero-Propulsion System Simulation (N-CMAPSS) data set, and high prognosis accuracy was achieved. Furthermore, this work expands on the model developed and used in the 2021 PHM Society Data Challenge, which received second place.

convolutional neural networks↗

Developing Science-based fueling protocols for 250-bar hydrogen tanks onboard hydrogen ferries: Experiments and modeling

Combined modeling and experimental studies are reported of the fueling of a large (28 kg capacity) 250-bar Type IV hydrogen tank of the type being deployed on early hydrogen ferries, such as the MV Sea Change. The primary goal was to determine how such tanks can be successfully fueled with hydrogen (state of charge greater than 97%) within 45 minutes without exceeding the 82 °C temperature limit for such tanks. The modeling studies show that a gas injector is needed to avoid thermal stratification during hydrogen fueling which can result in potential hot spots. Empirically, precooling of the hydrogen to 0 °C was found to be needed in some of the cases examined, as ambient conditions greatly affected the need for a precooling to achieve the 45-minute fill time desired by end users. The experimental results afforded a calibration of the engineering model SOFIL for these large 250-bar tanks, which now enables using SOFIL to predict volume-averaged hydrogen fueling temperatures to an accuracy of ±2.7°C for these tanks. The model can therefore be used to evaluate potential scenarios for development of a standardized fueling methodology for ferries utilizing large Type-IV tanks.

08 HYDROGEN↗

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗

PM 2.5 Is Insufficient to Explain Personal PAH Exposure

To understand how chemical exposure can impact health, researchers need tools that capture the complexities of personal chemical exposure. In practice, fine particulate matter (PM 2.5 ) air quality index (AQI) data from outdoor stationary monitors and Hazard Mapping System (HMS) smoke density data from satellites are often used as proxies for personal chemical exposure, but do not capture total chemical exposure. Silicone wristbands can quantify more individualized exposure data than stationary air monitors or smoke satellites. However, it is not understood how these proxy measurements compare to chemical data measured from wristbands. In this study, participants wore daily wristbands, carried a phone that recorded locations, and answered daily questionnaires for a 7-day period in multiple seasons. We gathered publicly available daily PM 2.5 AQI data and HMS data. We analyzed wristbands for 94 organic chemicals, including 53 polycyclic aromatic hydrocarbons. Wristband chemical detections and concentrations, behavioral variables (e.g., time spent indoors), and environmental conditions (e.g., PM 2.5 AQI) significantly differed between seasons. Machine learning models were fit to predict personal chemical exposure using PM 2.5 AQI only, HMS only, and a multivariate feature set including PM 2.5 AQI, HMS, and other environmental and behavioral information. On average, the multivariate models increased predictive accuracy by approximately 70% compared to either the AQI model or the HMS model for all chemicals modeled. This study provides evidence that PM 2.5 AQI data alone or HMS data alone is insufficient to explain personal chemical exposures. Our results identify additional key predictors of personal chemical exposure.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Stable Simulation of the Community Atmosphere Model Using Machine‐Learning Physical Parameterization Trained With Experience Replay

In recent years, machine learning (ML) models have been used to improve physical parameterizations of general circulation models (GCMs). A significant challenge of integrating ML models into GCMs is the online instability when they are coupled for long‐term simulation. We present a new strategy that demonstrates robust online stability when the physical parameterization package of an atmospheric GCM is replaced by a deep ML model. The method uses experience replay with a multistep training scheme of the ML model in which the model's own output at the previous time step is used in the training. Predicted physics tendencies in the replay buffer with the most recent errors in the training iterations are reused, making the ML model learn from its own errors. The training method reduces the gap between the offline and online environments of the ML model. The method is used to train the ML model as the physical parameterization of the Community Atmosphere Model (CAM5) with training data from the Multi‐scale Modeling Framework high resolution simulations. Three 6‐year online simulations of the CAM5 are carried out by using the ML physics package. The simulated spatial distributions of precipitation, surface temperature and zonally averaged atmospheric fields demonstrate overall better accuracy than that of the standard CAM5 and benchmark model even without the use of additional physical constraints or tuning. This work is the first to demonstrate a solution to address the online instability problem in climate modeling with ML physics by using experience replay.

54 ENVIRONMENTAL SCIENCES↗

Marine Hydrogen Demonstration

This report summarizes Phase 1 of a project involving the design of a Floating Hydrogen Production and Dispensing Barge destined for the Port of San Francisco (SF). The H 2 Barge is designed to produce renewable H 2 at the rate of ~ 530 kg/day, storing 512 kg of hydrogen at 517-bar, allowing fast refueling of hydrogen fuel cell vessels and land-side hydrogen delivery trailers for distribution into the nascent SF hydrogen ecosystem. The broader considerations that impacted the H2 Barge design are also described. An account is given of a new review process formulated by the United States Coast Guard (USCG) to review this first-of-its-kind maritime implementation of hydrogen technology. The immediate goals of the H 2 Barge Project are to 1) demonstrate the feasibility, viability and methods of hydrogen production, storage and fueling in a maritime context, 2) help shape (where needed) and navigate the required local, state and federal regulatory gauntlet and 3) catalyze a “green hydrogen ecosystem” (both marine and landside) with locally produced renewable hydrogen at the San Francisco waterfront. A summary is also given of the modeling and experimental activity of Phase 1 directed to the development of science-based refueling protocols for large marine Type IV 250-bar hydrogen tanks. Combined modeling and experimental studies are reported of the filling of large (28 kg) 250-bar Type IV hydrogen tanks of the type being deployed on early hydrogen ferries, such as the MV Sea Change. The primary question was to determine how such tanks can be successfully filled (state of charge greater than 97%) within 45 minutes without exceeding the 82 °C temperature limit historically set for such tanks. The studies show that a gas injector is needed avoid thermal stratification during filling which can result in potential hot spots. Pre-cooling of the hydrogen was found to be essential in most cases, as ambient conditions greatly affect the need for a pre-cooling to achieve the 45-minute fill time. Pre-cooling cannot be supplied by nearby water, such as that found in nature (bays, lakes, rivers, etc.) because pre-cooling cooling below 0 °C was found to be necessary to avoid excessive compression heating. The experimental results afforded a calibration of the engineering model (SOFIL) for these large 250-bar tanks, which now enables using SOFIL to predict volume-averaged hydrogen filling temperatures to an accuracy of +/- 2.7°C for these tanks. The model can therefore be used to evaluate potential scenarios for development of a standardized fueling methodology for ferries utilizing large Type-IV tanks of the type examined here.

08 HYDROGEN↗

Data to Accompany: PM2.5 is insufficient to explain personal PAH exposure

Fine particulate matter (PM2.5) air quality index (AQI) data from outdoor stationary monitors and Hazard Mapping System (HMS) smoke density data from satellites are often used as proxies for personal chemical exposure. Silicone wristbands can quantify more individualized exposure data than stationary air monitors or smoke satellites. However, it is not understood how these proxy measurements compare to chemical data measured from wristbands. We hypothesized that predictive models for personal chemical exposure would be significantly improved by expanding beyond stationary PM2.5 AQI data or satellite HMS data to also include environmental and behavioral information. In Eugene, Oregon, participants wore daily wristbands, carried a phone that recorded locations, and answered daily questionnaires for a seven-day period in multiple seasons. We gathered publicly available daily PM2.5 AQI data and HMS data. We analyzed wristbands for 94 organic chemicals, including 53 polycyclic aromatic hydrocarbons (PAHs). Wristband chemical detections and concentrations, behavioral variables (e.g., time spent indoors), and environmental conditions (e.g., PM2.5 AQI) significantly differed between seasons. Machine learning models were fit to predict personal chemical exposure using PM2.5 AQI only, HMS only, and a multivariate feature set including PM2.5 AQI, HMS, and other environmental and behavioral information. On average, the multivariate models increased predictive accuracy by approximately 70% compared to either the AQI model or the HMS model for all chemicals modeled. This study provides evidence that PM2.5 AQI data alone or HMS data alone is insufficient to explain personal chemical exposures. Our results identify additional key predictors of personal chemical exposure.

Bramer, Lisa M↗

Technical note: Using long short-term memory models to fill data gaps in hydrological monitoring networks

Abstract. Quantifying the spatiotemporal dynamics in subsurface hydrological flows over a long time window usually employs a network of monitoring wells. However, such observations are often spatially sparse with potential temporal gaps due to poor quality or instrument failure. In this study, we explore the ability of recurrent neural networks to fill gaps in a spatially distributed time-series dataset. We use a well network that monitors the dynamic and heterogeneous hydrologic exchanges between the Columbia River and its adjacent groundwater aquifer at the U.S. Department of Energy's Hanford site. This 10-year-long dataset contains hourly temperature, specific conductance, and groundwater table elevation measurements from 42 wells with gaps of various lengths. We employ a long short-term memory (LSTM) model to capture the temporal variations in the observed system behaviors needed for gap filling. The performance of the LSTM-based gap-filling method was evaluated against a traditional autoregressive integrated moving average (ARIMA) method in terms of error statistics and accuracy in capturing the temporal patterns of river corridor wells with various dynamics signatures. Our study demonstrates that the ARIMA models yield better average error statistics, although they tend to have larger errors during time windows with abrupt changes or high-frequency (daily and subdaily) variations. The LSTM-based models excel in capturing both high-frequency and low-frequency (monthly and seasonal) dynamics. However, the inclusion of high-frequency fluctuations may also lead to overly dynamic predictions in time windows that lack such fluctuations. The LSTM can take advantage of the spatial information from neighboring wells to improve the gap-filling accuracy, especially for long gaps in system states that vary at subdaily scales. While LSTM models require substantial training data and have limited extrapolation power beyond the conditions represented in the training data, they afford great flexibility to account for the spatial correlations, temporal correlations, and nonlinearity in data without a priori assumptions. Thus, LSTMs provide effective alternatives to fill in data gaps in spatially distributed time-series observations characterized by multiple dominant frequencies of variability, which are essential for advancing our understanding of dynamic complex systems.

54 ENVIRONMENTAL SCIENCES↗

Maintaining Trust in Reduction: Preserving the Accuracy of Quantities of Interest for Lossy Compression

As the growth of data sizes continues to outpace computational resources, there is a pressing need for data reduction techniques that can significantly reduce the amount of data and quantify the error incurred in compression. Compressing scientific data presents many challenges for reduction techniques since it is often on non-uniform or unstructured meshes, is from a high-dimensional space, and has many Quantities of Interests (QoIs) that need to be preserved. To illustrate these challenges, we focus on data from a large scale fusion code, XGC. XGC uses a Particle-In-Cell (PIC) technique which generates hundreds of PetaBytes (PBs) of data a day, from thousands of timesteps. XGC uses an unstructured mesh, and needs to compute many QoIs from the raw data, f.One critical aspect of the reduction is that we need to ensure that QoIs derived from the data (density, temperature, flux surface averaged momentums, etc.) maintain a relative high accuracy. We show that by compressing XGC data on the high-dimensional, nonuniform grid on which the data is defined, and adaptively quantizing the decomposed coefficients based on the characteristics of the QoIs, the compression ratios at various error tolerances obtained using a multilevel compressor (MGARD) increases more than ten times. We then present how to mathematically guarantee that the accuracy of the QoIs computed from the reduced f is preserved during the compression. We show that the error in the XGC density can be kept under a user-specified tolerance over 1000 timesteps of simulation using the mathematical QoI error control theory of MGARD, whereas traditional error control on the data to be reduced does not guarantee the accuracy of the QoIs.

Gong, Qian↗

Regional-scale soil carbon predictions can be enhanced by transferring global-scale soil–environment relationships

Accurate modelling and mapping soil organic carbon are crucial for supporting soil health restoration and climate change mitigation at both regional and global scales. However, regional soil predictions often suffer from data scarcity and high prediction uncertainty. Utilizing a pre-trained global-to-regional soil carbon predictive model can be a potential solution to address this challenge. Despite its promise, how to construct and apply the global-scale model to enhance regional-scale soil carbon mapping remains largely unexplored. Here, we propose the Global Soil Carbon Pre-trained Model (GSoilCPM), a deep-learning-based domain adaptative model, to enhance regional-scale soil carbon predictions. Based on large amount of environmental covariate data and 106,167 soil samples across the globe, we verify our hypothesis of the effectiveness of this 'global-to-regional' modelling strategy. The pre-trained model can be then transferred and fine-tuned to bridge the regional- and global-scale soil–environment relationships. We applied and validated this modelling strategy in four regional-scale study areas, three in the Northern Hemisphere and one in the Southern Hemisphere, each with distinct environmental background. Compared to traditional modelling approaches as a baseline, four case studies all demonstrated significant improvement in prediction accuracy across diverse environments and varying data availabilities. The average percentage improvement across all regions is 10.93% (absolute values decreased by 1.20 g kg−1 averagely) in MAE and 29.04% (absolute values increased by 0.10 averagely) in CCC. The applicability and future horizons of using GSoilCPM were further discussed. We further reveal that regions with fewer soil samples or lower baseline accuracy benefit more from the pre-trained global model. Our findings highlight the advantages of leveraging the generalized knowledge from global models to enhance specifically localized soil modelling, positioning a potential paradigm shift in digital soil mapping, and far-reaching implications for soil monitoring and land management.

Deep learning↗

Modeling of Supercritical CO2 Shell-and-Tube Heat Exchangers Under Extreme Conditions: Part II: Heat Exchanger Model

Abstract Heat exchangers play a critical role in supercritical CO2 Brayton cycles by providing necessary waste heat recovery. Supercritical CO2 thermal cycles potentially achieve higher energy density and thermal efficiency operating at elevated temperatures and pressures. Accurate and computationally efficient estimation of heat exchanger performance metrics at these conditions is important for the design and optimization of sCO2 systems and thermal cycles. In this paper (Part II), a computationally efficient and accurate numerical model is developed to predict the performance of shell-and-tube heat exchangers (STHXs). Highly accurate correlations reported in Part I of this study are utilized to improve the accuracy of performance predictions, and the concept of volume averaging is used to abstract the geometry and reduce computation time. The numerical model is validated by comparison with computational fluid dynamics (CFD) simulations and provides high accuracy and significantly lower computation time compared to existing numerical models. A preliminary optimization study is conducted and the advantage of using supercritical CO2 as a working fluid for energy systems is demonstrated.

Engineering↗

Toward Consistent High-Fidelity Quantum Learning on Unstable Devices via Efficient In-Situ Calibration

In the near-term noisy intermediate-scale quantum (NISQ) era, high noise will significantly reduce the fidelity of quantum computing. What's worse, recent works reveal that the noise on quantum devices is not stable, that is, the noise is dynamically changing over time. This leads to an imminent challenging problem: At run-time, is there a way to efficiently achieve a consistent high-fidelity quantum system on unstable devices? To study this problem, we take quantum learning (a.k.a., variational quantum algorithm) as a vehicle, which has a wide range of applications, such as combinatorial optimization and machine learning. A straightforward approach is to optimize a variational quantum circuit (VQC) with a parameter-shift approach on the target quantum device before using it; however, the optimization has an extremely high time cost, which is not practical at run-time. To address the pressing issue, in this paper, we proposed a novel quantum pulse-based noise adaptation framework, namely QuPAD. In the proposed framework, first, we identify that the CNOT gate is the fidelity bottleneck of the conventional VQC, and we employ a more robust parameterized multi-qubit gate (i.e., Rzx gate) to replace CNOT gate. Second, by benchmarking Rzx gate with different parameters, we build a fitting function for each coupling qubit pair, such that the deviation between the theoretic output of Rzx gate and its on-device output under a given pulse amplitude and duration can be efficiently predicted. On top of this, an evolutionary algorithm is devised to identify the pulse amplitude and duration of each Rzx gate (i.e., calibration) and find the quantum circuits with high fidelity. Experiments show that the runtime on quantum devices of QuPAD with 8–10 qubits is less than 15 minutes, which is up to 270 x faster than the parameter-shift approach. In addition, compared to the vanilla VQC as a baseline, QuPAD can achieve 59.33% accuracy gain on a classification task, and average 66.34% closer to ground state energy for molecular simulation.

Hu, Zhirui↗

Performance Evaluation of Intelligent Solar Control Software Through Hardware-in-the-Loop (CRADA Final Report)

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been developed by Latimer Controls, Inc. to estimate the headroom of large PV plants for grid operation and control; however, these technologies lack comprehensive validation under real-world application scenarios. Latimer Controls, Inc. received two voucher awards for research at a national laboratory from the Department of Energy American Made Solar Prize Round 6. The National Renewable Energy Laboratory (NREL) was selected to collaborate with Latimer staff to conduct a performance evaluation of Latimer PV control software. The NREL team will develop a hardware-in-the-loop (HIL) testbed to perform testing and validation of the Latimer PV control technology in a de-risked yet realistic testbed environment. Latimer and NREL worked together to analyze the test data, draw conclusions from the results, and disseminate the resulting scientific findings. In this CRADA work, we propose to test and validate the real-world application of the Latimer Control solution in an HIL environment. We evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. In particular, a data-driven potential high limit (PHL) estimation is developed for large solar plants to accurately estimate their headroom so that they have fast and short-time regulation and control capability to participate in grid services and respond to grid signals in real time (e.g., AGC). This PHL estimation algorithm is embedded in a hardware power plant controller (PPC) and tested with an IEEE-39 bus system model developed in RTDS. To account for the varying cloud conditions and diverse inverter dispatches, we developed a 135-MW PV plant with detailed modeling of 27 individual PV modules and inverters using RTDS. The real-world communications used in such big plants, such as ModBus TCP/IP for inverter level and DNP3 for plant level, were developed to emulate the real-world applications in big PV plants. The ML-based PHL estimation method is tested under nine separate weather scenarios against the ‘reference-control’ solution, hereafter referred to as the baseline solution. The baseline method reserves a subset of inverters (reference group) to operate at their PHL at all times and dispatches only the remaining inverters (control group) at curtailed levels to fulfill the flexibility need. Despite being successfully piloted by NREL in California in 2017 and Chile in 2020, there exist two gaps in the state of the art to fully unlock the flexibility of PV plants: a. There is a trade-off between the PHL estimation accuracy and the flexibility range. b. There lacks granularity in the PHL estimation to capture the variation across inverters. The Latimer solution seeks to address these gaps by applying machine learning methods to improve PHL estimation accuracy while accounting for variability at every inverter. Performance metrics were taken from the 2023 Georgia Power CARES utility-scale RFP. The results demonstrate that the ML-based approach outperforms the traditional baseline method in PHL estimation accuracy for 7 of 9 scenarios. The average PHL error across the nine scenarios was 7.40% for the ML-based method, 2.06% less than the 9.46% PHL error average across scenarios that was exhibited by the baseline method. Additionally, the PHL error was below 5% for at least 95% of the testing interval for 3 of 9 tested intervals with the ML approach, whereas it did not achieve this metric for any of the baseline tests. Overall, simulation results indicate the superior performance of an ML-based approach compared to the conventional baseline reference-control approach, showcasing its potential to support grid stability and operational efficiency. This laboratory HIL testing using real PPC, representative power system simulation models in real-time with detailed PV plant and inverter models, and real-world communication protocols gives us confidence that this machine learning based PHL estimation algorithm works well in the hardware PPC and therefore de-risks future field commissioning. The end goal of this project is to advance grid technology to address the grid operation challenges brought by solar plant’s variability and uncertainties in power generation.

14 SOLAR ENERGY↗

Abundance of Major Cell Wall Components in Natural Variants and Pedigrees of Populus trichocarpa

The rapid analysis of biopolymers including lignin and sugars in lignocellulosic biomass cell walls is essential for the analysis of the large sample populations needed for identifying heritable genetic variation in biomass feedstocks for biofuels and bioproducts. In this study, we reported the analysis of cell wall lignin content, syringyl/guaiacyl (S/G) ratio, as well as glucose and xylose content by high-throughput pyrolysis-molecular beam mass spectrometry (py-MBMS) for >3,600 samples derived from hundreds of accessions of Populus trichocarpa from natural populations, as well as pedigrees constructed from 14 parents (7 × 7). Partial Least Squares (PLS) regression models were built from the samples of known sugar composition previously determined by hydrolysis followed by nuclear magnetic resonance (NMR) analysis. Key spectral features positively correlated with glucose content consisted of m/z 126, 98, and 69, among others, deriving from pyrolyzates such as hydroxymethylfurfural, maltol, and other sugar-derived species. Xylose content positively correlated primarily with many lignin-derived ions and to a lesser degree with m/z 114, deriving from a lactone produced from xylose pyrolysis. Models were capable of predicting glucose and xylose contents with an average error of less than 4%, and accuracy was significantly improved over previously used methods. The differences in the models constructed from the two sample sets varied in training sample number, but the genetic and compositional uniformity of the pedigree set could be a potential driver in the slightly better performance of that model in comparison with the natural variants. Broad-sense heritability of glucose and xylose composition using these data was 0.32 and 0.34, respectively. In summary, we have demonstrated the use of a single high-throughput method to predict sugar and lignin composition in thousands of poplar samples to estimate the heritability and phenotypic plasticity of traits necessary to develop optimized feedstocks for bioenergy applications.

09 BIOMASS FUELS↗

Does Bayesian model averaging improve polynomial extrapolations? Two toy problems as tests

We assess the accuracy of Bayesian polynomial extrapolations from small parameter values, x, to large values of x. We consider a set of polynomials of fixed order, intended as a proxy for a fixed-order effective field theory (EFT) description of data. We employ Bayesian model averaging (BMA) to combine results from different order polynomials (EFT orders). Our study considers two 'toy problems' where the underlying function used to generate data sets is known. We use Bayesian parameter estimation to extract the polynomial coefficients that describe these data at low x. A 'naturalness' prior is imposed on the coefficients, so that they are $\mathcal{O}(1)$. We BMA different polynomial degrees by weighting each according to its Bayesian evidence and compare the predictive performance of this BMA with that of the individual polynomials. In conclusion, the credibility intervals on the BMA forecast have the stated coverage properties more consistently than does the highest evidence polynomial, though BMA does not necessarily outperform every polynomial.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Integration and validation of some modules for modelling of high-speed chemically reactive flows in two-phase gas-droplet mixtures

Three modules are integrated into the built-in OpenFOAM rhoCentralFoam solver towards accurate and efficient modelling of high-speed chemically reactive flows in two-phase gas-droplet mixtures within the OpenFOAM 10.0 framework. The first module is the mixture-averaged diffusion model. The second module is the built-in OpenFOAM Lagrangian solver coupled with optimised droplet drag coefficient and convective heat transfer coefficient sub-models. The last module is a sparse stiff chemistry solver based on dynamic adaptive hybrid integration (AHI-S). The optimised droplet sub-models are first verified in correct implementation for subsequent simulations in this work. Further, they show good accuracy against experimental and analytical data in the modelling of ammonia droplet acceleration and cooling in the flowing and/or low-temperature air. The accuracy and efficiency gains related to the mixture-averaged diffusion model and the AHI-S chemistry solver are examined by simulating 1-D detonation propagation in ammonia droplet-free/laden ammoniaoxygen mixtures. Numerical results of detonation propagation speed, gaseous temperature, density, and species distributions around the induction zone show good agreement with experimental data and analytical solutions. Compared to the built-in OpenFOAM diffusion model, the mixture-averaged diffusion model provides different numerical predictions of pulsating instabilities in detonation propagation. It shows better accuracy in depicting the detonation structure within the droplet-free section attributed to improved multi-component diffusion modelling. Compared to the built-in OpenFOAM solver EulerImplicit (backward Euler), the AHI-S chemistry solver reduces the computational cost by around 50%. It achieves satisfactory accuracy in calculating detonation propagation speed within the droplet-free section with the optimal efficiency when the safety factor, β, equals 0.5.

42 ENGINEERING↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Local Bayesian Dirichlet mixing of imperfect models

Abstract To improve the predictability of complex computational models in the experimentally-unknown domains, we propose a Bayesian statistical machine learning framework utilizing the Dirichlet distribution that combines results of several imperfect models. This framework can be viewed as an extension of Bayesian stacking. To illustrate the method, we study the ability of Bayesian model averaging and mixing techniques to mine nuclear masses. We show that the global and local mixtures of models reach excellent performance on both prediction accuracy and uncertainty quantification and are preferable to classical Bayesian model averaging. Additionally, our statistical analysis indicates that improving model predictions through mixing rather than mixing of corrected models leads to more robust extrapolations.

97 MATHEMATICS AND COMPUTING↗