Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Future Climate Projections for South Florida: Improving the Accuracy of Air Temperature and Precipitation Extremes With a Hybrid Statistical Bias Correction Technique

Projecting future climate variables is essential for comprehending the potential impacts on hydroclimatic hazards like floods and droughts. Evaluating these impacts is challenging due to the coarse spatial resolution of global climate models (GCMs); therefore, bias correction is widely used. Here, we applied two statistical methods—standard empirical quantile mapping (EQM) and a hybrid approach, EQM with linear correction (EQM-LIN)—to bias correct precipitation and air temperature simulated by nine GCMs. We used historical observations from 20 weather stations across South Florida to project future climate under three shared socioeconomic pathways (SSPs). Compared to the EQM, the hybrid EQM-LIN method improved R 2 of daily quantiles by up to 30% over the historical period and improved MAE up to 70% in months that contain most extreme values. Projected extreme precipitation at the weather stations showed that, compared to the EQM-LIN, the EQM method underestimates the high quantiles by up to 26% in SSP585. The projected changes in annual maximum precipitation from historical period (1985–2014) to near future (2040–2069) and far future (2070–2100) were between 2% and 16% across the study area. Projected future precipitation suggested a slight decrease during summer but an increase in fall. This, along with rising summer temperatures, suggested that South Florida can experience rapid oscillations from warmer summers and increased flooding in fall under future climate. Additionally, our comparative analyses with globally and nationally downscaled studies showed that such coarse scale studies do not represent the climatic extremes well, particularly for high quantile precipitation.

54 ENVIRONMENTAL SCIENCES↗

Projected income data under different shared socioeconomic pathways for Washington state

Abstract High-resolution income projections under different Shared Socioeconomic Pathways (SSPs) are essential for the climate change research communities to devise climate change adaptation and mitigation strategies. To generate income projections for Washington state, we obtain state-level GDP per capita projections and convert them into projected annual household income. The resulting state-level income projections are subsequently downscaled to the census block-level based on the Longitudinal Origin-Destination Employment Statistics (LODES) dataset. For accuracy assessment, we downscale historical income data from state- level to block- and block group-level and compare the downscaled results against the actual income data from LODES. County-level accuracy assessment is also conducted based on American Community Survey. The results demonstrate a good agreement (Average R 2 of 0.67, 0.8, and 0.99 for block-, block group-, and county-level, respectively) between the downscaled income data and the reference data, thereby validating the methodology employed. Our approach is applicable to other states for income projections, which can be utilized by a broader audience, including those involved in demographic analysis, economic research, and urban planning.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Predictive understanding of the surface tension and velocity of sound in ionic liquids using machine learning

Knowledge of the physical properties of ionic liquids (ILs), such as the surface tension and speed of sound, is important for both industrial and research applications. Unfortunately, technical challenges and costs limit exhaustive experimental screening efforts of ILs for these critical properties. Previous work has demonstrated that the use of quantum-mechanics-based thermochemical property prediction tools, such as the conductor-like screening model for real solvents, when combined with machine learning (ML) approaches, may provide an alternative pathway to guide the rapid screening and design of ILs for desired physiochemical properties. However, the question of which machine-learning approaches are most appropriate remains. In the present study, we examine how different ML architectures, ranging from tree-based approaches to feed-forward artificial neural networks, perform in generating nonlinear multivariate quantitative structure–property relationship models for the prediction of the temperature- and pressure-dependent surface tension of and speed of sound in ILs over a wide range of surface tensions (16.9–76.2 mN/m) and speeds of sound (1009.7–1992 m/s). The ML models are further interrogated using the powerful interpretation method, shapley additive explanations. We find that several different ML models provide high accuracy, according to traditional statistical metrics. The decision tree-based approaches appear to be the most accurate and precise, with extreme gradient-boosting trees and gradient-boosting trees being the best performers. However, our results also indicate that the promise of using machine-learning to gain deep insights into the underlying physics driving structure–property relationships in ILs may still be somewhat premature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection

Communication overhead in federated learning (FL) poses a significant challenge for network anomaly detection systems, where the myriad of client configurations and network conditions can severely impact system efficiency and detection accuracy. While existing approaches attempt to address this through individual optimization techniques, they often fail to maintain the delicate balance between reduced overhead and detection performance. This paper presents an adaptive FL framework that dynamically combines batch size optimization, client selection, and asynchronous updates to achieve efficient anomaly detection. Through extensive profiling and experimental analysis on two distinct datasets-UNSW-NBIS for general network traffic and ROAD for automotive networks-our framework reduces communication overhead by 97.6%; (from 700.0s to 16.8s) compared to synchronous baseline approaches while maintaining comparable detection accuracy (95.10%; vs. 95.12%;). Statistical validation using Mann-Whitney U test confirms significant improvements (p < 0.05) over existing FL approaches across both datasets, demonstrating the framework's adaptability to different network security contexts. Detailed profiling analysis reveals the efficiency gains through dramatic reductions in GPU operations and memory transfers while maintaining robust detection performance under varying client conditions.

Marfo, William [University of Texas at El Paso]↗

Single-shot picosecond pump coherent Rayleigh scattering thermometry

A single-shot coherent Rayleigh scattering (CRS) technique capable of measurement times less than 10 ns is presented. Here, the use of a mode-locked picosecond pump laser yields repeatable electrostrictive forcing compared to previous CRS experiments using unseeded nanosecond pump pulses, which are beset by shot-to-shot variations. Quantitative measurements are achieved by dispersing the CRS signal onto an EMCCD sensor using a virtually imaged phased array and comparing the experimental spectra to an existing kinetic model with a least-squares fitting routine. Measurements are demonstrated at ambient and low-pressure (2 Torr), low-temperature (100 K) conditions where the CRS measurement is within the collisionless regime. Single-shot statistics indicated precision and accuracy within 4% at ambient conditions and within 8% at low density and temperature conditions. This single-shot CRS technique is a powerful diagnostic tool with potential for multi-parameter measurements in complex flow environments, including high-speed aerodynamic ground test facilities.

Senior, William Charles Bowman [Sandia National La↗

Network Anomaly Detection in Distributed Edge Computing Infrastructure

As networks continue to grow in complexity and scale, detecting anomalies has become increasingly challenging, particularly in diverse and geographically dispersed environments. Traditional approaches often struggle with managing the computational burden associated with analyzing large-scale network traffic to identify anomalies. This paper introduces a distributed edge computing framework that integrates federated learning with Apache Spark and Kubernetes to address these challenges. We hypothesize that our approach, which enables collaborative model training across distributed nodes, significantly enhances the detection accuracy of network anomalies across different network types. We show that by leveraging distributed computing and containerization technologies, our framework not only improves scalability and fault tolerance but also achieves superior detection performance compared to state-of-the-art methods. Extensive experiments on the UNSW-NB15 and ROAD datasets validate the effectiveness of our approach, demonstrating statistically significant improvements in detection accuracy and training efficiency over baseline models, as confirmed by MannWhitney U and Kolmogorov-Smirnov tests (p<0.05).

Marfo, William [University of Texas at El Paso,Dep↗

KiDS-1000 cosmology: Combined second- and third-order shear statistics

Aims.In this work, we perform the first cosmological parameter analysis of the fourth release of Kilo Degree Survey (KiDS-1000) data with second- and third-order shear statistics. This paper builds on a series of studies aimed at describing the roadmap to third-order shear statistics. Methods.We derived and tested a combined model of the second-order shear statistic, namely, the COSEBIs and the third-order aperture mass statistics 〈ℳ ap 3 〉 in a tomographic set-up. We validated our pipeline withN-body mock simulations of the KiDS-1000 data release. To model the second- and third-order statistics, we used the latest version of HMCODE2020 for the power spectrum and BIHALOFITfor the bispectrum. Furthermore, we used an analytic description to model intrinsic alignments and hydro-dynamical simulations to model the effect of baryonic feedback processes. Lastly, we decreased the dimension of the data vector significantly by considering only equal smoothing radii for the 〈ℳ ap 3 〉 part of the data vector. This makes it possible to carry out a data analysis of the KiDS-1000 data release using a combined analysis of COSEBIs and third-order shear statistics. Results.We first validated the accuracy of our modelling by analysing a noise-free mock data vector, assuming the KiDS-1000 error budget, finding a shift in the maximum of the posterior distribution of the matter density parameter, ΔΩ m < 0.02 σ Ω m , and of the structure growth parameter, ΔS 8 < 0.05 σ S 8 . Lastly, we performed the first KiDS-1000 cosmological analysis using a combined analysis of second- and third-order shear statistics, where we constrained Ω m = 0.248 −0.055 +0.062 andS 8 = σ 8 √(Ω m /0.3 )= 0.772 ± 0.022. The geometric average on the errors of Ω m andS 8 of the combined statistics decreases, compared to the second-order statistic, by a factor of 2.2.

Astronomy & Astrophysics↗

Calibration and Commissioning of the LSST

Understanding the nature of dark energy and dark matter remains one of the fundamental questions in physics today; impacting our understanding of particle physics, cosmology, and possibly theories of gravity. Given the scale and complexity of the next generation of cosmology experiments (e.g., the Rubin Observatory, the Euclid satellite mission, and the Roman space telescope) we are entering an era where statistical noise no longer determines the accuracy to which we can measure cosmological parameters. Our ability to control and correct for systematics will ultimately determine the scientific impact of these experiments. This award addressed the challenge of how we determine what limits the accuracy of our cosmological measures, what techniques are appropriate for measuring and calibrating the properties of galaxies to best constrain cosmological models, how to develop statistical techniques that are insensitive to systematic errors, and how to optimize survey strategies in order to minimize systematics while maximizing the speed at which an experiment can achieve its science objectives. In this final technical report for award DE-SC0011635 we describe a set of open-source frameworks that simulate the characteristics and properties of current and planned cosmology surveys and the application of these frameworks to the development of new methodologies for estimating the properties and distances to galaxies that are robust to noisy and incomplete data.

79 ASTRONOMY AND ASTROPHYSICS↗

Improved Subseasonal Forecasting of Extreme Polar Vortices Using Machine Learning

Our research was focused on forecasting the position and shape of the winter stratospheric polar vortex at a subseasonal timescale of 15 days in advance. To achieve this, we employed both statistical and neural network machine learning techniques. The analysis was performed on 42 winter seasons of reanalysis data provided by NASA giving us a total of 6,342 days of data. The state of the polar vortex for determined by using geometric moments to calculate the centroid latitude and the aspect ratio of an ellipse fit onto the vortex. Timeseries for thirty additional precursors were calculated to help improve the predictive capabilities of the algorithm. Feature importance of these precursors was performed using random forest to measure the predictive importance and the ideal number of precursors. Then, using the precursors identified as important, various statistical methods were tested for predictive accuracy with random forest and nearest neighbor performing the best. An echo state network, a type of recurrent neural network that features sparsely connected hidden layer and a reduced number of trainable parameters that allows for rapid training and testing, was also implemented for the forecasting problem. Hyperparameter tuning was performed for each methods using a subset of the training data. The algorithms were trained and tuned on the first 41 years of data, then tested for accuracy on the final year. In general, the centroid latitude of the polar vortex proved easier to predict than the aspect ratio across all algorithms. Random forest outperformed other statistical forecasting algorithms overall but struggled to predict extreme values. Forecasting from echo state network suggested a strong predictive capability past 15 days, but further work is required to fully realize the potential of recurrent neural network approaches.

54 ENVIRONMENTAL SCIENCES↗

Experimental validation of a high fidelity Monte Carlo neutron transport model of the MIT graphite exponential pile

High-fidelity modeling and simulation were performed for the MIT graphite exponential pile (MGEP) using Monte Carlo neutron transport codes OpenMC and MCNP, and the results were validated by experimental data. The MGEP is being used as the test bed for the design of an autonomous control system for the pile's neutron flux distribution. The main contribution of this work is to generate the training data sets of neutron flux distributions with different locations of control rods that perturb the neutron flux profiles. First, code -to-code cross verification between OpenMC and MCNP was performed to ensure consistency of the numerical modeling within statistical uncertainties. To validate the accuracy of this high-fidelity model, a series of neutron flux measurements were conducted using a Helium-3 (He-3) neutron detector on a mobile platform that is placed inside the pile. Second, the neutron flux profiles were measured in four vertical layers of interest, and compared to the corresponding simulation results. The comparison results shows that the root mean square error is less than 2.5% in the two upper layers, and less than 4.5% in all four measured layers. Here the results validated the accuracy of the modeling and simulation. Finally, the relative change of the neutron flux profiles from moving control rods was analyzed, which identified the layer that has the best sensitivity regarding the control rods movements. Thus, this work identified and provided training data sets of both simulated and experimental neutron flux profiles in the most sensitive layer, paving the path forward to the real-time experimental demonstration of the autonomous control system.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Hierarchical-embedding autoencoder with a predictor as efficient architecture for learning time-evolution in multi-scale turbulent flows

We introduce a scale-aware, data-driven deep learning modeling framework for accurately predicting the time evolution of multi-scale turbulent plasma and liquid flows. The approach is motivated by the idea of scale separation. Structures of vastly different length scales emerge in these systems, and interactions between these structures occur only locally. To exploit this structure, the flow state is transformed by a hierarchical, fully convolutional autoencoder, not into a single embedding layer as in conventional convolutional surrogate models, but into a series of embedding layers. A stepwise training strategy ensures that fine-scale features are encoded on a high-resolution grid, while larger structures are represented on progressively coarser layers. The time evolution predictor advances all embedding layers in sync, capturing local interactions between features at the same scale as well as between all scales. This approach enables efficient modeling of multi-scale systems since negligible interactions between distant, small-scale structures do not need to be directly modeled. Our hierarchical-embedding autoencoder with a predictor framework is evaluated on canonical examples of multi-scale turbulence: two-dimensional Kolmogorov flow and Hasegawa–Wakatani plasma turbulence. In both cases, the proposed framework significantly improves predictive accuracy relative to conventional convolutional network architectures. A significant improvement in prediction accuracy was observed for crucial statistical characteristics of the Hasegawa–Wakatani plasma as well as for individual trajectories of the Kolmogorov flow turbulence. Importantly, the model's rollout for the Hasegawa–Wakatani problem demonstrates a four-order-of-magnitude speedup compared to traditional numerical solvers.

Khrabry, Alexander I. [Princeton Univ., NJ (United↗

Explicit simulation of the Brownian rotation of arbitrary shaped aerosol particles using quaternions

The shape of an aerosol particle strongly influences its mass and momentum transfer cross-sections, charging properties, and other physical properties. Here, we present an explicit time-stepping procedure to simulate the rotational Brownian motion of arbitrary shaped aerosol particles by solving Euler’s equation of rotation. A Langevin formulation of the rotation equations is used, wherein Brownian motion due to thermal collisions between a particle and background gas molecules is represented using a stochastic fluctuating torque and fluid resistance is included as a drag torque. To avoid singularities associated with describing the orientation of a shape with Euler angles, we employ a quaternion formulation that leads to first-order stochastic differential equations to describe the evolution of the angular position and angular velocity of a rigid body. We perform all the rotational dynamics calculations in the body-fixed frame of reference attached to the rotating shape whose basis vectors are the normalized eigenvectors of the inertia tensor of the particle. Numerical solutions to rotation under torque-free conditions, damped rotation without Brownian motion, and stochastic rotation for arbitrary shapes are presented and discussed. The presented method enables time-resolved simulation of Brownian rotation for direct comparison with experimentally measured trajectories or statistical measures. The second order accuracy of the used time-stepping procedure places a severe restriction on the timestep that can be used for obtaining accurate results. Animations of presented simulations are included for visualizing rotational motion at various gas pressures. To aid implementation, MATLAB ® codes are also provided. Extension to include translation Brownian motion is straightforward.

Roy, Mrittika↗

The development and application of the stirred‐reactor coupon analysis (SRCA) test method

A new technique, termed the stirred‐reactor coupon analysis (SRCA) method, has been developed to measure the rate of glass dissolution in forward‐rate conditions. Monolithic glass coupons are partially masked with an inert material before placement in a large volume of well‐mixed solution with known chemistry and temperature for a predetermined duration. After the test, the mask is removed, and the difference in step height between the protected area and the exposed corroded portions of the sample coupon is measured to determine the extent of glass dissolution. The step height is converted to a rate measurement using the test duration and glass density. Test parameters such as sample surface preparation and test duration were evaluated to determine their effects on the measured rates. Additionally, results from an interlaboratory study (ILS) consisting of 12 laboratories from 11 different institutions are presented, where each laboratory performed 12 independent tests. When removing experimental outlier data, the 95% reproducibility limits for the SRCA method has no statistical difference with previously published standardized test methods used to determine the forward rate of glass dissolution. Overall, this paper describes steps necessary to perform the test method and provides the statistical calculations to evaluate test accuracy.

chemical durability↗

Optimization of the deep neural network parameters for generating homogenized fuel assembly data for nodal codes

Homogenized fuel assembly (FA) data is a typical input data for nodal codes. Generating that data, however, could be time-consuming. One of promising ways to mitigate the computational burden of generating macroscopic cross-sections is to use trained artificial neural network (ANN) models for predicting nuclear data. However, there is a challenge to make the model support variable FA geometry. In this work, two most common types of FA were combined in one ANN model. Since there could be multiple ways of converting 2-dimensional FA data into 1-dimensional input vector for ANN, three different approaches of data flattening were evaluated. The input parameters included each fuel pin enrichment, fuel temperature, moderator temperature and boron concentration. The output parameters were 2-group macroscopic cross-sections (XS) and pin power distribution (HFF). A fully connected deep neural network (DNN) model was trained and tested using pre-generated data obtained with lattice physics code STREAM. The results of this study showed no statistically significant difference in the accuracy of XS and HFF generation for all 3 tested input vector orders. This means that fully connected DNN for XS generation demonstrated input sequence invariance. Results of comparing predicted XS data with reference solutions were found sufficiently close considering the reduction of computation time offered by ANN. Mean relative difference (MRD) for all output XS parameters was found below 0.7%, while HFF MRD was found higher compared to XS values, in some cases slightly exceeding 1%, mostly near guide tube locations. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

DESI mock challenge: Halo and galaxy catalogues with the bias assignment method

We present a novel approach to the construction of mock galaxy catalogues for large-scale structure analysis based on the distribution of dark matter halos obtained with effective bias models at the field level. We aim to produce mock galaxy catalogues capable of generating accurate covariance matrices for a number of cosmological probes that are expected to be measured in current and forthcoming galaxy redshift surveys (e.g. two- and three-point statistics). The construction of the catalogues shown in this paper is part of a mock-comparison project within the Dark Energy Spectroscopic Instrument (DESI) collaboration. We use the bias assignment method ( BAM ) to model the statistics of halo distribution through a learning algorithm using a few detailed N-body simulations, and approximated gravity solvers based on Lagrangian perturbation theory. We introduce cosmic-web-dependent corrections to modelling redshift-space distortions at the N-body level – both in the halo and galaxy distributions –, as well as a multi-scale approach for accurate assignment of halo properties. Using specific models of halo occupation distributions to populate halos, we generate galaxy mocks with the expected number density and central-satellite fraction of emission-line galaxies, which are a key target of the DESI experiment. BAM generates mock catalogues with per cent accuracy in a number of summary statistics, such as the abundance, the two- and three-point statistics of halo distributions, both in real and redshift space. In particular, the mock galaxy catalogues display ~3%-10% accuracy in the multipoles of the power spectrum up to scales of k ~ 0.4 h -1 Mpc. We show that covariance matrices of two- and three-point statistics obtained with BAM display a similar structure to the reference simulation. BAM offers an efficient way to produce mock halo catalogues with accurate two- and three-point statistics and is able to generate a variety of multi-tracer catalogues with precise covariance matrices of several cosmological probes. We discuss future developments of the algorithm towards mock production in DESI and other galaxy-redshift surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Dark Energy Survey: Modeling strategy for multiprobe cluster cosmology and validation for the Full Six-year Dataset

We introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3x2pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (1) identification of modeling and parameterization choices and impact studies using simulated analyses with each possible model misspecification (2) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3x2pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy--galaxy lensing and cosmic shear two-point statistics) and the CL+3x2pt+BAO+SN (combination of CL+3x2pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dark energy survey: Modeling strategy for multiprobe cluster cosmology and validation for the full six-year dataset

Here, we introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3×2 pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (i) identification of modeling and parametrization choices and impact studies using simulated analyses with each possible model misspecification and (ii) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3×2 pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy–galaxy lensing and cosmic shear two-point statistics) and the CL+3×2 pt+BAO+SN (combination of CL+3×2 pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

79 ASTRONOMY AND ASTROPHYSICS↗

Nonlinear encoding in diffractive information processing using linear optical materials

Nonlinear encoding of optical information can be achieved using various forms of data representation. Here, we analyze the performances of different nonlinear information encoding strategies that can be employed in diffractive optical processors based on linear materials and shed light on their utility and performance gaps compared to the state-of-the-art digital deep neural networks. For a comprehensive evaluation, we used different datasets to compare the statistical inference performance of simpler-to-implement nonlinear encoding strategies that involve, e.g., phase encoding, against data repetition-based nonlinear encoding strategies. We show that data repetition within a diffractive volume (e.g., through an optical cavity or cascaded introduction of the input data) causes the loss of the universal linear transformation capability of a diffractive optical processor. Therefore, data repetition-based diffractive blocks cannot provide optical analogs to fully connected or convolutional layers commonly employed in digital neural networks. However, they can still be effectively trained for specific inference tasks and achieve enhanced accuracy, benefiting from the nonlinear encoding of the input information. Our results also reveal that phase encoding of input information without data repetition provides a simpler nonlinear encoding strategy with comparable statistical inference accuracy to data repetition-based diffractive processors. Our analyses and conclusions would be of broad interest to explore the push-pull relationship between linear material-based diffractive optical systems and nonlinear encoding strategies in visual information processors.

42 ENGINEERING↗