Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Likelihood Methods for CMB Experiments

A great deal of experimental effort is currently being devoted to the precise measurements of the cosmic microwave background (CMB) sky in temperature and polarization. Satellites, balloon-borne, and ground-based experiments scrutinize the CMB sky at multiple scales, and therefore enable to investigate not only the evolution of the early Universe, but also its late-time physics with unprecedented accuracy. The pipeline leading from time ordered data as collected by the instrument to the final product is highly structured. Moreover, it has also to provide accurate estimates of statistical and systematic uncertainties connected to the specific experiment. In this paper, we review likelihood approaches targeted to the analysis of the CMB signal at different scales, and to the estimation of key cosmological parameters. We consider methods that analyze the data in the spatial (i.e., pixel-based) or harmonic domain. We highlight the most relevant aspects of each approach and compare their performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Sheaf Theoretical Approach to Uncertainty Quantification of Heterogeneous Geolocation Information

Integration of multiple, heterogeneous sensors is a challenging problem across a range of applications. Prominent among these are multi-target tracking, where one must combine observations from different sensor types in a meaningful and efficient way to track multiple targets. Because different sensors have differing error models, we seek a theoretically justified quantification of the agreement among ensembles of sensors, both overall for a sensor collection, and also at a fine-grained level specifying pairwise and multi-way interactions among sensors. We demonstrate that the theory of mathematical sheaves provides a unified answer to this need, supporting both quantitative and qualitative data. Furthermore, the theory provides algorithms to globalize data across the network of deployed sensors, and to diagnose issues when the data do not globalize cleanly. We demonstrate and illustrate the utility of sheaf-based tracking models based on experimental data of a wild population of black bears in Asheville, North Carolina. A measurement model involving four sensors deployed among the bears and the team of scientists charged with tracking their location is deployed. This provides a sheaf-based integration model which is small enough to fully interpret, but of sufficient complexity to demonstrate the sheaf’s ability to recover a holistic picture of the locations and behaviors of both individual bears and the bear-human tracking system. A statistical approach was developed in parallel for comparison, a dynamic linear model which was estimated using a Kalman filter. This approach also recovered bear and human locations and sensor accuracies. When the observations are normalized into a common coordinate system, the structure of the dynamic linear observation model recapitulates the structure of the sheaf model, demonstrating the canonicity of the sheaf-based approach. However, when the observations are not so normalized, the sheaf model still remains valid.

97 MATHEMATICS AND COMPUTING↗

Selecting Critical Scenarios of DER Adoption in Distribution Grids Using Bayesian Optimization

We develop a new methodology to select scenarios of DER adoption most critical for distribution grids. Anticipating risks of future voltage and line flow violations due to additional PV adopters is central for utility investment planning but continues to rely on deterministic or ad hoc scenario selection. We propose a highly efficient search framework based on multi-objective Bayesian Optimization. We treat underlying grid stress metrics as computationally expensive black-box functions, approximated via Gaussian Process surrogates and design an acquisition function based on probability of scenarios being Pareto-critical across a collection of line- and bus-based violation objectives. Our approach provides a statistical guarantee and offers an order of magnitude speed-up relative to a conservative exhaustive search. Case studies on realistic feeders with 200-400 buses demonstrate the effectiveness and accuracy of our approach.

Mulkin, Olivier↗

Testing biasedness of self-reported microbusiness innovation in the annual business survey

This study tests for potential bias in self-reported innovation due to the inclusion of a research and development (R&D) module that only microbusinesses (less than 10 employees) receive in the Annual Business Survey (ABS). Previous research found that respondents to combined innovation/R&D surveys reported innovation at lower rates than respondents to innovation-only surveys. A regression discontinuity design is used to test whether microbusinesses, which constitute a significant portion of U.S. firms with employees, are less likely to report innovation compared to other small businesses. In the vicinity of the 10-employee threshold, the study does not detect statistically significant biases for new-to-market and new-to-business product innovation. Statistical power analysis confirms the nonexistence of biases with a high power. Comparing the survey design of ABS to earlier combined innovation/R&D surveys provides valuable insights for the proposed integration of multiple Federal surveys into a single enterprise platform survey. The findings also have important implications for the accuracy and reliability of innovation data used as an input to policymaking and business development strategies in the United States.

99 GENERAL AND MISCELLANEOUS↗

A Statistical Evaluation of WRF-LES Trace Gas Dispersion Using Project Prairie Grass Measurements

In recent years, new measurement systems have been deployed to monitor and quantify methane emissions from the natural gas sector. Large-eddy simulation (LES) has complemented measurement campaigns by serving as a controlled environment in which to study plume dynamics and sampling strategies. However, with few comparisons with controlled-release experiments, the accuracy of LES for modeling natural gas emissions is poorly characterized. In this paper, we evaluate LES from the Weather Research and Forecasting (WRF) Model against Project Prairie Grass campaign measurements and surface layer similarity theory. Using WRF-LES, we simulate continuous emissions from 30 near-surface trace gas sources in two stability regimes: strong convection and weak convection. We examine the impact of grid resolutions ranging from 6.25 to 52 m in the horizontal dimension on model results. We evaluate performance in a statistical framework, calculating fractional bias and conducting Welch’s t tests. WRF-LES accurately simulates observed surface concentrations at 100 m and beyond under strong convection; simulated concentrations pass t tests in this region irrespective of grid resolution. However, in weakly convective conditions with strong winds, WRF-LES substantially overpredicts concentrations—the magnitude of fractional bias often exceeds 30%, and all but one t test fails. The good performance of WRF-LES under strong convection correlates with agreement with local free convection theory and a minimal amount of parameterized turbulent kinetic energy. The poor performance under weak convection corresponds to misalignment with Monin–Obukhov similarity theory and a significant amount of parameterized turbulent kinetic energy.

17 WIND ENERGY↗

Leveraging interpolation models and error bounds for verifiable scientific machine learning

Effective verification and validation techniques for modern scientific machine learning workflows are challenging to devise. Statistical methods are abundant and easily deployed, but often rely on speculative assumptions about the data and methods involved. Error bounds for classical interpolation techniques can provide mathematically rigorous estimates of accuracy, but often are difficult or impractical to determine computationally. Here, in this work, we present a best-of-both-worlds approach to verifiable scientific machine learning by demonstrating that (1) multiple standard interpolation techniques have informative error bounds that can be computed or estimated efficiently; (2) comparative performance among distinct interpolants can aid in validation goals; (3) deploying interpolation methods on latent spaces generated by deep learning techniques enables some interpretability for black-box models. We present a detailed case study of our approach for predicting lift-drag ratios from airfoil images. Code developed for this work is available in a public Github repository.

97 MATHEMATICS AND COMPUTING↗

Generating high-resolution total canopy SIF emission from TROPOMI data: Algorithm and application

Solar-induced chlorophyll fluorescence (SIF) is a rapidly advancing front in modeling global terrestrial gross primary production (GPP). Canopy total SIF emissions (SIF total ) are mechanistically linked to the plant photosynthesis, and can be estimated from satellite observed SIF (SIF obs ) through radiative transfer modeling. However, the current satellite SIF obs and thus SIF total are available only at coarse spatial resolutions from several kilometers to tens of kilometers, inhibiting the application at fine spatial scales. Here, in this work, we proposed an algorithm to generate both global high-resolution SIF total (HSIF total ) and high-resolution SIF obs (HSIF obs ) at 1 km from low-resolution SIF obs (LSIF obs ) from the TROPOspheric Monitoring Instrument (TROPOMI), which has a spatial resolution at nadir of 3.5 km by 5.6–7 km. Our statistical method is based on the law of energy conservation and uses satellite derived fraction of absorbed photosynthetically active radiation, fluorescence efficiency, and the escape probability of fluorescence. We evaluated the accuracy of our HSIF total using the Orbiting Carbon Observatory-2 SIF (R 2 = 0.78). We found that the spatial resolution had clear effects on the relationship between HSIF total and GPP. We also compared HSIF total to 8-day averaged tower GPP from 135 flux sites and found that they were better correlated when HSIF total was averaged over a 1-km radius around the tower than when averaged over a larger radius. Our study provided a unique high-resolution HSIF total product, which will advance the estimation of GPP by extrapolating site-level relationships to the global scale.

54 ENVIRONMENTAL SCIENCES↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

DeepMerge: Classifying high-redshift merging galaxies with deep neural networks

In this work, we investigate and demonstrate the use of convolutional neural networks (CNNs) for the task of distinguishing between merging and non-merging galaxies in simulated images, and for the first time at high redshifts (i.e. $z=2$). We extract images of merging and non-merging galaxies from the Illustris-1 cosmological simulation and apply observational and experimental noise that mimics that from the Hubble Space Telescope; the data without noise form a "pristine" data set and that with noise form a "noisy" data set. The test set classification accuracy of the CNN is $79\%$ for pristine and $76\%$ for noisy. The CNN outperforms a Random Forest classifier, which was shown to be superior to conventional one- or two-dimensional statistical methods (Concentration, Asymmetry, the Gini, $M_{20}$ statistics etc.), which are commonly used when classifying merging galaxies. We also investigate the selection effects of the classifier with respect to merger state and star formation rate, finding no bias. Finally, we extract Grad-CAMs (Gradient-weighted Class Activation Mapping) from the results to further assess and interrogate the fidelity of the classification model.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A hybrid deep neural operator/finite element method for ice-sheet modeling

One of the most challenging and consequential problems in climate modeling is to provide probabilistic projections of sea level rise. A large part of the uncertainty of sea level projections is due to uncertainty in ice sheet dynamics. At the moment, accurate quantification of the uncertainty is hindered by the cost of ice sheet computational models. In this work we develop a hybrid approach to approximate existing ice sheet models at a fraction of their cost. Our approach consists of replacing the finite element model for the momentum equations for the ice velocity, the most expensive part of an ice sheet model, with a Deep Operator Network, while we retain a classic finite element discretization for the evolution of the ice thickness. We show that the resulting hybrid model is very accurate and it is an order of magnitude faster than the traditional finite element model. Further, a distinctive feature of the proposed model, compared to other neural network approaches, is that it can handle high-dimensional parameter spaces (parameter fields) such as the basal friction at the bed of the glacier and can therefore be used for generating samples for uncertainty quantification. Further, we study the impact of hyper-parameters, number of unknowns and correlation length of the parameter distribution on the training and accuracy of the Deep Operator Network on a synthetic ice sheet model. We then target the evolution of the Humboldt glacier in Greenland and show that our hybrid model can provide accurate statistics of the glacier mass loss and can be effectively used to accelerate the quantification of uncertainty.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A deep generative model enables automated structure elucidation of novel psychoactive substances

Over the past decade, the illicit drug market has been reshaped by the proliferation of clandestinely produced designer drugs. These agents, referred to as new psychoactive substances (NPSs), are designed to mimic the physiological actions of better-known drugs of abuse while skirting drug control laws. The public health burden of NPS abuse obliges toxicological, police and customs laboratories to screen for them in law enforcement seizures and biological samples. However, the identification of emerging NPSs is challenging due to the chemical diversity of these substances and the fleeting nature of their appearance on the illicit market. Here, in this study, we present DarkNPS, a deep learning-enabled approach to automatically elucidate the structures of unidentified designer drugs using only mass spectrometric data. Our method employs a deep generative model to learn a statistical probability distribution over unobserved structures, which we term the structural prior. We show that the structural prior allows DarkNPS to elucidate the exact chemical structure of an unidentified NPS with an accuracy of 51% and a top-10 accuracy of 86%. Our generative approach has the potential to enable de novo structure elucidation for other types of small molecules that are routinely analysed by mass spectrometry.

Cheminformatics↗

Proposed Analytical Methods for Determining Filter Media Properties

High Efficiency Particulate Air (HEPA) filters, commonly used in nuclear filtration applications, are an integral part of the waste management processes in nuclear plants. HEPA filters are 99.97% efficient filtration devices characterized by their high resistance to air flow, or pressure drop. From theoretical models, the initial pressure drop across a clean filter is proven to be a function of the filter media fiber diameters and porosity of the media. These filter properties are relatively difficult to obtain with traditional manual methods; therefore, there is a need to develop analytical methods to find the fiber diameter and porosity to determine the pressure drop. Typically, these values can be found by substituting a calculated representative value based upon measured values, such as equivalent fiber diameter based upon a clean pressure drop. However, these values may also be directly recorded and measured by incorporating Scanning Electron Microscope (SEM) image analysis and other measurement methods to analyze the filter media and determine its physical properties without the need for media testing. DiameterJ, an open source Java plug-in, used with ImageJ, can process images of filter media taken by an SEM to find statistical data such as mean fiber diameter and porosity. To produce the raw data, SEM images of two filter media type samples are taken and segmented in DiameterJ using the traditional and statistical region merging segmentation algorithms. Manual segmentation is necessary after the initial segmentation by the algorithms as the images tend to be too complex for the algorithms to output with the necessary accuracy. However, complications exist in the manual segmentation process as these methods can be time intensive and prone to the individual bias of the user. This in turn can skew the final mean fiber diameter result, and lead to either an over or under prediction of the pressure drop. It was also discovered that the porosity data produced by DiameterJ is inaccurate, as the SEM analyzes the three-dimensional filter media by projecting its geometry onto a plane and analyzing it as a two-dimensional binary image. Thus, the porosity is artificially inflated through the segmentation process, rendering this result incorrect. Alternatively, density determination, gravimetric analysis, and thickness testing of the filter media is collectively used to determine the filter fiber porosity. Together, the SEM image analysis and analytical lab methods produce results through direct measurements which allow for the prediction of the initial pressure drop from clean filter media without the need for prior media testing to collect pressure data. By improving upon this proposed analytical method in the future, there is potential to streamline the process of finding these filter properties into a more direct methodology for determining the pressure drop across HEPA filters.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Experimentally validated high-fidelity simulations of a liquid jet in supersonic crossflow

Here, utilizing recent advancements in computational schemes for compressible, multiphase flows, this work features a parametric study of a pure liquid jet in supersonic crossflow that involves simulating the atomization process for four values of momentum-flux ratio. These simulations are validated against experimental results measured with high-speed X-ray imaging, which confirm the accuracy of the numerical approach. Also, the effect of numerical resolution on some flow behavior is investigated, revealing convergence of the jet shape and surface instability wavelength. Analysis of the resulting sprays includes statistical descriptions of the liquid distribution, liquid structures created through breakup, interfacial instabilities, and dominant flow features. As the flowrate increases, the spray penetrates further, it becomes more disperse, and less liquid impacts the wall, but the droplet size distribution changes little. The wavelength of instabilities on the windward side of the jet diverges from measured trends in subsonic crossflows. In a visualization of the time-averaged flow, counter-rotating vortices are observed along the jet core and in the wake, affecting the process of primary atomization and early droplet trajectories.

42 ENGINEERING↗

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE↗

Validation and parameterization of a novel physics-constrained neural dynamics model applied to turbulent fluid flow

We report, in fluid physics, data-driven models to enhance or accelerate time to solution are becoming increasingly popular for many application domains, such as alternatives to turbulence closures, system surrogates, or for new physics discovery. In the context of reduced order models of high-dimensional time-dependent fluid systems, machine learning methods grant the benefit of automated learning from data, but the burden of a model lies on its reduced-order representation of both the fluid state and physical dynamics. In this work, we build a physics-constrained, data-driven reduced order model for Navier–Stokes equations to approximate spatiotemporal fluid dynamics in the canonical case of isotropic turbulence in a triply periodic box. The model design choices mimic numerical and physical constraints by, for example, implicitly enforcing the incompressibility constraint and utilizing continuous neural ordinary differential equations for tracking the evolution of the governing differential equation. We demonstrate this technique on a three-dimensional, moderate Reynolds number turbulent fluid flow. In assessing the statistical quality and characteristics of the machine-learned model through rigorous diagnostic tests, we find that our model is capable of reconstructing the dynamics of the flow over large integral timescales, favoring accuracy at the larger length scales. More significantly, comprehensive diagnostics suggest that physically interpretable model parameters, corresponding to the representations of the fluid state and dynamics, have attributable and quantifiable impact on the quality of the model predictions and computational complexity.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Validation of semi-analytical, semi-empirical covariance matrices for two-point correlation function for early DESI data

ABSTRACT We present an extended validation of semi-analytical, semi-empirical covariance matrices for the two-point correlation function (2PCF) on simulated catalogs representative of luminous red galaxies (LRGs) data collected during the initial 2 months of operations of the Stage-IV ground-based Dark Energy Spectroscopic Instrument (DESI). We run the pipeline on multiple effective Zel’dovich (EZ) mock galaxy catalogs with the corresponding cuts applied and compare the results with the mock sample covariance to assess the accuracy and its fluctuations. We propose an extension of the previously developed formalism for catalogs processed with standard reconstruction algorithms. We consider methods for comparing covariance matrices in detail, highlighting their interpretation and statistical properties caused by sample variance, in particular, non-trivial expectation values of certain metrics even when the external covariance estimate is perfect. With improved mocks and validation techniques, we confirm a good agreement between our predictions and sample covariance. This allows one to generate covariance matrices for comparable data sets without the need to create numerous mock galaxy catalogs with matching clustering, only requiring 2PCF measurements from the data itself. The code used in this paper is publicly available at https://github.com/oliverphilcox/RascalC.

79 ASTRONOMY AND ASTROPHYSICS↗

Online Detection of Inter-Turn Winding Faults in Single-Phase Distribution Transformers Using Smart Meter Data

Turn-to-turn faults between primary windings due to insulation degradation are a major cause of distribution transformer failure, and occur due to high levels of stress such as overloading and overheating. An additional consequence of these faults is an increased voltage on the transformer secondary due to effective change in turns ratio. This paper develops a novel method for early detection of insulation degradation and subsequent inter-turn winding failure by monitoring the transformer secondary voltage. The algorithm is based on a cumulative sum (CUSUM) statistic and compares voltages on neighbouring transformers to flag degrading assets. Results obtained from simulation as well as experimental data show that smart meter measurements can be utilized to achieve very high detection accuracy while keeping costs low. Here, the paper also demonstrates the validity of the algorithm in the presence of measurement noise, residential solar power injection etc.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Do Machine Learning Approaches Offer Skill Improvement for Short-Term Forecasting of Wind Gust Occurrence and Magnitude?

Abstract Wind gusts, and in particular intense gusts, are societally relevant but extremely challenging to forecast. This study systematically assesses the skill enhancement that can be achieved using artificial neural networks (ANNs) for forecasting of wind gust occurrence and magnitude. Geophysical predictors from the ERA5 reanalysis are used in conjunction with an autoregressive term in regression and ANN models with different predictors, and varying model complexity. Models are derived and assessed for the warm (April–September) and cold (October–March) seasons for three high passenger volume airports in the United States. Model uncertainty is assessed by deriving models for 1000 different randomly selected training (70%) and testing (30%) subsets. Gust prediction fidelity in independent test samples is critically dependent on inclusion of an autoregressive term. Gust occurrence probabilities derived using five-layer ANNs exhibit consistently higher fidelity than those from regression models and shallower ANNs. Inclusion of the autoregressive term and increasing the number of hidden layers in ANNs from 1 to 5 also improve the model performance for gust magnitudes (lower RMSE, increased correlation, and model standard deviations that more closely approximate observed values). Deeper ANNs (e.g., 20 hidden layers) exhibit higher skill in forecasting strong (17–25.7 m s −1 ) and damaging (≥25.7 m s −1 ) wind gusts. However, such deep networks exhibit evidence of overfitting and still substantially underestimate (by 50%) the frequency of strong and damaging wind gusts at the three airports considered herein. Significance Statement Improved short-term forecasting of wind gusts will enhance aviation safety and logistics and may offer other societal benefits. Here we present a rigorous investigation of the relative skill of models of wind gust occurrence and magnitude that employ different statistical methods. It is shown that artificial neural networks (ANNs) offer considerable skill enhancement over regression methods, particularly for strong and damaging wind gusts. For wind gust magnitudes in particular, application of deeper learning networks (e.g., five or more hidden layers) offers tangible improvements in forecast accuracy. However, deeper networks are vulnerable to overfitting and exhibit substantial variability with the specific training and testing data subset used. Also, even deep ANNs reproduce only half of strong and damaging wind gusts. These results indicate the need for future work to elucidate the dynamical mechanisms of intense wind gusts and advance solutions to their prediction.

54 ENVIRONMENTAL SCIENCES↗