Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical sampling techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Preliminary design of an Earth-based debris detection system using current technology and existing installations

Assessment of debris hazard requires the determination of debris down to mm sizes for near-Earth orbits and near-stationary points. It is necessary to obtain reasonable orbits for a statistically significant sample of the debris population. Several ground-based techniques for detection are available. Radar detection was used to obtain information of existing debris population. Another technique is optical detection. The possibilities and application of optical detection with state-of-the-art instrumentation is studied.

Morgan, T. H.↗

Shuttle Hypervelocity Impact Database

With three missions outstanding, the Shuttle Hypervelocity Impact Database has nearly 3000 entries. The data is divided into tables for crew module windows, payload bay door radiators and thermal protection system regions, with window impacts compromising just over half the records. In general, the database provides dimensions of hypervelocity impact damage, a component level location (i.e., window number or radiator panel number) and the orbiter mission when the impact occurred. Additional detail on the type of particle that produced the damage site is provided when sampling data and definitive analysis results are available. Details and insights on the contents of the database including examples of descriptive statistics will be provided. Post flight impact damage inspection and sampling techniques that were employed during the different observation campaigns will also be discussed. Potential enhancements to the database structure and availability of the data for other researchers will be addressed in the Future Work section. A related database of returned surfaces from the International Space Station will also be introduced.

Hyde, James L.↗

Multiple-Beam Detection of Fast Transient Radio Sources

A method has been designed for using multiple independent stations to discriminate fast transient radio sources from local anomalies, such as antenna noise or radio frequency interference (RFI). This can improve the sensitivity of incoherent detection for geographically separated stations such as the very long baseline array (VLBA), the future square kilometer array (SKA), or any other coincident observations by multiple separated receivers. The transients are short, broadband pulses of radio energy, often just a few milliseconds long, emitted by a variety of exotic astronomical phenomena. They generally represent rare, high-energy events making them of great scientific value. For RFI-robust adaptive detection of transients, using multiple stations, a family of algorithms has been developed. The technique exploits the fact that the separated stations constitute statistically independent samples of the target. This can be used to adaptively ignore RFI events for superior sensitivity. If the antenna signals are independent and identically distributed (IID), then RFI events are simply outlier data points that can be removed through robust estimation such as a trimmed or Winsorized estimator. The alternative "trimmed" estimator is considered, which excises the strongest n signals from the list of short-beamed intensities. Because local RFI is independent at each antenna, this interference is unlikely to occur at many antennas on the same step. Trimming the strongest signals provides robustness to RFI that can theoretically outperform even the detection performance of the same number of antennas at a single site. This algorithm requires sorting the signals at each time step and dispersion measure, an operation that is computationally tractable for existing array sizes. An alternative uses the various stations to form an ensemble estimate of the conditional density function (CDF) evaluated at each time step. Both methods outperform standard detection strategies on a test sequence of VLBA data, and both are efficient enough for deployment in real-time, online transient detection applications.

Thompson, David R.↗

IAEA Facility-Level Safeguards and Implementation and Advanced Verification Technologies

IAEA Facility-Level Safeguards and Implementation - This presentation is intended to introduce the audience to the application of international safeguards by the International Atomic Energy Agency at nuclear facilities around the world. It covers the State Level Approach, random statistical sampling processes, material balance area structure, non-destructive assay techniques, and inspection areas. Addition of advanced verification technologies for the IAEA

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

A Provably Accurate Randomized Sampling Algorithm for Logistic Regression

In statistics and machine learning, logistic regression is a widely-used supervised learning technique primarily employed for binary classification tasks. When the number of observations greatly exceeds the number of predictor variables, we present a simple, randomized sampling-based algorithm for logistic regression problem that guarantees high-quality approximations to both the estimated probabilities and the overall discrepancy of the model. Our analysis builds upon two simple structural conditions that boil down to randomized matrix multiplication, a fundamental and well-understood primitive of randomized numerical linear algebra. We analyze the properties of estimated probabilities of logistic regression when leverage scores are used to sample observations, and prove that accurate approximations can be achieved with a sample whose size is much smaller than the total number of observations. To further validate our theoretical findings, we conduct comprehensive empirical evaluations. Overall, our work sheds light on the potential of using randomized sampling approaches to efficiently approximate the estimated probabilities in logistic regression, offering a practical and computationally efficient solution for large-scale datasets.

Chowdhury, Agniva↗

Graph Neural Networks for Parameterized Quantum Circuits Expressibility Estimation (Rev.1)

Parameterized quantum circuits (PQCs) are fundamental to quantum machine learning (QML), quantum optimization, and variational quantum algorithms (VQAs). The expressibility of PQCs is a measure that determines their capability to harness the full potential of the quantum state space. It is thus a crucial guidepost to know when selecting a particular PQC ansatz. However, the existing technique for expressibility computation through statistical estimation requires a large number of samples, which poses significant challenges due to time and computational resource constraints. This paper introduces a novel approach for expressibility estimation of PQCs using Graph Neural Networks (GNNs). We demonstrate the predictive power of our GNN model with a dataset consisting of 25,000 samples from the noiseless IBM QASM Simulator and 12,000 samples from three distinct noisy quantum backends. The model accurately estimates expressibility, with root mean square errors (RMSE) of 0.05 and 0.06 for the noiseless and noisy backends, respectively. We compare our model’s predictions with reference circuits from Sim et al. and IBM Qiskit’s hardwareefficient ansatz sets to further evaluate our model’s performance. Our experimental evaluation in noiseless and noisy scenarios reveals a close alignment with ground truth expressibility values, highlighting the model’s efficacy. Moreover, our model exhibits promising extrapolation capabilities, predicting expressibility values with low RMSE for out-of-range qubit circuits trained solely on only up to 5-qubit circuit sets. This work thus provides a reliable means of efficiently evaluating the expressibility of diverse PQCs on noiseless simulators and hardware.

97 MATHEMATICS AND COMPUTING↗

Characterizing the Uncertainty of Measurement of Traceable Isotope Ratios with Bayesian Statistical Techniques

Analytical techniques such as multicollector—inductively coupled plasma—mass spectrometry (MC-ICP-MS) are routinely employed at SRNL, other National Laboratories, and in academia to determine the precise isotopic composition of diverse natural and anthropogenic samples (e.g., rocks and nuclear materials). Quantifying and reporting uncertainty in such analyses, while regularly performed, have a rigorous statistical foundation. The Guide to the Expression of Uncertainty in Measurement 4 (GUM) outlines conventional techniques used to assess such uncertainty. As the accessibility and speed of statistical computing increase, there is a need to modernize conventional techniques. For example, Supplement 1 to the 3rd to the GUM suggests the use of approximation methods as an updated approach to the GUM.

McLarty, Ellis C.↗

The Dark Energy Survey supernova program: cosmological biases from supernova photometric classification

ABSTRACT Cosmological analyses of samples of photometrically identified type Ia supernovae (SNe Ia) depend on understanding the effects of ‘contamination’ from core-collapse and peculiar SN Ia events. We employ a rigorous analysis using the photometric classifier SuperNNova on state-of-the-art simulations of SN samples to determine cosmological biases due to such ‘non-Ia’ contamination in the Dark Energy Survey (DES) 5-yr SN sample. Depending on the non-Ia SN models used in the SuperNNova training and testing samples, contamination ranges from 0.8 to 3.5 per cent, with a classification efficiency of 97.7–99.5 per cent. Using the Bayesian Estimation Applied to Multiple Species (BEAMS) framework and its extension BBC (‘BEAMS with Bias Correction’), we produce a redshift-binned Hubble diagram marginalized over contamination and corrected for selection effects, and use it to constrain the dark energy equation-of-state, w. Assuming a flat universe with Gaussian ΩM prior of 0.311 ± 0.010, we show that biases on w are <0.008 when using SuperNNova, with systematic uncertainties associated with contamination around 10 per cent of the statistical uncertainty on w for the DES-SN sample. An alternative approach of discarding contaminants using outlier rejection techniques (e.g. Chauvenet’s criterion) in place of SuperNNova leads to biases on w that are larger but still modest (0.015–0.03). Finally, we measure biases due to contamination on w0 and wa (assuming a flat universe), and find these to be <0.009 in w0 and <0.108 in wa, 5 to 10 times smaller than the statistical uncertainties for the DES-SN sample.

79 ASTRONOMY AND ASTROPHYSICS↗

Reduction of flow-measurement uncertainties in laser velocimeters with nonorthogonal channels

An analysis of certain geometrical limitations inherent in the application of laser velocimeters with nonorthogonal channels has led to the development of advanced-LDA-calibration and data-acquisition techniques that minimize systematic and statistical errors, respectively. The data-acquisition technique optimizes the number of velocity samples collected from three velocimeter channels as a function of local turbulence intensity, vector direction, and prescribed confidence interval. Linear velocity surveys and streamline traces measured in a turbulent flow field with a three-dimensional laser velocimeter are presented and the validity and accuracy of the theoretical analysis are discussed.

Snyder, P. K.↗

A deep x-ray survey of the Pleiades cluster and the B6-A3 main sequence stars in Orion

We have obtained deep ROSAT images of three regions within the Pleiades open cluster. We have detected 317 X-ray sources in these ROSAT PSPC images, 171 of which we associate with certain probable members of the Pleiades cluster. We detect nearly all Pleiades members with spectral types later than G0 and within 25 arcminutes of our three field centers where our sensitivity is highest. This has allowed us to derive for the first time the luminosity function for the G, K, and M dwarfs of an open cluster without the need to use statistical techniques to account for the presence of upper limits in the data sample. Because of our high X-ray detection frequency down to the faint limit of the optical catalog, we suspect that some of our unidentified X-ray sources are previously unknown, very low-mass members of the Pleiades. A large fraction of the Pleiades members detected with ROSAT have published rotational velocities. Plots of L(sub x)/L(sub bol) versus spectroscopic rotational velocity show tightly correlated 'saturation' type relations for stars with (B - V)(sub O) greater than 0.60. For each of several color ranges, X-ray luminosities rise rapidly with increasing rotation rate until v sin i approximately equals 15 km/s, and then remain essentially flat for rotation rates up to at least v sin i approximately equal to 100 km/s. The dispersion in rotation among low-mass stars in the Pleiades is by far the dominant contributor to the dispersion in L(subx) at a given mass. Only about 35 percent of the B.A. and early F stars in the Pleiades are detected as X-ray sources in our survey. There is no correlation between X-ray flux and rotation for these stars. The X-ray luminosity function for the early-type Pleiades stars appears to be bimodal, with only a few exceptions. We either detect these stars at fluxes in the range found for low-mass stars or we derive X-ray limits below the level found for most Pleiades dwarfs. The X-ray spectra for the early-type Pleiades stars detected by ROSAT are indistinguishable from the spectra of the low-mass Pleiades members. We believe that the simple explanation for this behavior is that the early-type Pleiades stars are not themselves intrinsic X-ray sources and that the X-ray sources and that the X-ray emission actually arises from low-mass companions to these stars.

Caillault, Jean-Pierre↗

A deep imaging survey of the Pleiades with ROSAT

We have obtained deep ROSAT images of three regions within the Pleiades open cluster. We have detected 317 X-ray sources in these ROSAT Position Sensitive Proportional Counter (PSPC) images, 171 of which we associate with certain or probable members of the Pleiades cluster. We detect nearly all Pleiades members with spectral types later than G0 and within 25 arcminutes of our three field centers where our sensitivity is highest. This has allowed us to derive for the first time the luminosity function for the G, K, amd M dwarfs of an open cluster without the need to use statistical techniques to account for the presence of upper limits in the data sample. Because of our high X-ray detection frequency down to the faint limit of the optical catalog, we suspect that some of our unidentified X-ray sources are previously unknown, very low-mass members of Pleiades. A large fraction of the Pleiades members detected with ROSAT have published rotational velocities. Plots of L(sub X)/L(sub Bol) versus spectroscopic rotational velocity show tightly correlated `saturation' type relations for stars with ((B - V)(sub 0)) greater than or equal to 0.60. For each of several color ranges, X-ray luminosities rise rapidly with increasing rotation rate until c sin i approximately equal to 15 km/sec, and then remains essentially flat for rotation rates up to at least v sin i approximately equal to 100 km/sec. The dispersion in rotation among low-mass stars in the Pleiades is by far the dominant contributor to the dispersion in L(sub X) at a given mass. Only about 35% of the B, A, and early F stars in the Pleiades are detected as X-ray sources in our survey. There is no correlation between X-ray flux and rotation for these stars. The X-ray luminosity function for the early-type Pleiades stars appears to be bimodal -- with only a few exceptions, we either detect these stars at fluxes in the range found for low-mass stars or we derive X-ray limits below the level found for most Pleiades dwarfs. The X-ray spectra for the early-type Pleiades stars detected by ROSAT are indistinguishable from the spectra of the low-mass Pleiades members. We believe that the simplest explanation for this behavior is that the early-type Pleiades stars are not themselves intrinsic X-ray sources and that the X-ray emission actually arises from low-mass companions to these stars.

Stauffer, J. R.↗

Analyzing and Predicting Effort Associated with Finding and Fixing Software Faults

Context: Software developers spend a significant amount of time fixing faults. However, not many papers have addressed the actual effort needed to fix software faults. Objective: The objective of this paper is twofold: (1) analysis of the effort needed to fix software faults and how it was affected by several factors and (2) prediction of the level of fix implementation effort based on the information provided in software change requests. Method: The work is based on data related to 1200 failures, extracted from the change tracking system of a large NASA mission. The analysis includes descriptive and inferential statistics. Predictions are made using three supervised machine learning algorithms and three sampling techniques aimed at addressing the imbalanced data problem. Results: Our results show that (1) 83% of the total fix implementation effort was associated with only 20% of failures. (2) Both safety critical failures and post-release failures required three times more effort to fix compared to non-critical and pre-release counterparts, respectively. (3) Failures with fixes spread across multiple components or across multiple types of software artifacts required more effort. The spread across artifacts was more costly than spread across components. (4) Surprisingly, some types of faults associated with later life-cycle activities did not require significant effort. (5) The level of fix implementation effort was predicted with 73% overall accuracy using the original, imbalanced data. Using oversampling techniques improved the overall accuracy up to 77%. More importantly, oversampling significantly improved the prediction of the high level effort, from 31% to around 85%. Conclusions: This paper shows the importance of tying software failures to changes made to fix all associated faults, in one or more software components and/or in one or more software artifacts, and the benefit of studying how the spread of faults and other factors affect the fix implementation effort.

software fix implementation effort↗

A statistical technique for determining rainfall over land employing Nimbus-6 ESMR measurements

Statistical analysis is performed by first sampling three categories of Nimbus 6 ESMR brightness temperatures (representing rain over land, wet land surfaces without rain, and dry land surfaces), then testing these populations for uniqueness. A classification algorithm to delineate rain over land is developed. It is found that synoptic-scale rainfall over land, where surface thermodynamic temperatures are greater than 5 C and the vegetation is bereft of dew, can indeed be delineated despite the large ESMR-6 instantaneous field of view. However, some ambiguity exists in distinguishing between rainfall areas and wet land surfaces.

Rodgers, E.↗

Uncertainty Quantification for Polynomial Systems via Bernstein Expansions

This paper presents a unifying framework to uncertainty quantification for systems having polynomial response metrics that depend on both aleatory and epistemic uncertainties. The approach proposed, which is based on the Bernstein expansions of polynomials, enables bounding the range of moments and failure probabilities of response metrics as well as finding supersets of the extreme epistemic realizations where the limits of such ranges occur. These bounds and supersets, whose analytical structure renders them free of approximation error, can be made arbitrarily tight with additional computational effort. Furthermore, this framework enables determining the importance of particular uncertain parameters according to the extent to which they affect the first two moments of response metrics and failure probabilities. This analysis enables determining the parameters that should be considered uncertain as well as those that can be assumed to be constants without incurring significant error. The analytical nature of the approach eliminates the numerical error that characterizes the sampling-based techniques commonly used to propagate aleatory uncertainties as well as the possibility of under predicting the range of the statistic of interest that may result from searching for the best- and worstcase epistemic values via nonlinear optimization or sampling.

Crespo, Luis G.↗

Deep inference of simulated strong lenses in ground-based surveys

The large number of strong lenses discoverable in future astronomical surveys will likely enhance the value of strong gravitational lensing as a cosmic probe of dark energy and dark matter. However, leveraging the increased statistical power of such large samples will require further development of automated lens modeling techniques. We show that deep learning and simulation-based inference (SBI) methods produce informative and reliable estimates of parameter posteriors for strong lensing systems in ground-based surveys. We present the examination and comparison of two approaches to lens parameter estimation for strong galaxy-galaxy lenses — Neural Posterior Estimation (NPE) and Bayesian Neural Networks (BNNs). We perform inference on 1-, 5-, and 12-parameter lens models for ground-based imaging data that mimics the Dark Energy Survey (DES). We find that NPE outperforms BNNs, producing posterior distributions that are more accurate, precise, and well-calibrated for most parameters. For the 12-parameter NPE model, the calibration is consistently within <10% of optimal calibration for all parameters, while the BNN is rarely within 20% of optimal calibration for any of the parameters. Similarly, residuals for most of the parameters are smaller (by up to an order of magnitude) with the NPE model than the BNN model. This work takes important steps in the systematic comparison of methods for different levels of model complexity.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The Surface-Topography Challenge: A Multi-Laboratory Benchmark Study to Advance the Characterization of Topography

Surface performance is critically influenced by topography in virtually all real-world applications. The current standard practice is to describe topography using one of a few industry-standard parameters. The most commonly reported number is Ra, the average absolute deviation of the height from the mean line (at some, not necessarily known or specified, lateral length scale). However, other parameters, particularly those that are scale-dependent, influence surface and interfacial properties; for example the local surface slope is critical for visual appearance, friction, and wear. The present Surface-Topography Challenge was launched to raise awareness for the need of a multi-scale description, but also to assess the reliability of different metrology techniques. In the resulting international collaborative effort, 153 scientists and engineers from 64 research groups and companies across 20 countries characterized statistically equivalent samples from two different surfaces: a “rough” and a “smooth” surface. The results of the 2088 measurements constitute the most comprehensive surface description ever compiled. We find wide disagreement across measurements and techniques when the lateral scale of the measurement is ignored. Consensus is established through scale-dependent parameters while removing data that violates an established resolution criterion and deviates from the majority measurements at each length scale. Our findings suggest best practices for characterizing and specifying topography. The public release of the accumulated data and presented analyses enables global reuse for further scientific investigation and benchmarking.

42 ENGINEERING↗

Multiclass Bayes error estimation by a feature space sampling technique

A general Gaussian M-class N-feature classification problem is defined. An algorithm is developed that requires the class statistics as its only input and computes the minimum probability of error through use of a combined analytical and numerical integration over a sequence simplifying transformations of the feature space. The results are compared with those obtained by conventional techniques applied to a 2-class 4-feature discrimination problem with results previously reported and 4-class 4-feature multispectral scanner Landsat data classified by training and testing of the available data.

Mobasseri, B. G.↗

Operations for Learning with Graphical Models

This paper is a multidisciplinary review of empirical, statistical learning from a graphical model perspective. Well-known examples of graphical models include Bayesian net- works, directed graphs representing a Markov chain, and undirected networks representing a Markov field. These graphical models are extended to model data analysis and empirical learning using the notation of plates. Graphical operations for simplifying and manipulating a problem are provided including decomposition, differentiation, and the manipulation of probability models from the exponential family. These operations adapt existing techniques from statistics and automatic differentiation to graphs. Two standard algorithm schemes for learning are reviewed in a graphical framework: Gibbs sampling and the expectation maximization algorithm. Some algorithms are developed in this graphical framework including a generalized version of linear regression, techniques for feed-forward networks, and learning Gaussian and discrete Bayesian networks from data. The paper concludes by sketching some implications for data analysis and summarizing some popular algorithms that fall within the framework presented. The main original contributions here are the decomposition techniques and the demonstration that graphical models provide a framework for understanding and developing complex learning algorithms.

Buntine, Wray L.↗