Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Boundaries of quantum supremacy via random circuit sampling

Abstract Google’s quantum supremacy experiment heralded a transition point where quantum computers can evaluate a computational task, random circuit sampling, faster than classical supercomputers. We examine the constraints on the region of quantum advantage for quantum circuits with a larger number of qubits and gates than experimentally implemented. At near-term gate fidelities, we demonstrate that quantum supremacy is limited to circuits with a qubit count and circuit depth of a few hundred. Larger circuits encounter two distinct boundaries: a return of a classical advantage and practically infeasible quantum runtimes. Decreasing error rates cause the region of a quantum advantage to grow rapidly. At error rates required for early implementations of the surface code, the largest circuit size within the quantum supremacy regime coincides approximately with the smallest circuit size needed to implement error correction. Thus, the boundaries of quantum supremacy may fortuitously coincide with the advent of scalable, error-corrected quantum computing.

97 MATHEMATICS AND COMPUTING↗

Effect of Nonunital Noise on Random-Circuit Sampling

In this work, drawing inspiration from the type of noise present in real hardware, we study the output distribution of random quantum circuits under practical nonunital noise sources with constant noise rates. We show that even in the presence of unital sources such as the depolarizing channel, the distribution, under the combined noise channel, never resembles a maximally entropic distribution at any depth. To show this, we prove that the output distribution of such circuits never anticoncentrates—meaning that it is never too “flat”—regardless of the depth of the circuit. This is in stark contrast to the behavior of noiseless random quantum circuits or those with only unital noise, both of which anticoncentrate at sufficiently large depths. As a consequence, our results shows that the complexity of random-circuit sampling under realistic noise is still an open question, since anticoncentration is a critical property exploited by both state-of-the-art classical hardness and easiness results. Published by the American Physical Society 2024

Physics↗

Comparison of theory with the experimental characterization of the spatial frequency response of interferometers using a binary pseudo-random array sample

Experimental evaluations of the surface height response of an interference microscope using a binary pseudo-random array test sample are compared with a theory based on a Fourier optics model. Measurements of key instrument characteristics, including the illumination, imaging, and obscuring apertures of three different Mirau objectives, support the theoretical calculations. Agreement between experimental and theoretical modeling confirms the predictability of the spatial frequency response for the purpose of specification and optimization of instrument configuration for specific metrology tasks. The results also provide confidence in methods of compensating for the decrease in instrument response with spatial frequency.

Calibration↗

Investigating the ecological fallacy through sampling distributions constructed from finite populations

Correlation coefficients and linear regression values computed from group averages can differ from correlation coefficients and linear regression values computed using individual scores. This observation known as the ecological fallacy often assumes that all the individual scores are available from a population. In many situations, one must use a sample from the larger population. In such cases, the computed correlation coefficient and linear regression values will depend on the sample that is chosen and the underlying sampling distribution. The sampling distribution of correlation coefficients and linear regression values for group averages will be identical to the sampling distribution for individuals for normally distributed variables for random samples drawn from infinitely large continuous distributions. However, data that is acquired in practice is often acquired when sampling without replacement from a finite population. Our objective is to demonstrate through Monte Carlo simulations that the sampling distributions for correlation and linear regression will also be similar for individuals and group averages when sampling without replacement from normally distributed variables. These simulations suggest that when a random sample from a population is selected, the correlation coefficients and linear regression values computed from individual scores will not be more accurate in estimating the entire population values compared to samples when group averages are used as long as the sample size is the same.

97 MATHEMATICS AND COMPUTING↗

Bootstrapping outperforms community‐weighted approaches for estimating the shapes of phenotypic distributions

Abstract Estimating phenotypic distributions of populations and communities is central to many questions in ecology and evolution. These distributions can be characterized by their moments (mean, variance, skewness and kurtosis) or diversity metrics (e.g. functional richness). Typically, such moments and metrics are calculated using community‐weighted approaches (e.g. abundance‐weighted mean). We propose an alternative bootstrapping approach that allows flexibility in trait sampling and explicit incorporation of intraspecific variation, and show that this approach significantly improves estimation while allowing us to quantify uncertainty. We assess the performance of different approaches for estimating the moments of trait distributions across various sampling scenarios, taxa and datasets by comparing estimates derived from simulated samples with the true values calculated from full datasets. Simulations differ in sampling intensity (individuals per species), sampling biases (abundance, size), trait data source (local vs. global) and estimation method (two types of community‐weighting, two types of bootstrapping). We introduce the traitstrap R package, which contains a modular and extensible set of bootstrapping and weighted‐averaging functions that use community composition and trait data to estimate the moments of community trait distributions with their uncertainty. Importantly, the first function in the workflow, trait_fill , allows the user to specify hierarchical structures (e.g. plot within site, experiment vs. control, species within genus) to assign trait values to each taxon in each community sample. Across all taxa, simulations and metrics, bootstrapping approaches were more accurate and less biased than community‐weighted approaches. With bootstrapping, a sample size of 9 or more measurements per species per trait generally included the true mean within the 95% CI. It reduced average percent errors by 26%–74% relative to community‐weighting. Random sampling across all species outperformed both size‐ and abundance‐biased sampling. Our results suggest randomly sampling ~9 individuals per sampling unit and species, covering all species in the community and analysing the data using nonparametric bootstrapping generally enable reliable inference on trait distributions, including the central moments, of communities. By providing better estimates of community trait distributions, bootstrapping approaches can improve our ability to link traits to both the processes that generate them and their effects on ecosystems.

Maitner, Brian S.↗

Quantifying uncertainty in uranium concentration measurements via K-edge densitometry

This study quantifies the uncertainty in uranium concentration predictions of fluoride and chloride-based salts within a steel pipe using K-edge densitometry. Modeling and simulation was conducted with the Monte Carlo N-Particle Transport (MCNP) code. The quality of of this technique’s prediction in a pipe requires proper characterization of the pipe’s thickness, which is dependent on the source size and axial offset from the pipe centerline. The thickness was determined as either the center-line thickness seen by the X-ray source or an average value determined through random sampling. Generally, the predicted concentrations were slightly better at lower offset with the random sampling thickness and using the center-line thickness for the highest offsets. For a line-beam source and varying axial offsets, the relative error of concentration was within 1% of the true value but uncertainty increased by 2 orders of magnitude. Similarly, for no axial offset, the relative error was significantly less than 1% while no trend for uncertainty was found. However, at the largest possible offset for a given source size, the concentrations become erroneous and greater than the allowable 1% relative error. Furthermore, high offsets tended to increase the variance of the transmission spectra by 3 orders of magnitude.

Characterization and Analytical Technique↗

Iterative self-organizing SCEne-LEvel sampling (ISOSCELES) for large-scale building extraction

Convolutional neural networks (CNN) provide state-of-the-art performance in many computer vision tasks, including those related to remote-sensing image analysis. Successfully training a CNN to generalize well to unseen data, however, requires training on samples that represent the full distribution of variation of both the target classes and their surrounding contexts. With remote sensing data, acquiring a sufficiently representative training set is a challenge due to both the inherent multi-modal variability of satellite or aerial imagery and the general high cost of labeling data. To address this challenge, we have developed ISOSCELES, an Iterative Self-Organizing SCEne LEvel Sampling method for hierarchical sampling of large image sets. Using affinity propagation, ISOSCELES automates the selection of highly representative training images. Compared to random sampling or using available reference data, the distribution of the training is principally data driven, reducing the chance of oversampling uninformative areas or undersampling informative ones. In comparison to manual sample selection by an analyst, ISOSCELES exploits descriptive features, spectral and/or textural, and eliminates human bias in sample selection. Using a hierarchical sampling approach, ISOSCELES can obtain a training set that reflects both between-scene variability, such as in viewing angle and time of day, and within-scene variability at the level of individual training samples. We verify the method by demonstrating its superiority to stratified random sampling in the challenging task of adapting a pre-trained model to a new image and spatial domain for country-scale building extraction. Using a pair of hand-labeled training sets comprising 1,987 sample image chips, a total of 496,000,000 individually labeled pixels, we show, across three distinct model architectures, an increase in accuracy, as measured by F1-score, of 2.2–4.2%.

42 ENGINEERING↗

Active Learning‐Driven Inkless Additive Nanomanufacturing for Printed Electronics

Inkless additive nanomanufacturing for printed electronics promises broad material and substrate versatility, yet the high-dimensional print parameter space makes tuning print parameters time-intensive. We present a Bayesian optimization study that constructs a digital twin from printed-silver data to benchmark surrogate models, acquisition functions, and batch sizes head-to-head to achieve user-specified target resistance. Tested surrogate models included Gaussian process, random forest, and Bayesian neural network surrogates with expected improvement and confidence bound acquisition functions. In total, we evaluate 48 unique model configurations alongside a random sampling baseline for comparison. For printed silver, the Bayesian neural network with a batch size of one achieved the lowest average cumulative regret, approximately four times more efficient on average than random sampling. To balance performance and substrate space, a random forest model with expected improvement and a batch size of four was chosen as the model for validation testing. Applying this chosen configuration to copper with an additional print parameter, the model achieved a resistance within 0.15 Ω of a 1 Ω target in fewer than 30 printed lines across five validation sets. Altogether, the workflow yields a tuned and validated model that efficiently guides experiments toward the target while simultaneously learning the parameter space.

Bevel, Colton [Auburn University, AL (United State↗

Estimating the Adequacy of a Multi-Objective Optimization

Multi-objective optimization methods can be criticized for lacking a statistically valid measure of the quality and representativeness of a solution. This stance is especially relevant to metaheuristic optimization approaches but can also apply to other methods that typically might only report a small representative subset of a Pareto frontier. Here we present a method to address this deficiency based on random sampling of a solution space to determine, with a specified level of confidence, the fraction of the solution space that is surpassed by an optimization. The Superiority of Multi-Objective Optimization to Random Sampling, or SMORS method, can evaluate quality and representativeness using dominance or other measures, e.g., a spacing measure for high-dimensional spaces. SMORS has been tested in a combinatorial optimization context using a genetic algorithm but could be useful for other optimization methods.

42 ENGINEERING↗

Statistical analysis on random quantum circuit sampling by Sycamore and Zuchongzhi quantum processors

Random quantum circuit sampling, a task to sample bit strings from a random quantum circuit, is considered a suitable benchmark task to demonstrate the outperformance of quantum computers even with noisy qubits. Recently, random quantum circuit sampling was performed on the Sycamore quantum processor with 53 qubits [Nature (London) 574, 505 (2019)] and on the Zuchongzhi quantum processor with 56 qubits [Phys. Rev. Lett. 127, 180501 (2021)]. Here, we analyze and compare the statistical properties of the outputs of the random quantum circuit sampling by the Sycamore and Zuchongzhi processors. Using the Marchenko-Pastur law of random matrices of bit strings and the Wasssertein distances between bit strings, we find that the statistical properties of Sycamore bit strings are quite different from those of Zuchongzhi bit strings, while both processors score similar values of linear cross-entropy fidelity for random circuit sampling. Some bit strings sampled by the Zuchongzhi processor pass the NIST random number tests while both Sycamore and Zuchongzhi processors show similar patterns in the heat maps of bit strings. Zuchongzhi bit strings are much closer to classical uniform random bits than those of Sycamore. It is shown that the statistical properties of bit strings of both random quantum circuits change little as the depth of the random quantum circuits increases. Our findings raise a question about the computational reliability of noisy quantum processors because two quantum processors with similar noise levels and similar qubit structures produced statistically different outputs for the same random quantum circuit sampling.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Entropy-driven Optimal Sub-sampling of Fluid Dynamics for Developing Machine-learned Surrogates

Optimal sub-sampling of large datasets from fluid dynamics simulations is essential for training reduced-order machine learned models. A method using Shannon entropy was developed to weight flow features according to their level of information content, such that the most informative features can be extracted and used for training a surrogate model. The method is demonstrated in the canonical flow over a cylinder problem simulated with OpenFOAM. Both time-independent predictions and temporal forecasting were investigated as well as two types of prediction targets: local per-grid-point predictions and global per-time-step predictions. When tested on training a surrogate model, results indicate that our entropy-based sampling method typically outperforms random sampling and yields more reproducible results in less iterations. Finally, the method was used to train a surrogate model for modeling turbulence in magnetohydrodynamic flows, which revealed various challenges and opportunities for future research.

Brewer, Wes↗

Artificial Intelligence and Machine Learning Support for Probabilistic Fracture Mechanics

In this research, artificial intelligence and machine learning (ML) methods are used to search an uncertain parameter space more efficiently for the most important inputs with respect to response sensitivities. These methods are applied to the Extremely Low Probability of Rupture (xLPR) probabilistic fracture mechanics code used at the U.S. Nuclear Regulatory Commission (NRC) in support of nuclear regulatory research. This report documents two separate but related sub-tasks: (1) ranking important uncertain input features with respect to target outputs, determined by convergence in confidence intervals for increasing sample sizes using simple random sampling; and (2) implementation of a reduced-order surrogate model for fast, approximate sample generation. Unoptimized readily available off-the-shelf ML models were used in both sub-tasks.

97 MATHEMATICS AND COMPUTING↗

A globally sampled high-resolution hand-labeled validation dataset for evaluating surface water extent maps

Effective monitoring of global water resources is increasingly critical due to climate change and population growth. Advancements in remote sensing technology, specifically in spatial, spectral, and temporal resolutions, are revolutionizing water resource monitoring, leading to more frequent and high-quality surface water extent maps using various techniques such as traditional image processing and machine learning algorithms. However, satellite imagery datasets contain trade-offs that result in inconsistencies in performance, such as disparities in measurement principles between optical (e.g., Sentinel-2) and radar (e.g., Sentinel-1) sensors and differences in spatial and spectral resolutions among optical sensors. Therefore, developing accurate and robust surface water mapping solutions requires independent validations from multiple datasets to identify potential biases within the imagery and algorithms. However, high-quality validation datasets are expensive to build, and few contain information on water resources. For this purpose, we introduce a globally sampled, high-spatial-resolution dataset labeled using 3 m PlanetScope imagery. Our surface water extent dataset comprises 100 images, each with a size of 1024×1024 pixels, which were sampled using a stratified random sampling strategy covering all 14 biomes. We highlighted urban and rural regions, lakes, and rivers, including braided rivers and coastal regions. We evaluated two surface water extent mapping methods using our dataset – Dynamic World, based on Sentinel-2, and the NASA IMPACT model, based on Sentinel-1. Dynamic World achieved a mean intersection over union (IoU) of 72.16 % and F1 score of 79.70 %, while the NASA IMPACT model had a mean IoU of 57.61 % and F1 score of 65.79 %. Performance varied substantially across biomes, highlighting the importance of evaluating models on diverse landscapes to assess their generalizability and robustness. Our dataset can be used to analyze satellite products and methods, providing insights into their advantages and drawbacks. Our dataset offers a unique tool for analyzing satellite products, aiding the development of more accurate and robust surface water monitoring solutions. The dataset can be accessed via https://doi.org/10.25739/03nt-4f29.

54 ENVIRONMENTAL SCIENCES↗

Smart Scattering Scanning Near-Field Optical Microscopy

Scattering scanning near-field optical microscopy (s-SNOM) provides spectroscopic imaging from molecular to quantum materials with few nanometer deep subdiffraction limited spatial resolution. However, in its conventional implementation s-SNOM is slow to effectively acquire a series of spatio-spectral images, especially with large fields of view. This problem is further exacerbated for weak resonance contrast or when using light sources with limited spectral irradiance. Indeed, the generally limited signal-to-noise ratio prevents sampling a weak signal at the Nyquist sampling rate. Here, we demonstrate how acquisition time and sampling rate can be significantly reduced by using compressed sampling, matrix completion, and adaptive random sampling, while maintaining or even enhancing the physical or chemical image content. We use fully sampled real data sets of molecular, biological, and quantum materials as ground-truth physical data and show how deep under-sampling with a corresponding reduction of acquisition time by 1 order of magnitude or more retains the core s-SNOM image information. We demonstrate that a sampling rate of up to 6× smaller than the Nyquist criterion can be applied, which would provide a 30-fold reduction in the data required under typical experimental conditions. Furthermore, our smart s-SNOM approach is generally applicable and provides systematic full spatio-spectral s-SNOM imaging with a large field of view at high spectral resolution and reduced acquisition time.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

How to estimate soil organic carbon stocks of agricultural fields? perspectives using ex-ante evaluation

Estimating soil organic carbon (SOC) stocks of agricultural fields has a range of important applications from development of sustainable management practices to monitoring carbon stocks. There are many estimation strategies with the potential for more reliable estimates of SOC stock and more efficient use of soil sampling and analysis resources, especially by leveraging readily available auxiliary information such as remote sensing. However, concrete guidance for strategy selection is lacking. This study narrows this gap with a comparison of strategies for estimating deep SOC stock (0–60 cm) in a prototypical field. Using high density SOC stock measurements and simulation, we built on past studies by 1) ex-ante evaluating a large number of strategy options, 2) using a Bayesian approach to quantify the uncertainty of the comparison, and 3) considering multiple Bayesian models to assess sensitivity to this modeling choice. We found that, using readily available auxiliary information, both balanced and stratified sampling offer substantial improvements over simple random sampling. The auxiliary information most important for this improvement is a Sentinel-2 SOC index = blue / (green × red), followed by the topographic wetness index. We found that these results are robust to the choice of mapping method, but that there is uncertainty in the magnitude of improvement. Here, we recommend future studies implement this Bayesian approach for simulated ex-ante evaluation of SOC stock estimation strategies across more fields to investigate the generalizability of these findings.

54 ENVIRONMENTAL SCIENCES↗

Deterministic Linear Time for Maximal Poisson‐Disk Sampling using Chocks without Rejection or Approximation

Abstract We show how to sample uniformly within the three‐sided region bounded by a circle, a radial ray, and a tangent, called a “chock.” By dividing a 2D planar rectangle into a background grid, and subtracting Poisson disks from grid squares, we are able to represent the available region for samples exactly using triangles and chocks. Uniform random samples are generated from chock areas precisely without rejection sampling. This provides the first implemented algorithm for precise maximal Poisson‐disk sampling in deterministic linear time. We prove O(n · M(b) log b), where n is the number of samples, b is the bits of numerical precision and M is the cost of multiplication. Prior methods have higher time complexity, take expected time, are non‐maximal, and/or are not Poisson‐disk distributions in the most precise mathematical sense. We fill this theoretical lacuna.

Mitchell, Scott A.↗

Extension of SCALE/Sampler’s sensitivity analysis

Nuclear data are a major source of uncertainties in reactor physics calculations. The propagation of nuclear data uncertainties to important system responses is instrumental when determining appropriate safety margins in reactor safety analyses. It is also important to understand the major contributors to the observed uncertainties to make recommendations for further measurements and evaluations and aid in the understanding of the studied system. The SCALE code system allows for nuclear data uncertainty analysis based on the random sampling approach as implemented in SCALE’s Sampler sequence. Sampler was recently extended by a sensitivity analysis in terms of the calculation of two correlation-based sensitivity indices. This analysis allows for the identification of the top contributing nuclear reactions to any analyzed output uncertainty. This paper presents the sensitivity indices, along with their interpretation and limitations. It demonstrates the application in an eigenvalue and decay heat analysis for a boiling water reactor fuel assembly.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

2015 Madison County, Indiana, In the Moment Travel Study

The 2015 In the Moment Travel Study—a pilot study—captured the travel behavior and characteristics of residents in Madison County, Indiana. The Madison County Council of Governments sponsored the study, which was administered by Resource Systems Group and conducted from February to March 2015. It used an activity sampling or "random moments" sampling approach via a smartphone application to capture travel behavior and characteristics from the survey participants. This approach included brief smartphone interactions, e.g., a few minutes per interaction, conducted multiple times a day over multiple days, which was considered less burdensome than traditional household travel diary surveys, which often require 20-30 minutes in one sitting. This proof-of-concept study included households that also participated in the 2014 Heartland in Motion household travel diary survey. Because of this, an assessment of the accuracy and completeness of the collected smartphone application data and comparisons between the "random moments" sample method and traditional household travel surveys can conceivably be drawn.

1Hz data↗