Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Random projection using random quantum circuits

The random sampling task performed by Google's Sycamore processor gave us a glimpse of the “quantum supremacy era.” This has definitely shed some light on the power of random quantum circuits in this abstract task of sampling outputs from the (pseudo)random circuits. In this paper, we explore a practical near-term use of local random quantum circuits in dimensional reduction of large low-rank data sets. We make use of the well-studied dimensionality reduction technique called the random projection method. This method has been extensively used in various applications such as image processing, logistic regression, entropy computation of low-rank matrices, etc. We prove that the matrix representations of local random quantum circuits with sufficiently shorter depths [ ∼ O ( n ) ] serve as good candidates for random projection. We demonstrate numerically that their projection abilities are not far off from the computationally expensive classical principal components analysis on MNIST and CIFAR-100 image datasets. We also benchmark the performance of quantum random projection against the commonly used classical random projection in the tasks of dimensionality reduction of image data sets and computing von Neumann entropies of large low-rank density matrices. And finally, using variational quantum singular value decomposition, we demonstrate a near-term implementation of extracting the singular vectors with dominant singular values after quantum random projecting a large low-rank matrix to lower dimensions. All such numerical experiments unequivocally demonstrate the ability of local random circuits to randomize a large Hilbert space at sufficiently shorter depths with robust retention of properties of large data sets in reduced dimensions. Published by the American Physical Society 2024

Kumaran, Keerthi (ORCID:0009000949125721)↗

A comprehensive study of non-adaptive and residual-based adaptive sampling for physics-informed neural networks

Physics-informed neural networks (PINNs) have shown to be effective tools for solving both forward and inverse problems of partial differential equations (PDEs). PINNs embed the PDEs into the loss of the neural network using automatic differentiation, and this PDE loss is evaluated at a set of scattered spatio-temporal points (called residual points). The location and distribution of these residual points are highly important to the performance of PINNs. However, in the existing studies on PINNs, only a few simple residual point sampling methods have mainly been used. Here, we present a comprehensive study of two categories of sampling for PINNs: non-adaptive uniform sampling and adaptive nonuniform sampling. We consider six uniform sampling methods, including (1) equispaced uniform grid, (2) uniformly random sampling, (3) Latin hypercube sampling, (4) Halton sequence, (5) Hammersley sequence, and (6) Sobol sequence. We also consider a resampling strategy for uniform sampling. To improve the sampling efficiency and the accuracy of PINNs, we propose two new residual-based adaptive sampling methods: residual-based adaptive distribution (RAD) and residual-based adaptive refinement with distribution (RAR-D), which dynamically improve the distribution of residual points based on the PDE residuals during training. Hence, we have considered a total of 10 different sampling methods, including six non-adaptive uniform sampling, uniform sampling with resampling, two proposed adaptive sampling, and an existing adaptive sampling. We extensively tested the performance of these sampling methods for four forward problems and two inverse problems in many setups. Our numerical results presented in this study are summarized from more than 6000 simulations of PINNs. Here, we show that the proposed adaptive sampling methods of RAD and RAR-D significantly improve the accuracy of PINNs with fewer residual points for both forward and inverse problems. Furthermore, the results obtained in this study can also be used as a practical guideline in choosing sampling methods.

97 MATHEMATICS AND COMPUTING↗

Say where you sample: Increasing site selection transparency in urban ecology

Urban ecological studies have the potential to extend our understanding of socio-ecological systems beyond that of an individual city or region. Cross-comparative empirical work and synthesis are imperative to develop a general urban ecological theory. This can be achieved only if studies are replicable and generalizable. Transparency in methods reporting facilitates generalizability and replicability by documenting the decisions scientists make during the various steps of research design; this is particularly true for sampling design and selection because of their impact on both internal and external validity and the potential to unintentionally introduce bias. Three interdependent aspects of sample design are study sample selection (e.g., specific organisms, soils, or water), sample specification (measurement of specific variable of interest), and site selection (locations sampled). Of these, documentation of site selection—the where component of sample design—is underrepresented in the urban ecology literature. Using a stratified random sample of 158 papers from 12 major urban ecology journals, we investigated how researchers selected study sites in urban ecosystems and evaluated whether their site selection methods were transparent. We extracted data from these papers using a 50-question, theory-based questionnaire and a multiple-reviewer approach. Our sample represented almost 45 years of urban ecology research across 40 different countries. We found that more than 80% of the papers read were not transparent in their site selection methodology. We do not believe site selection methods are replicable for 70% of the papers read. Key weaknesses include incomplete descriptions of populations and sampling frames, urban gradients, sample selection methods, and property access. Low transparency in reporting the where methodology limits urban ecologists’ ability to assess the internal and external validity of studies’ findings and to replicate published studies; it also limits the generalizability of existing studies. The challenges of low transparency are particularly relevant in urban ecology, a field where standard protocols for site selection and delineation are still being developed. These limitations interfere with the fields’ ability to build theory and inform policy. We conclude by offering a set of recommendations to increase transparency, replicability, and generalizability.

54 ENVIRONMENTAL SCIENCES↗

Multi-site evaluation of stratified and balanced sampling of soil organic carbon stocks in agricultural fields

Estimating soil organic carbon (SOC) stocks in agricultural fields is essential for environmental and agronomic research, management, and policy. Stratified sampling is a classic strategy for estimating mean soil properties, and has recently been codified in SOC monitoring protocols. However, for the specific task of estimating the SOC stock of an agricultural field, concrete guidance is needed for which covariates to stratify on and how much stratification can improve estimation efficiency. It is also unknown how stratified sampling of SOC stocks compares to modern alternatives, notably doubly balanced sampling. To address these gaps, we collected high-density (average of 7 samples ha -1 ) and deep (average of 75 cm) measurements of SOC stocks at eight commercial fields under maize-soybean production in two US Midwestern states. We combined these measurements with a Bayesian geostatistical model to evaluate stratified and balanced sampling strategies that use a set of readily-available geographic, topographic, spectroscopic, and soil survey data. We examined the number of samples needed to achieve a given level of SOC stock estimation accuracy. While stratified sampling using these variables enables an average sample size reduction of 17% (95% CI, 11% to 23%) compared to simple random sampling, doubly balanced sampling is consistently more efficient, reducing sample sizes by 32% (95% CI, 25% to 37%). The data most important to these efficiency gains are a remotely-sensed SOC index, SSURGO estimates of SOC stocks, and the topographic wetness index. We conclude that in order to meet the urgent challenge of climate change, SOC stocks in agricultural fields could be more efficiently estimated by taking advantage of this readily-available data, especially with doubly balanced sampling.

54 ENVIRONMENTAL SCIENCES↗

Deployment of salt sample extraction system at an engineering-scale electrorefiner

The goal of the salt sampling program at Argonne is to develop and deploy automated molten salt sampling approaches for interfacing relevant unit operations with salt analysis to improve the timeliness of sampling-based accountancy measurements. Two technologies under development in support of this goal are a vacuum sampling loop module and a high-throughput pneumatic sample generator module. Compared to traditional point sampling approaches (i.e., dip probes), the vacuum sampling loop facilitates the collection of a larger cross-section of the bulk salt in order to collect more representative samples. The vacuum sampling approach also eliminates the risk of dross contamination of samples and avoids the use of moving parts in the salt. The pneumatic sample generator module is used to facilitate high-throughput sample analysis to improve the measurement precision of any given analytical technique by averaging out random sampling and measurement errors. In FY21, two methods for integrating these two modules were tested including direct fluidic coupling and coupling using a solid salt transfer mechanism. Solid salt transfer was ultimately selected over fluidic coupling, primarily to enable the transport of samples over longer distances to support automated at-line integration with high-precision techniques (such as microcalorimetry) that cannot withstand the conditions near an electrorefining process. To facilitate rapid solid salt coupling, new mechanisms were developed for rapidly charging and discharging salt sample tubes at the vacuum sampling loop and pneumatic sample generator modules, respectively. While the charging mechanism will be deployed in FY22, the discharge mechanism was tested in FY21 and is described here. The solid salt tube transfer method was deployed at one of Argonne’s engineering-scale electrorefiners to implement at-line high-throughput pneumatic micro-sample generation capabilities. The approach was used to generate precise uranium- and lanthanide-bearing electrorefiner micro-samples with the specific dimensions requested by researchers at Los Alamos National Laboratory for use in testing their novel microcalorimeter x-ray techniques. The solid salt transfer mechanism proved not only to be an effective means of integrating the precision sample generator with vacuum sampling, but also improved the performance of the sampler generator. To discharge salt from the sample tubes at the sampler generator, tube segments were inserted directly into the sample generator’s Helmholtz chamber and pressure pulse actuations were used to generate precision molten salt samples directly from the tube segments. The direct insertion of sample tubes into the sample generator enabled rapid loading of the salt and prevented salt from contacting most of the interior surfaces of the sample generator, which eliminated cross-contamination between runs. The vacuum sampling-loop tube charging mechanism will support high-throughput tube sampling operations by employing a dynamic vacuum filling process to fill short charge tubes that are configured to be rapidly connected and disconnected from the loop. The dynamic vacuum sampling operation will be automated, and sample tube handling can be executed with simple overhead actuation. Because the modular sampling approach described here eliminates the need for new high-radiation sample handling capabilities, salt-wetted seals, salt-wetted moving parts, and heated transfer lines outside the electrorefiner, it will address most of the remaining technical challenges for the at-line deployment of high-precision analytical techniques which will enable significant reductions in the time delay for sampling-based high-precision accountancy measurements.

42 ENGINEERING↗

COBRA:COMPUTED-TOMOGRAPHY BASED RANDOM-FIELD APPROXIMATION

SF-25-115 COBRA (COmputed-tomography Based Random-field Approximation) is a Python application for generating statistically equivalent random fields from CT-scan imagery. It leverages Karhunen–Loève expansions to model microstructural variability, enabling users to: Preprocess CT scans (filtering and Gaussian transformation); Fit covariance kernels fromempirical data; Solve eigenproblems to obtain KL modes; Sample random fields onsistent with fitted statistics; Postprocess samples back into the physical domain.

Hu, Tianchen↗

Self-Supervised Anomaly Detection via Neural Autoregressive Flows with Active Learning

Many self-supervised methods have been proposed with the target of image anomaly detection. These methods often rely on the paradigm of data augmentation with predefined transformations such as flipping, cropping, and rotations. However, it is not straightforward to apply these techniques for non-image data, such as time series or tabular data, while the performance of the existing deep approaches has been under our expectation on tasks beyond images. In this work, we propose a novel active learning (AL) scheme that relied on neural autoregressive flows (NAF) for self-supervised anomaly detection, specifically on small-scale data. Unlike other generative models such as GANs or VAEs, flow-based models allow to explicitly learn the probability density and thus can assign accurate likelihoods to normal data which makes it usable to detect anomalies. The proposed NAF-AL method is achieved by efficiently generating random samples from latent space and transforming them into feature space along with likelihoods via invertible mapping. The samples with lower likelihoods are selected and further checked by outlier detection using Mahalanobis distance. The augmented samples incorporating with normal samples are used for training a better detector so as to approach decision boundaries. Compared with random transformations, NAF-AL can be interpreted as a likelihood-oriented data augmentation that is more efficient and robust. Extensive experiments show that our approach outperforms existing baselines on multiple time series and tabular datasets, and a real-world application in advanced manufacturing, with significant improvement on anomaly detection accuracy and robustness over the state-of-the-art.

Zhang, Jiaxin↗

Rapid estimation of photosynthetic leaf traits of tropical plants in diverse environmental conditions using reflectance spectroscopy

Tropical forests are one of the main carbon sinks on Earth, but the magnitude of CO 2 absorbed by tropical vegetation remains uncertain. Terrestrial biosphere models (TBMs) are commonly used to estimate the CO 2 absorbed by forests, but their performance is highly sensitive to the parameterization of processes that control leaf-level CO 2 exchange. Direct measurements of leaf respiratory and photosynthetic traits that determine vegetation CO 2 fluxes are critical, but traditional approaches are time-consuming. Reflectance spectroscopy can be a viable alternative for the estimation of these traits and, because data collection is markedly quicker than traditional gas exchange, the approach can enable the rapid assembly of large datasets. However, the application of spectroscopy to estimate photosynthetic traits across a wide range of tropical species, leaf ages and light environments has not been extensively studied. Here, we used leaf reflectance spectroscopy together with partial least-squares regression (PLSR) modeling to estimate leaf respiration ( R dark25 ), the maximum rate of carboxylation by the enzyme Rubisco ( V cmax25 ), the maximum rate of electron transport ( J max25 ), and the triose phosphate utilization rate ( T p25 ), all normalized to 25°C. We collected data from three tropical forest sites and included leaves from fifty-three species sampled at different leaf phenological stages and different leaf light environments. Our resulting spectra-trait models validated on randomly sampled data showed good predictive performance for V cmax25 , J max25 , T p25 and R dark25 (RMSE of 13, 20, 1.5 and 0.3 μmol m -2 s -1 , and R 2 of 0.74, 0.73, 0.64 and 0.58, respectively). The models showed similar performance when applied to leaves of species not included in the training dataset, illustrating that the approach is robust for capturing the main axes of trait variation in tropical species. We discuss the utility of the spectra-trait and traditional gas exchange approaches for enhancing tropical plant trait studies and improving the parameterization of TBMs.

54 ENVIRONMENTAL SCIENCES↗

i- flow: High-dimensional integration and sampling with normalizing flows

In many fields of science, high-dimensional integration is required. Numerical methods have been developed to evaluate these complex integrals. We introduce the code i-flow, a python package that performs high-dimensional numerical integration utilizing normalizing flows. Normalizing flows are machine-learned, bijective mappings between two distributions. i-flow can also be used to sample random points according to complicated distributions in high dimensions. We compare i-flow to other algorithms for high-dimensional numerical integration and show that i-flow outperforms them for high dimensional correlated integrals. The i-flow code is publicly available on gitlab at https://gitlab.com/i-flow/i-flow.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

(U) A General-Purpose Code for Correlated Sampling Using Batch Statistics with MCNP6 for Fixed-Source Problems

Correlated sampling can be used to reduce the uncertainty of a difference of tallies by taking advantage of the negative covariance term in the sandwich formula. Booth first showed how correlated sampling can be applied with batch statistics using MCNP’s tally fluctuation chart (TFC) to reduce the uncertainty of a difference of tallies in fixed-source problems. Booth presented a problem in which a 1273% uncertainty in a difference was reduced to 8% by accounting for correlations. Researchers He and Su recently studied correlated sampling using the TFC in MCNP version 5. They determined that the code did not print enough digits in the TFC tally means for accurate batch statistics in some cases. After modifying the source code, they concluded that “correlated sampling can yield a standard deviation of about one magnitude smaller than that predicted by the direct, un-correlated simulation when the changes in system response are small (say about 1%), which is equivalent to saving in CPU time by a factor of 100. Such saving [sic] becomes less significant as the change in system response becomes larger.” He and Su provided the formulas needed to apply batch statistics to compute the correlated uncertainty of a difference of tallies. In this report, we follow up on their work by providing the formulas needed to apply batch statistics to compute the correlated uncertainty of a ratio of tallies and of a difference of two tallies divided by a third tally. We extend these formulas to differences and ratios of ratios. These formulas are applied to reduce the uncertainty associated with calculating a relative sensitivity. He and Su did not investigate the accuracy of their correlated sampling uncertainty estimates. We use their test problems and evaluate the accuracy of the uncertainty estimates by comparing with results obtained from random sampling, and, in simple cases, with theoretical values of the “exact” uncertainties. We find that the uncertainties obtained from batch statistics are accurate as long as at least 100 batches are used. We present a new computer code, COSUBS (COrrelated Sampling Using Batch Statistics), that reads MCNP6 TFCs and applies correlated sampling using batch statistics for the tally combinations that the user specifies. COSUBS is a very general tool that compares all TFCs for a base case and one or two perturbed cases. It computes uncertainties for ratios if given only a base case. This report is organized as follows. The equations to apply batch statistics to the difference of random tallies are reviewed in Sec. II. Section III presents the equations for applying batch statistics to a ratio of random tallies; this is useful for computing relative sensitivities using a one-sided finite difference and the relative sensitivity using the differential operator method. Section IV presents the equations for applying batch statistics to a difference of two random tallies divided by a third; this is useful for computing a relative sensitivities using a central difference. Section V presents the equations for applying batch statistics to a difference of two ratios with four random tallies. Section VI presents the equations for applying batch statistics to a one-sided finite difference estimate of the relative sensitivity of a ratio (this uses four random tallies). Section VII presents the equations for applying batch statistics to a central difference estimate of the relative sensitivity of a ratio (this uses six random tallies). Section VIII presents the equations for applying batch statistics to a sum of random tallies. Section IX discusses how to apply batch statistics using MCNP6. Section X presents COSUBS, describing its command-line options and logic. Sections XI through XVI present numerical results for various test problems. Section XVII is a summary and conclusions. Appendix A derives the theoretical Monte Carlo tally variance given certain assumptions; these variances are used to verify the batch statistics for some of the problems. Appendix B lists the MCNP6 input for the unperturbed example problem. Appendix C presents modifications made to MCNP6.3 to support this work.

97 MATHEMATICS AND COMPUTING↗

Adaptive Data-Driven Deep-Learning Surrogate Model for Frontal Polymerization in Dicyclopentadiene

Frontal polymerization (FP) is a self-sustaining curing process that enables rapid and energy-efficient manufacturing of thermoset polymers and composites. Computational methods conventionally used to simulate the FP process are time-consuming, and repeating simulations are required for sensitivity analysis, uncertainty quantification, or optimization of the manufacturing process. Here, in this work, we develop an adaptive surrogate deep-learning model for FP of dicyclopentadiene (DCPD), which predicts the evolution of temperature and degree of cure orders of magnitude faster than the finite-element method (FEM). The adaptive algorithm provides a strategy to select training samples efficiently and save computational costs by reducing the redundancy of FEM-based training samples. The adaptive algorithm calculates the residual error of the FP governing equations using automatic differentiation of the deep neural network. A probability density function expressed in terms of the residual error is used to select training samples from the Sobol sequence space. The temperature and degree of cure evolution of each training sample are obtained by a 2D FEM simulation. The adaptive method is more efficient and has a better prediction accuracy than the random sampling method. With the well-trained surrogate neural network, the FP characteristics (front speed, shape, and temperature) can be extracted quickly from the predicted temperature and degree-of-cure fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Milky Way mass with K giants and BHB stars using LAMOST, SDSS/SEGUE, and Gaia : 3D spherical Jeans equation and tracer mass estimator

ABSTRACT We measure the enclosed Milky Way mass profile to Galactocentric distances of ∼70 and ∼50 kpc using the smooth, diffuse stellar halo samples of Bird et al. The samples are Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST) and Sloan Digital Sky Survey/Sloan Extension for Galactic Understanding and Exploration (SDSS/SEGUE) K giants (KG) and SDSS/SEGUE blue horizontal branch (BHB) stars with accurate metallicities. The 3D kinematics are available through LAMOST and SDSS/SEGUE distances and radial velocities and Gaia DR2 proper motions. Two methods are used to estimate the enclosed mass: 3D spherical Jeans equation and Evans et al. tracer mass estimator (TME). We remove substructure via the Xue et al. method based on integrals of motion. We evaluate the uncertainties on our estimates due to random sampling noise, systematic distance errors, the adopted density profile, and non-virialization and non-spherical effects of the halo. The tracer density profile remains a limiting systematic in our mass estimates, although within these limits we find reasonable agreement across the different samples and the methods applied. Out to ∼70 and ∼50 kpc, the Jeans method yields total enclosed masses of 4.3 ± 0.95 (random) ±0.6 (systematic) × 1011 M⊙ and 4.1 ± 1.2 (random) ±0.6 (systematic) × 1011 M⊙ for the KG and BHB stars, respectively. For the KG and BHB samples, we find a dark matter virial mass of $M_{200}=0.55^{+0.15}_{-0.11}$ (random) ±0.083 (systematic) × 1012 M⊙ and $M_{200}=1.00^{+0.67}_{-0.33}$ (random) ±0.15 (systematic) × 1012 M⊙, respectively.

79 ASTRONOMY AND ASTROPHYSICS↗

Grover-QAOA for 3-SAT: quadratic speedup, fair-sampling, and parameter clustering

Abstract The SAT problem is a prototypical NP-complete problem of fundamental importance in computational complexity theory with many applications in science and engineering; as such, it has long served as an essential benchmark for classical and quantum algorithms. This study shows numerical evidence for a quadratic speedup of the Grover Quantum Approximate Optimization Algorithm (G-QAOA) over random sampling for finding all solutions to 3-SAT (All-SAT) and Max-SAT problems. G-QAOA is less resource-intensive and more adaptable for these problems than Grover’s algorithm, and it surpasses conventional QAOA in its ability to sample all solutions. We show these benefits by classical simulations of many-round G-QAOA on thousands of random 3-SAT instances. We also observe G-QAOA advantages on the IonQ Aria quantum computer for small instances, finding that current hardware suffices to determine and sample all solutions. Interestingly, a single-angle-pair constraint that uses the same pair of angles at each G-QAOA round greatly reduces the classical computational overhead of optimizing the G-QAOA angles while preserving its quadratic speedup. We also find parameter clustering of the angles. The single-angle-pair protocol and parameter clustering significantly reduce obstacles to classical optimization of the G-QAOA angles.

Zhang, Zewen (ORCID:000000032258613X)↗

Power Grid Reliability Estimation via Adaptive Importance Sampling

Electricity production currently generates approximately 25% of greenhouse gas emissions in the USA. Thus, increasing the amount of renewable energy is a key step to carbon neutrality. However, integrating a large amount of fluctuating renewable generation is a significant challenge for power grid operating and planning. Grid reliability, i.e., an ability to meet operational constraints under power fluctuations, is probably the most important of them. In this letter, we propose computationally efficient and accurate methods to estimate the probability of line overflow, i.e., reliability constraints violation, under a known distribution of renewable energy generation. To this end, we investigate an importance sampling approach, a flexible extension of Monte-Carlo methods, which adaptively changes the sampling distribution to generate more samples near the reliability boundary. The approach allows to estimate overload probability in real-time based only on a few dozens of random samples, compared to thousands required by the plain Monte-Carlo. Our study focuses on high voltage direct current power transmission grids with linear reliability constraints on power injections and line currents. Herein, we propose a novel theoretically justified physics-informed adaptive importance sampling algorithm and compare its performance to state-of-the-art methods on multiple IEEE power grid test cases.

power system faults↗

Curiosity driven exploration to optimize structure–property learning in microscopy

Rapidly determining structure–property correlations in materials is an important challenge in better understanding fundamental mechanisms and greatly assists in materials design. In microscopy, imaging data provides a direct measurement of the local structure, while spectroscopic measurements provide relevant functional property information. Deep kernel active learning approaches have been utilized to rapidly map local structure to functional properties in microscopy experiments, but are computationally expensive for multi-dimensional and correlated output spaces. Here, we present an alternative lightweight curiosity algorithm which actively samples regions with unexplored structure–property relations, utilizing a deep-learning based surrogate model for error prediction. We show that the algorithm outperforms random sampling for predicting properties from structures, and provides a convenient tool for efficient mapping of structure–property relationships in materials science.

36 MATERIALS SCIENCE↗

The Role of Data Filtering in Open Source Software Ranking and Selection

Faced with more than 100M open source projects, a more manageable small subset is needed for most empirical investigations. More than half of the research papers in leading venues investigated filtering projects by some measure of popularity with explicit or implicit arguments that unpopular projects are not of interest, may not even represent "real" software projects, or that less popular projects are not worthy of study. However, such filtering may have enormous effects on the results of the studies if and precisely because the sought-out response or prediction is in any way related to the filtering criteria.This paper exemplifies the impact of this common practice on research outcomes, specifically how filtering of software projects on GitHub based on inherent characteristics affects the assessment of their popularity. Using a dataset of over 100,000 repositories, we used multiple regression to model the number of stars -a commonly used proxy for popularity- based on factors such as the number of commits, the duration of the project, the number of authors and the number of core developers. Our control model included the entire dataset, while a second filtered model considered only projects with ten or more authors. The results indicated that while certain characteristics of the repository consistently predict popularity, the filtering process significantly alters the relationships between these characteristics and the response. We found that the number of commits exhibited a positive correlation with popularity in the control sample but showed a negative correlation in the filtered sample. These findings highlight the potential biases introduced by data filtering and emphasize the need for careful sample selection in empirical research of mining software repositories. We recommend that empirical work should either analyze complete datasets such as World of Code, or employ stratified random sampling from a complete dataset to ensure that filtering is not biasing the results.

Malviya Thakur, Addi↗

A Case Study in Assessing a Potential Severity Framework for Incidents from a Decadal Sample

In this study, the primary objective of this case study is to determine the applicability and feasibility of a framework that leverages occupational incident details to prospectively identify “potential Serious Injury or Fatality” (pSIF) cases. This study comprehensively reviewed a random sample of 1,081 injury and illness cases across 21 generalized incident types spanning over a decade at Lawrence Livermore National Laboratory (LLNL), a U.S. Department of Energy research and development facility with more than 9,000 employees. The review applied a general framework that classified each case on information suitability, potential severity, and future incident mitigation. The findings from the study indicate that 86.6% of the cases had sufficient information to make a high-confidence determination on potential severity, underscoring the feasibility of applying this general framework. Additionally, cases with a higher pSIF score had, on average, a higher level of institutional response. Implementing a simplified methodology for incident classification that emphasizes incidents that pose high potential severity, regardless of incident type, can help LLNL prioritize resources and tailor responses to such incidents using a graded approach. LLNL has recognized the value of this capability and is integrating the framework into their injury and illness process in the 2024 calendar year.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Computing the QRPA level density with the finite amplitude method

Here, we describe a new algorithm to calculate the vibrational nuclear level density of an atomic nucleus. Fictitious perturbation operators that probe the response of the system are generated by drawing their matrix elements from some probability distribution function. We use the Finite Amplitude Method to explicitly compute the response for each such sample. With the help of the Kernel Polynomial Method, we build an estimator of the vibrational level density and provide the upper bound of the relative error in the limit of infinitely many random samples. The new algorithm can give accurate estimates of the vibrational level density. Since it is based on drawing multiple samples of perturbation operators, its computational implementation is naturally parallel and scales like the number of available processing units.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗