Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Soybean ( Glycine max ) Haplotype Map (GmHapMap): a universal resource for soybean translational and functional genomics

Here, we describe a worldwide haplotype map for soybean (GmHapMap) constructed using whole-genome sequence data for 1007 Glycine max accessions and yielding 14.9 million variants as well as 4.3 M tag single-nucleotide polymorphisms (SNPs). When sampling random subsets of these accessions, the number of variants and tag SNPs plateaued beyond approximately 800 and 600 accessions, respectively. This suggests extensive coverage of diversity within the cultivated soybean. GmHapMap variants were imputed onto 21 618 previously genotyped accessions with up to 96% success for common alleles. A local association analysis was performed with the imputed data using markers located in a 1-Mb region known to contribute to seed oil content and enabled us to identify a candidate causal SNP residing in the NPC1 gene. We determined gene-centric haplotypes (407 867 GCHs) for the 55 589 genes and showed that such haplotypes can help to identify alleles that differ in the resulting phenotype. Finally, we predicted 18 031 putative loss-of-function (LOF) mutations in 10 662 genes and illustrated how such a resource can be used to explore gene function. The GmHapMap provides a unique worldwide resource for applied soybean genomics and breeding.

54 ENVIRONMENTAL SCIENCES↗

Stochastic Gradients for Large-Scale Tensor Decomposition

Tensor decomposition is a well-known tool for multiway data analysis. This work proposes using stochastic gradients for efficient generalized canonical polyadic (GCP) tensor decomposition of large-scale tensors. GCP tensor decomposition is a recently proposed version of tensor decomposition that allows for a variety of loss functions such as Bernoulli loss for binary data or Huber loss for robust estimation. Here, the stochastic gradient is formed from randomly sampled elements of the tensor and is efficient because it can be computed using the sparse matricized-tensor times Khatri--Rao product tensor kernel. For dense tensors, we simply use uniform sampling. For sparse tensors, we propose two types of stratified sampling that give precedence to sampling nonzeros. Numerical results demonstrate the advantages of the proposed approach and its scalability to large-scale problems.

97 MATHEMATICS AND COMPUTING↗

Tidal Disruption Event Galaxy Binner

This software simulates astronomical survey detections of tidal disruptions of stars by super-massive black holes. It begins with the synthetic galaxy catalogue described in van Velzen 2008 (https://arxiv.org/abs/1707.03458). The stellar disruption rate in each galaxy is estimated based on Stone & Metzger 2016 (https://arxiv.org/abs/1410.7772). Based on these rates, and the present-day stellar mass function in the galaxy, disruptions are randomly sampled, and the properties of the resulting flares are sampled based on empirical distributions. The code also accounts for obscuration by dust in the host galaxy. Finally, the survey selection effects are applied. The detectable simulated flares are stored in a database, allowing histograms of their properties to be created.

Roth, NathanielJ.↗

Reliability of Open Public Electric Vehicle Direct Current Fast Chargers

The aim was to systematically evaluate the usability of all public electric vehicles (EV) direct current fast chargers (DCFC) in the San Francisco region. To achieve a rapid transition to EVs, a highly reliable and easy to use charging infrastructure is critical to building confidence among consumers. The functionality and usability of all 182 open, public DCFC charging stations with CCS connectors (combined charging system) in the 9 counties of the Bay Area were tested (655 electric vehicle service equipment (EVSE) ports). An EVSE was classified as functional if it charged an EV for 2 minutes. Overall, 73.3% of the 655 EVSEs were functional. The causes of the nonfunctioning EVSEs (23.5%) were blank or unresponsive screens or error messages; payment system failures; charge initiation failures; network failures; or broken connectors. In addition, the cable was too short to reach the EV inlet for 3.2% of the EVSEs. A random sampling of 10% of the EVSEs, approximately 8 days after the first evaluation, found no overall change in functionality. The level of functionality found with field testing conflicts with the 95–98% uptime reported by the EV service providers (EVSPs) who operate the EV charging stations. There is a need for precise and verifiable definitions of uptime, downtime, and excluded time, as applied to public EV chargers. In conclusion, the level of failure of the existing public EV DCFC charge infrastructure highlights the importance of improving the system design and maintenance to improve adoption of EVs.

33 ADVANCED PROPULSION SYSTEMS↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

COVID-19 prevention at institutions of higher education, United States, 2020–2021: implementation of nonpharmaceutical interventions

Background, In early 2020, following the start of the coronavirus disease 2019 (COVID-19) pandemic, institutions of higher education (IHEs) across the United States rapidly pivoted to online learning to reduce the risk of on-campus virus transmission. We explored IHEs’ use of this and other nonpharmaceutical interventions (NPIs) during the subsequent pandemic-affected academic year 2020–2021. Methods, From December 2020 to June 2021, we collected publicly available data from official webpages of 847 IHEs, including all public (n = 547) and a stratified random sample of private four-year institutions (n = 300). Abstracted data included NPIs deployed during the academic year such as changes to the calendar, learning environment, housing, common areas, and dining; COVID-19 testing; and facemask protocols. We performed weighted analysis to assess congruence with the October 29, 2020, US Centers for Disease Control and Prevention (CDC) guidance for IHEs. For IHEs offering ≥50% of courses in person, we used weighted multivariable linear regression to explore the association between IHE characteristics and the summated number of implemented NPIs. Results, Overall, 20% of IHEs implemented all CDC-recommended NPIs. The most frequently utilized NPI was learning environment changes (91%), practiced as one or more of the following modalities: distance or hybrid learning opportunities (98%), 6-ft spacing (60%), and reduced class sizes (51%). Additionally, 88% of IHEs specified facemask protocols, 78% physically changed common areas, and 67% offered COVID-19 testing. Among the 33% of IHEs offering ≥50% of courses in person, having < 1000 students was associated with having implemented fewer NPIs than IHEs with ≥ 1000 students. Conclusions, Only 1 in 5 IHEs implemented all CDC recommendations, while a majority implemented a subset, most commonly changes to the classroom, facemask protocols, and COVID-19 testing. IHE enrollment size and location were associated with degree of NPI implementation. Additional research is needed to assess adherence to NPI implementation in IHE settings.

59 BASIC BIOLOGICAL SCIENCES↗

Designing an Optimal Sensor Network via Minimizing Information Loss

Optimal experimental design is a classic topic in statistics, with many well-studied problems, applications, and solutions. The design problem we study is the placement of sensors to monitor spatiotemporal processes, explicitly accounting for the temporal dimension in our modeling and optimization. We observe that recent advancements in computational sciences often yield large datasets based on physics-based simulations, which are rarely leveraged in experimental design. We introduce a novel model-based sensor placement criterion, along with a highly-efficient optimization algorithm, which integrates physics-based simulations and Bayesian experimental design principles to identify sensor networks that “minimize information loss” from simulated data. Our technique relies on sparse variational inference and (separable) Gauss-Markov priors, and thus may adapt many techniques from Bayesian experimental design. We validate our method through a case study monitoring air temperature in Phoenix, Arizona, using state-of-the-art physics-based simulations. Our results show our framework to be superior to random or quasi-random sampling, particularly with a limited number of sensors. We conclude by discussing practical considerations and implications of our framework, including more complex modeling tools and real-world deployments.

54 ENVIRONMENTAL SCIENCES↗

Long Term Per-Component Power and Thermal Measurements of the OLCF Summit System

As we move into the exascale era, the power and energy footprints of high-performance computing (HPC) systems have grown significantly larger. Due to the harsh power and thermal conditions the system, components are exposed to extreme operating conditions. Operation of such modern HPC systems requires deep insights into long term system behavior to maintain its efficiency as well as its longevity. To help the HPC community to gain such insights, we provide a dataset that records the long-term power and thermal behavior of the 200PF pre-exascale supercomputer at the Oak Ridge Leadership Computing Facility (OLCF), Summit. This system is an IBM AC922 based system that has 9,252 IBM Power9 CPUs and 27,756 Nvidia V100 GPUs and can consume up to 13MW power at peak. Heat removal is performed using medium temperature direct liquid cooling and rear-door heat exchanger based secondary cooling loop. Originally extracted from a high-resolution (1Hz) per-component (GPUs, CPUs) measurements from the system, we primarily provide a dataset that has 10-second and 1-minute mean power and thermal measurements selected from five month-long segments over the course of 2020 (January and August), 2021 (February and August), and 2022 (January). For convenience, we also provide various sub datasets randomly sampled from the time and space (hosts) of the cluster. Further details and example code for analysis can be found in the following GitHub repository: https://github.com/at-aaims/summit_power_and_thermal_data

97 MATHEMATICS AND COMPUTING↗

Proposal and application of ROM-Lasso method for sensitivity coefficient evaluation

We propose a novel method for evaluating sensitivity coefficients of neutronics parameters to cross sections, so-called the reduced-order modeling technique ROM-Lasso. In this method, cross sections of interest are randomly sampled, and corresponding perturbed core analyses are performed. Then, the sensitivity coefficient vector of the higher-level model is expanded via the active subspace bases obtained with the lower-level model whose dimensional complexity is smaller than that of the higher-level model, and the expansion coefficients are estimated by the Lasso regression. A unique feature of the ROM-Lasso method allows the use of different bases optimized for each neutronics parameter. We conducted a verification calculation for an accelerator-driven system and demonstrated that the ROM-Lasso method can reproduce the sensitivity coefficients with a much smaller number of forward calculations than the direct method. The proposed method can be used to practically evaluate sensitivity coefficients. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Description and Use of SCALE Sampler Parametric Capability for Engineering Analysis and Optimization

The Sampler sequence was introduced into the SCALE nuclear modeling and simulation suite in SCALE 6.2 to perform uncertainty quantification via random sampling of nuclear data, material number densities, and dimensions. Sampler was expanded with the introduction of a parametric capability in SCALE 6.2.2. This paper discusses input for the Sampler parametric sequence and presents two case studies of analyses performed using the sequence. These case studies include preconceptual design of a package for transporting high assay low-enriched uranium (HALEU) oxide and scoping calculations to support subcritical limit development for a future update of the ANSI/ANS-8.1 (ANS-8.1) standard. The parametric capability within Sampler provides many benefits to analysts. For instance, parametric sweeps are frequently used to identify optimum parameter values as part of safety analysis or system design, but such sweeps can require substantial engineering time or may rely on custom-written scripts or scripts such as Write One, Run Many (or WORM) developed outside of any software quality assurance program. With the parametric capabilities in Sampler, however, a large number of inputs can be generated automatically without recourse to scripting by individual analysts. The parametric capability can also be used in lieu of the CSAS5S search sequence to identify optimum parameters more simply with straightforward inputs and outputs. Sampler can also be used to calculate input parameters from engineering specifications. For example, diameters can be converted to radii, or masses can be used to calculate number densities. Overall, the Sampler parametric capability provides a robust feature within SCALE, eliminating the need for user-developed scripting.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ptychographic wavefront characterization for single-particle imaging at x-ray lasers

A well-characterized wavefront is important for many x-ray free-electron laser (XFEL) experiments, especially for single-particle imaging (SPI), where individual biomolecules randomly sample a nanometer region of highly focused femtosecond pulses. We demonstrate high-resolution multiple-plane wavefront imaging of an ensemble of XFEL pulses, focused by Kirkpatrick–Baez mirrors, based on mixed-state ptychography, an approach letting us infer and reduce experimental sources of instability. From the recovered wavefront profiles, we show that while local photon fluence correction is crucial and possible for SPI, a small diversity of phase tilts likely has no impact. Our detailed characterization will aid interpretation of data from past and future SPI experiments and provides a basis for further improvements to experimental design and reconstruction algorithms.

47 OTHER INSTRUMENTATION↗

Transit Rider/Travel Behavior Inventory Survey - Minneapolis-St. Paul Metro - 1990

The 1990 transit on-board survey aimed to update the 1988 survey, which was conducted as part of the Preliminary Engineering Study for the Hennepin County Light Rail Transit System. The 1988 study received Regional Transit Board funding, and survey results could be applied to a mode split model for projecting ridership on the proposed Light Rail Transit System. However, this survey was designed with the 1990 Travel Behavior Inventory in mind. The intention had been to update the 1988 survey in 1990 to be compatible with the 1990 Travel Behavior Inventory data. The results of the 1990 survey were used primarily to create a table of observed transit trips between each of the 1,200 traffic analysis zones in the region. This trip table was used to calibrate a new mode split model, which estimated future year travel by mode. The 1990 update survey focused on new routes and routes that had changed significantly since 1988. To preserve compatibility with the 1988 survey, the same survey questionnaire card was used, together with the same survey procedures for data collection. The procedures randomly sampled bus patrons during the transit trip, asking key questions about the patron and the transit trip. The survey card was intended for patrons to fill out quickly so it could be completed during the transit trip. The questions focused on conditions that have proven over time to significantly influence ridership. In all, a total of 20,126 valid survey records were processed. Adding surrogate trips, the survey data file is composed of a total of 27,159 trip records . About 10% of the records were filled out by persons who had answered more than one questionnaire.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Quantifying Uncertainty in All-to-All Estimates of Space Object Conjunction Probabilities using U-Statistics

Predicting space object conjunctions is inherently probabilistic due to initial state and orbit model uncertainty. A commonly considered Monte Carlo estimator of the conjunction probability is the ’all-to-all’ estimator. Given independent random samples of the trajectories of both objects, the estimator is the percentage of all pairs of trajectories that result in a conjunction. Intuitively, the all-to-all estimator is the best possible estimator of the conjunction probability since it considers all pairs of Monte Carlo samples. However, its distribution is not available in closed-form, which limits its use in practice and makes this intuition difficult to make rigorous. In this paper, the all-to-all estimator is identified as a U-statistic, which implies that it has several favorable properties. Specifically, the estimator is the minimum variance unbiased estimator of the conjunction probability and is asymptotically Gaussian distributed. An approximate confidence interval for the conjunction probability is obtained from an estimate of the asymptotic Gaussian distribution. We show how to efficiently compute the confidence interval and demonstrate that the interval has the nominal coverage level. The confidence intervals are also seen to be narrower than those based on the commonly-used each-to-each estimator. Furthermore, the all-to-all estimator is shown to allow different Monte Carlo sample sizes, whereas the each-to-each estimator requires equal sample sizes.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sensitivity and Uncertainty of the IFR-1 BISON Benchmark

The fuel performance code BISON is being used to evaluate metallic fuel for a new fast-spectrum test reactor called the Versatile Test Reactor, which is being considered by the US Department of Energy. To quantify the accuracy of BISON predictions, researchers at Oak Ridge National Laboratory have been developing a series of benchmarks based on legacy metallic fuel experiments. As part of this effort, the sensitivity of BISON predictions to variations in model inputs and the uncertainties associated with BISON predictions must be established. This report summarizes efforts to perform a comprehensive sensitivity analysis (SA) and uncertainty quantification (UQ) on a benchmark based on the IFR-1 experiment. For the SA, at least one input was chosen from every BISON model and physics module used in the benchmark. The inputs were varied individually in a series of BISON simulations. The resulting variations in benchmark predictions were normalized to calculate sensitivities. The strongest sensitivities were identified and used to inform input selections for the UQ. The UQ was performed using the Monte Carlo UQ method. A literature review was conducted to estimate uncertainty distributions for the selected inputs, and values were sampled randomly from each distribution in a series of BISON simulations. Variations in the benchmark predictions were used to estimate uncertainty distributions and confidence intervals. It was found that nearly 100% of benchmark predictions matched the corresponding legacy values within the confidence intervals. However, this is at least partially because of the wide confidence intervals associated with the benchmark predictions. The uncertainty contributions of assumptions in the benchmark, experimental uncertainties, and BISON models were quantified. Some analysis was performed to identify inputs that contributed to the uncertainties. Finally, recommendations are made for future benchmark development and future BISON development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Monte Carlo Transport: Computational Physics Summer Workshop [Slides]

In general, Monte Carlo methods simulate large numbers of random trials in order to observe numerical behavior of systems described by probabilistic behavior. In radiation transport, pseudo-random number generators are used to randomly sample individual particle lives. Information about the particles are tallied.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Simultaneous Discovery of Positive and Negative Interactions Among Rhizosphere Bacteria Using Microwell Recovery Arrays

Understanding microbe-microbe interactions is critical to predict microbiome function and to construct communities for desired outcomes. Investigation of these interactions poses a significant challenge due to the lack of suitable experimental tools available. Here we present the microwell recovery array (MRA), a new technology platform that screens interactions across a microbiome to uncover higher-order strain combinations that inhibit or promote the function of a focal species. One experimental trial generates 10 4 microbial communities that contain the focal species and a distinct random sample of uncharacterized cells from plant rhizosphere. Cells are sequentially recovered from individual wells that display highest or lowest levels of focal species growth using a high-resolution photopolymer extraction system. Interacting species are then identified and putative interactions are validated. Using this approach, we screen the poplar rhizosphere for strains affecting the growth of Pantoea sp. YR343, a plant growth promoting bacteria isolated from Populus deltoides rhizosphere. In one screen, we montiored 3,600 microwells within the array to uncover multiple antagonistic Stenotrophomonas strains and a set of Enterobacter strains that promoted YR343 growth. The later demonstrates the unique ability of the platform to discover multi-membered consortia that generate emergent outcomes, thereby expanding the range of phenotypes that can be characterized from microbiomes. This knowledge will aid in the development of consortia for Populus production, while the platform offers a new approach for screening and discovery of microbial interactions, applicable to any microbiome.

59 BASIC BIOLOGICAL SCIENCES↗

A fast Monte Carlo cell-by-cell simulation for radiobiological effects in targeted radionuclide therapy using pre-calculated single-particle track standard DNA damage data

Introduction: We developed a new method that drastically speeds up radiobiological Monte Carlo radiation-track-structure (MC-RTS) calculations on a cell-by-cell basis. Methods: The technique is based on random sampling and superposition of single-particle track (SPT) standard DNA damage (SDD) files from a “pre-calculated” data library, constructed using the RTS code TOPAS-nBio, with “time stamps” manually added to incorporate dose-rate effects. This time-stamped SDD file can then be input into MEDRAS, a mechanistic kinetic model that calculates various radiation-induced biological endpoints, such as DNA double-strand breaks (DSBs), misrepairs and chromosomal aberrations, and cell death. As a benchmark validation of the approach, we calculated the predicted energy-dependent DSB yield and the ratio of direct-to-total DNA damage, both of which agreed with published in vitro experimental data. We subsequently applied the method to perform a superfast cell-by-cell simulation of an experimental in vitro system consisting of neuroendocrine tumor cells uniformly incubated with 177 Lu. Results and discussion: The results for residual DSBs, both at 24 and 48 h post-irradiation, are in line with the published literature values. Our work serves as a proof-of-concept demonstration of the feasibility of a cost-effective “in silico clonogenic cell survival assay” for the computational design and development of radiopharmaceuticals and novel radiotherapy treatments more generally.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Coreset Clustering on Small Quantum Computers

Many quantum algorithms for machine learning require access to classical data in superposition. However, for many natural data sets and algorithms, the overhead required to load the data set in superposition can erase any potential quantum speedup over classical algorithms. Recent work by Harrow introduces a new paradigm in hybrid quantum-classical computing to address this issue, relying on coresets to minimize the data loading overhead of quantum algorithms. We investigated using this paradigm to perform k-means clustering on near-term quantum computers, by casting it as a QAOA optimization instance over a small coreset. We used numerical simulations to compare the performance of this approach to classical k-means clustering. We were able to find data sets with which coresets work well relative to random sampling and where QAOA could potentially outperform standard k-means on a coreset. However, finding data sets where both coresets and QAOA work well—which is necessary for a quantum advantage over k-means on the entire data set—appears to be challenging.

42 ENGINEERING↗