Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Computing the QRPA level density with the finite amplitude method

Here, we describe a new algorithm to calculate the vibrational nuclear level density of an atomic nucleus. Fictitious perturbation operators that probe the response of the system are generated by drawing their matrix elements from some probability distribution function. We use the Finite Amplitude Method to explicitly compute the response for each such sample. With the help of the Kernel Polynomial Method, we build an estimator of the vibrational level density and provide the upper bound of the relative error in the limit of infinitely many random samples. The new algorithm can give accurate estimates of the vibrational level density. Since it is based on drawing multiple samples of perturbation operators, its computational implementation is naturally parallel and scales like the number of available processing units.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

hiposa v0.1.0

Fast construction of hierarchical N-dimensional experimental design. The software allows construction of sampling strategies that are self-avoiding, tunable in resolution in N dimensions. This code is a drop-in replacement for seeding optimization runs often performed using uniform random sampling. The hierarchical approach allows for a rapid, top down method for construction the most effective design that drives towards a global optimum.

Zwart, PetrusH [Lawrence Berkeley National Laborat↗

Energy Delivery Systems with Verifiable Trustworthiness (Final Report)

Energy Delivery Systems (EDS) must be verified to be free from intrusive and malicious software. One way of verifying this software is to perform device scans to detect malicious code. Because it is possible to have “fileless” malware that exists only in device (volatile) memory, offline scanning and even many forms of online scanning is insufficient for detection. This project (“Verify”) addresses this need by performing direct sampling of memory during device operation to detect unexpected or modified software while not interfering with device operation. The Verify project provides a proof-of-concept of detection by random sampling combined with remote software- and timing-based attestation methods for robust detection of in-memory threats. An external review of Verify was performed by our partner, General Electric (GE), and a summary of their findings is provided.

97 MATHEMATICS AND COMPUTING↗

A Machine Learning-Based Vulnerability Analysis for Cascading Failures of Integrated Power-Gas Systems

This article proposes a cascading failure simulation (CFS) method and a hybrid machine learning method for vulnerability analysis of integrated power-gas systems (IPGSs). The CFS method is designed to study the propagating process of cascading failures between the two systems, generating data for machine learning with initial states randomly sampled. The proposed method considers generator and gas well ramping, transmission line and gas pipeline tripping, island issue handling and load shedding strategies. Then, a hybrid machine learning model with a combined random forest (RF) classification and regression algorithms is proposed to investigate the impact of random initial states on the vulnerability metrics of IPGSs. Extensive case studies are carried out on three test IPGSs to verify the proposed models and algorithms. Simulation results show that the proposed models and algorithms can achieve high accuracy for the vulnerability analysis of IPGSs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deconsolidation and Leach Burn Leach of Seven As Irradiated AGR 5/6/7 TRISO Fuel Compacts from Capsules 2, 3, 4, and 5

Seven as-irradiated Advanced Gas Reactor (AGR) 5/6/7 compacts underwent destructive post-irradiation examination via deconsolidation-leach-burn leach at Idaho National Laboratory (INL). The selection of the compacts extended the upper and lower limits of time-average volume-average (TAVA) temperature for compacts that had gone through deconsolidation-leach-burn leach so far. The measured inventories of fission products and actinides in the compact matrix and outer pyrolytic carbon were reported. Results indicated unexpectedly higher rates of fuel kernel leaching compared to compacts from AGR-1 and AGR-2. These failure rates were attributed to damage during post-irradiation sample handling, rather than irradiation itself. AGR-5/6/7 compacts have little or no matrix coverage for some particles at the top and bottom ends of cylindrical fuel compacts, making them more fragile. Sixty particles were randomly sampled from each compact, and the gamma results were reported. Three SiC shells from were identified, one of which showed signs of chemical attack in the high-irradiation temperature compact.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Intelligent Sampling for Surrogate Modeling, Hyperparameter Optimization, and Data Analysis

Sampling techniques are used in many fields, including design of experiments, image processing, and graphics. The techniques in each field are designed to meet the constraints specific to that field such as uniform coverage of the range of each dimension or random samples that are at least a certain distance apart from each other. When an application imposes new constraints, for example, by requiring samples in a non-rectangular domain or the addition of new samples to an existing set, a common solution is to modify the algorithm currently in use, often with less than satisfactory results. As an alternative, we propose the concept of intelligent sampling, where we devise algorithms specifically tailored to meet our sampling needs, either by creating new algorithms or by modifying suitable algorithms from other fields. Surprisingly, both qualitative and quantitative comparisons indicate that some relatively simple algorithms can be easily modified to meet the many sampling requirements of surrogate modeling, hyperparameter optimization, and data analysis; these algorithms outperform their more sophisticated counterparts currently in use, resulting in better use of time and computer resources.

97 MATHEMATICS AND COMPUTING↗

Computational catalyst discovery: Active classification through myopic multiscale sampling

We report the recent boom in computational chemistry has enabled several projects aimed at discovering useful materials or catalysts. We acknowledge and address two recurring issues in the field of computational catalyst discovery. First, calculating macro-scale catalyst properties is not straightforward when using ensembles of atomic-scale calculations [e.g., density functional theory (DFT)]. We attempt to address this issue by creating a multi-scale model that estimates bulk catalyst activity using adsorption energy predictions from both DFT and machine learning models. The second issue is that many catalyst discovery efforts seek to optimize catalyst properties, but optimization is an inherently exploitative objective that is in tension with the explorative nature of early-stage discovery projects. In other words, why invest so much time finding a “best” catalyst when it is likely to fail for some other, unforeseen problem? We address this issue by relaxing the catalyst discovery goal into a classification problem: “What is the set of catalysts that is worth testing experimentally?” Here, we present a catalyst discovery method called myopic multiscale sampling, which combines multiscale modeling with automated selection of DFT calculations. It is an active classification strategy that seeks to classify catalysts as “worth investigating” or “not worth investigating” experimentally. Our results show an ~7–16 times speedup in catalyst classification relative to random sampling. These results were based on offline simulations of our algorithm on two different datasets: a larger, synthesized dataset and a smaller, real dataset.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

New Capabilities for Sampling Tools

This report will discuss new capabilities that have been added to the sample.py and uniform_sampler.py codes. The codes have been updated to allow for both log-uniform sampling and categorical variables. The categorical variables do not have to be numeric. The different values for a categorical variable are specified in a limits file by having spaces between them. The difference between sample.py and uniform_sampler.py is that sample.py is for generating random samples and uniform_sampler.py is for generating samples or points on a fixed grid. After sourcing a file to set the environment, execute the codes with the commands: sample.py and uniform sampler.py .

97 MATHEMATICS AND COMPUTING↗

Hybrid-BPR (Bayesian Personalized Ranking with Feature Embeddings and Explicit Negative Sampling) [SWR-26-039]

Hybrid-BPR is a Python library for Bayesian Personalized Ranking (BPR) with two key capabilities that go beyond standard BPR implementations: 1. User and item feature embeddings - incorporate content-based signals (genres, tags, metadata) alongside collaborative filtering. 2. Implicit negative interactions - use observed non-interactions (e.g. viewed-but-not-clicked) as negative training signal instead of random sampling from the full item space. The software is built for recommender systems research with MLflow experiment tracking, parallel hyperparameter sweeps, and standard ranking metrics.

Sandhu, Rimple [National Laboratory of the Rockies↗

Boundary-Aware Adversarial Learning Domain Adaption and Active Learning for Cross-Sensor Building Extraction

The use of convolutional neural networks (CNNs) for building extraction from remote sensing images has been widely studied and many public datasets have been made available for accelerating development of these CNN models. Yet adapting pretrained models at scale in real-world scenarios remains a challenging task. The main barrier is that certain new labels are still needed to compensate for domain shifting between the labeled data and new images that potentially cover new geographic locations or that are from a different sensor. In this article, we propose to add informatively labeled samples from a new image pool under the paradigm of active learning. To select the most useful samples based on model uncertainty, we first tackle the problem of uncalibrated uncertainty estimation due to distribution shifting by adapting feature extractors with boundary-based adversarial learning. Calibrated uncertainty is used as the query criterion in the active learning process, where the most uncertain samples are selected for annotation and included for model retraining. The proposed workflow was tested with three data pairs in which each workflow represents a scenario often encountered in real-world applications, including adapting pretrained models to new images collected with different sensors or to new geographic areas where appearances and types of buildings are very different. Compared to several baselines, including random sampling, temperature scaling (a well-known uncertainty calibration technique), different query strategies, and active domain adaptation methods, the proposed workflow shows that strategically querying a smaller set of samples for labeling achieves comparable or better building extraction performance. The proposed method reduces the number of labeled samples required to achieve sufficient model accuracy, thus significantly reducing hundreds of person-hours for labeled data creation. In addition, we include a few considerations when deploying this workflow in a GPU cluster that can be easily adapted to achieve operational building extraction model retraining.

97 MATHEMATICS AND COMPUTING↗

Uncertainty analysis for VERA problem 2 using the cell-code Condor v2.8.05

Condor is a cell-level neutronic calculation code that applies multi-group collision probabilities with heterogeneous response coupling method within generic geometry configurations. Under the Condor's code continuous development, the incorporation of up-to-date methodologies and state-of-the-art practices in reactor analysis represents a driving force. In this work, the capabilities of Condor v2.8.05 to develop an uncertainty analysis for realistic PWR-kind fuel assemblies are studied. The Total Monte Carlo approach is applied to quantify the impact of fabrication tolerances in the code's results for reactivity and power distributions, by means of randomly sampled input values using the VERA problem 2 as basis. The VERA problem 2 proposes a series of Westinghouse 2D 17 x 17-type fuel lattices, to be calculated reflected at beginning-of-life without Xe. The configurations correspond to a modern PWR. Selected neutronic parameters from Condor runs are thus analyzed in terms of the observed spread as well as the obtained distributions for the randomly perturbed cases, showing the capability of the code to handle the required input data, as well as its ability to provide valuable insights regarding uncertainty quantification.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Detection of significant antiviral drug effects on COVID-19 with reasonable sample sizes in randomized controlled trials: A modeling study

Development of an effective antiviral drug for Coronavirus Disease 2019 (COVID-19) is a global health priority. Although several candidate drugs have been identified through in vitro and in vivo models, consistent and compelling evidence from clinical studies is limited. The lack of evidence from clinical trials may stem in part from the imperfect design of the trials. We investigated how clinical trials for antivirals need to be designed, especially focusing on the sample size in randomized controlled trials. A modeling study was conducted to help understand the reasons behind inconsistent clinical trial findings and to design better clinical trials. We first analyzed longitudinal viral load data for Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) without antiviral treatment by use of a within-host virus dynamics model. The fitted viral load was categorized into 3 different groups by a clustering approach. Comparison of the estimated parameters showed that the 3 distinct groups were characterized by different virus decay rates (p-value < 0.001). The mean decay rates were 1.17 d -1 (95% CI: 1.06 to 1.27 d -1 ), 0.777 d -1 (0.716 to 0.838 d -1 ), and 0.450 d -1 (0.378 to 0.522 d -1 ) for the 3 groups, respectively. Such heterogeneity in virus dynamics could be a confounding variable if it is associated with treatment allocation in compassionate use programs (i.e., observational studies). Subsequently, we mimicked randomized controlled trials of antivirals by simulation. An antiviral effect causing a 95% to 99% reduction in viral replication was added to the model. To be realistic, we assumed that randomization and treatment are initiated with some time lag after symptom onset. Using the duration of virus shedding as an outcome, the sample size to detect a statistically significant mean difference between the treatment and placebo groups (1:1 allocation) was 13,603 and 11,670 (when the antiviral effect was 95% and 99%, respectively) per group if all patients are enrolled regardless of timing of randomization. The sample size was reduced to 584 and 458 (when the antiviral effect was 95% and 99%, respectively) if only patients who are treated within 1 day of symptom onset are enrolled. We confirmed the sample size was similarly reduced when using cumulative viral load in log scale as an outcome. We used a conventional virus dynamics model, which may not fully reflect the detailed mechanisms of viral dynamics of SARS-CoV-2. The model needs to be calibrated in terms of both parameter settings and model structure, which would yield more reliable sample size calculation. In this study, we found that estimated association in observational studies can be biased due to large heterogeneity in viral dynamics among infected individuals, and statistically significant effect in randomized controlled trials may be difficult to be detected due to small sample size. The sample size can be dramatically reduced by recruiting patients immediately after developing symptoms. We believe this is the first study investigated the study design of clinical trials for antiviral treatment using the viral dynamics model.

60 APPLIED LIFE SCIENCES↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Cross-national analysis of food security drivers: comparing results based on the Food Insecurity Experience Scale and Global Food Security Index

Abstract The second UN Sustainable Development Goal establishes food security as a priority for governments, multilateral organizations, and NGOs. These institutions track national-level food security performance with an array of metrics and weigh intervention options considering the leverage of many possible drivers. We studied the relationships between several candidate drivers and two response variables based on prominent measures of national food security: the 2019 Global Food Security Index (GFSI) and the Food Insecurity Experience Scale’s (FIES) estimate of the percentage of a nation’s population experiencing food security or mild food insecurity (FI ). We compared the contributions of explanatory variables in regressions predicting both response variables, and we further tested the stability of our results to changes in explanatory variable selection and in the countries included in regression model training and testing. At the cross-national level, the quantity and quality of a nation’s agricultural land were not predictive of either food security metric. We found mixed evidence that per-capita cereal production, per-hectare cereal yield, an aggregate governance metric, logistics performance, and extent of paid employment work were predictive of national food security. Household spending as measured by per-capita final consumption expenditure (HFCE) was consistently the strongest driver among those studied, alone explaining a median of 92% and 70% of variation (based on out-of-sample R 2 ) in GFSI and FI , respectively. The relative strength of HFCE as a predictor was observed for both response variables and was independent of the countries used for model training, the transformations applied to the explanatory variables prior to model training, and the variable selection technique used to specify multivariate regressions. The results of this cross-national analysis reinforce previous research supportive of a causal mechanism where, in the absence of exceptional local factors, an increase in income drives increase in food security. However, the strength of this effect varies depending on the countries included in regression model fitting. We demonstrate that using multiple response metrics, repeated random sampling of input data, and iterative variable selection facilitates a convergence of evidence approach to analyzing food security drivers.

42 ENGINEERING↗

Megadrought: A Series of Unfortunate La Niña Events?

Megadroughts are multidecadal periods of aridity more persistent than most droughts during the instrumental period. Paleoclimate evidence suggests that megadroughts occur in many parts of the world, including North America, Central America, western Europe, eastern Asia, and northern Africa. It remains unclear to what extent such megadroughts require external forcing or whether they can arise from internal climate variability alone. A novel statistical–dynamical approach is used to evaluate the possibility that such events arise solely as a function of interannual tropical sea surface temperature (SST) variations. A statistical emulator of tropical SST variations is constructed by using an empirical moving-blocks bootstrap approach that randomly samples multiyear sequences of the observational SST record. This approach preserves the power spectrum, seasonal cycle, and spatial pattern of El Niño-Southern Oscillation (ENSO) but removes longer timescale fluctuations embedded in the observational record. These resampled SST anomalies are then used to force an atmospheric model (the Community Atmosphere Model Version 5). As megadroughts emerge in this run, they should, therefore, be solely a consequence of La Niña sequences combined with internal atmospheric variability and persistence driven by soil moisture storage and other land-surface processes. We indeed find that megadroughts in this simulation have an amplitude-duration rate that is generally indistinguishable from the rate documented in paleoclimate records of the western United States. Our findings support the idea that megadroughts may occur randomly when the unforced climate system evolves freely over a sufficiently long period of time, implying that an unforced unusual but statistically plausible series of La Niña events may be sufficient to generate megadrought.

54 ENVIRONMENTAL SCIENCES↗

Statistics of base polytopes in F-theory

We propose a new statistical ensemble of toric bases for elliptic Calabi-Yaus used in F-theory models, by focusing on only the convex hull of the base, i.e., the base polytope. This physically motivated coarse-graining greatly simplifies the combinatorial complexity of the part of the 4d F-theory landscape with toric bases. We develop a Monte Carlo approach that randomly samples the base polytopes within fixed boxes, with proper statistical weights. We first apply the algorithm to the set of 2d base polytopes, generating an enlarged set of toric 2d bases that include certain types of codimension-two (4,6) points, and we validate our approach against exact numbers. We then explore the set of 3d base polytopes which fit in a set of “maximal” 3d boxes, and estimate the total number of inequivalent 3d base polytopes to be 10 85 –10 90 . We provide statistical data such as the distribution of non-Higgsable gauge groups on these bases. Amusingly, a similar method can also be applied to generate reflexive polytopes in various dimensions. In both the reflexive and base polytope cases, the number of relevant polytopes obeys a Gaussian distribution as a function of the number of vertices, which can be understood in terms of other results on random polytopes in the math literature.

Differential and algebraic geometry↗

Real classical shadows

Efficiently learning expectation values of a quantum state using classical shadow tomography has become a fundamental task in quantum information theory. In a classical shadows protocol, one measures a state in a chosen basis $\mathcal{W}$ after it has evolved under a unitary transformation randomly sampled from a chosen distribution $\mathcal{U}$. In this work we study the case where $\mathcal{U}$ corresponds to either local or global orthogonal Clifford gates, and $\mathcal{W}$ consists of real-valued vectors. Our results show that for various situations of interest, this ‘real’ classical shadow protocol improves the sample complexity over the standard scheme based on general Clifford unitaries. For example, when one is interested in estimating the expectation values of arbitrary real-valued observables, global orthogonal Cliffords typically decrease the required number of samples by a factor of two. More dramatically, for k-local observables composed only of real-valued Pauli operators, sampling local orthogonal Cliffords leads to a reduction by an exponential-in-k factor in the sample complexity over local unitary Cliffords. Finally, we show that by measuring in a basis containing complex-valued vectors, orthogonal shadows can, in the limit of large system size, exactly reproduce the original unitary shadows protocol.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Similarity Downselection: Finding the n Most Dissimilar Molecular Conformers for Reference-Free Metabolomics

Computational methods for creating in silico libraries of molecular descriptors (e.g., collision cross sections) are becoming increasingly prevalent due to the limited number of authentic reference materials available for traditional library building. These so-called “reference-free metabolomics” methods require sampling sets of molecular conformers in order to produce high accuracy property predictions. Due to the computational cost of the subsequent calculations for each conformer, there is a need to sample the most relevant subset and avoid repeating calculations on conformers that are nearly identical. The goal of this study is to introduce a heuristic method of finding the most dissimilar conformers from a larger population in order to help speed up reference-free calculation methods and maintain a high property prediction accuracy. Finding the set of the n items most dissimilar from each other out of a larger population becomes increasingly difficult and computationally expensive as either n or the population size grows large. Because there exists a pairwise relationship between each item and all other items in the population, finding the set of the n most dissimilar items is different than simply sorting an array of numbers. For instance, if you have a set of the most dissimilar n = 4 items, one or more of the items from n = 4 might not be in the set n = 5. An exact solution would have to search all possible combinations of size n in the population exhaustively. We present an open-source software called similarity downselection (SDS), written in Python and freely available on GitHub. SDS implements a heuristic algorithm for quickly finding the approximate set(s) of the n most dissimilar items. We benchmark SDS against a Monte Carlo method, which attempts to find the exact solution through repeated random sampling. We show that for SDS to find the set of n most dissimilar conformers, our method is not only orders of magnitude faster, but it is also more accurate than running Monte Carlo for 1,000,000 iterations, each searching for set sizes n = 3–7 out of a population of 50,000. We also benchmark SDS against the exact solution for example small populations, showing that SDS produces a solution close to the exact solution in these instances. Using theoretical approaches, we also demonstrate the constraints of the greedy algorithm and its efficacy as a ratio to the exact solution.

97 MATHEMATICS AND COMPUTING↗