Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Randomized methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Treating random sequential addition via the replica method

While many physical processes are non-equilibrium in nature, the theory and modeling of such phenomena lag behind theoretical treatments of equilibrium systems. The diversity of powerful theoretical tools available to describe equilibrium systems has inspired strategies that map non-equilibrium systems onto equivalent equilibrium analogs so that interrogation with standard statistical mechanical approaches is possible. In this work, we revisit the mapping from the non-equilibrium random sequential addition process onto an equilibrium multi-component mixture via the replica method, allowing for theoretical predictions of non-equilibrium structural quantities. We validate the above approach by comparing the theoretical predictions to numerical simulations of random sequential addition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using porous random fields to predict the elastic modulus of unoxidized and oxidized superfine graphite

Nuclear graphite is a candidate material for Generation IV nuclear power plants. Porous materials such as graphite can contain complex networks of pores that influence the material's mechanical and irradiation response. A methodology known as the random finite element method (RFEM) was adapted to create synthetic microstructures and predict the influence of porosity on the elastic properties of graphite during oxidation. RFEM combines random field theory and the finite element method in a Monte Carlo framework to estimate the mechanical response of a given grade of graphite. In this research, the random fields were verified through experimental characterization to predict the elastic response of three nuclear graphite grades, ETU-10, IG-110, and 2114. Finite element models (FEM) were generated using segmentations of x-ray computed tomography (XCT) data known as image-based models (IBMs) to validate and compare with the RFEM results and better understand the effects of uniform oxidation in these graphite grades. The RFEM predictions appear to correlate well with the experimental values of the measured Young’s modulus of the three graphite grades and display the same trends as IBMs.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A randomized multiplex CRISPRi-Seq approach for the identification of critical combinations of genes

Identifying virulence-critical genes from pathogens is often limited by functional redundancy. To rapidly interrogate the contributions of combinations of genes to a biological outcome, we have developed a multiplex, randomized CRISPR interference sequencing (MuRCiS) approach. At its center is a new method for the randomized self-assembly of CRISPR arrays from synthetic oligonucleotide pairs. When paired with PacBio long-read sequencing, MuRCiS allowed for near-comprehensive interrogation of all pairwise combinations of a group of 44 Legionella pneumophila virulence genes encoding highly conserved transmembrane proteins for their role in pathogenesis. Both amoeba and human macrophages were challenged with L. pneumophila bearing the pooled CRISPR array libraries, leading to the identification of several new virulence-critical combinations of genes. lpg2888 and lpg3000 were particularly fascinating for their apparent redundant functions during L. pneumophila human macrophage infection, while lpg3000 alone was essential for L. pneumophila virulence in the amoeban host Acanthamoeba castellanii. Thus, MuRCiS provides a method for rapid genetic examination of even large groups of redundant genes, setting the stage for application of this technology to a variety of biological contexts and organisms.

79 ASTRONOMY AND ASTROPHYSICS↗

Entropy-driven Optimal Sub-sampling of Fluid Dynamics for Developing Machine-learned Surrogates

Optimal sub-sampling of large datasets from fluid dynamics simulations is essential for training reduced-order machine learned models. A method using Shannon entropy was developed to weight flow features according to their level of information content, such that the most informative features can be extracted and used for training a surrogate model. The method is demonstrated in the canonical flow over a cylinder problem simulated with OpenFOAM. Both time-independent predictions and temporal forecasting were investigated as well as two types of prediction targets: local per-grid-point predictions and global per-time-step predictions. When tested on training a surrogate model, results indicate that our entropy-based sampling method typically outperforms random sampling and yields more reproducible results in less iterations. Finally, the method was used to train a surrogate model for modeling turbulence in magnetohydrodynamic flows, which revealed various challenges and opportunities for future research.

Brewer, Wes↗

Accelerating multicanonical sampling with irreversibility

Flat-histogram Monte Carlo simulations are well-established, robust methods to perform random walks in a physical observable or parameter space, making them suitable for finding ground states or studying phase transitions in complex systems in statistical physics. However, their efficiency can be limited by the time to attain the desired flat distribution, which is generally unknown prior to the simulations. In particular, they might suffer from slowing down towards the end of a simulation due to the diffusive nature of random walks. In this work we apply irreversibility to the multicanonical Monte Carlo method via the lifting approach to alleviate this behavior. We achieve a 2–4 times speedup in ground-state search for a two-dimensional (2D) Ising model, and up to an order of magnitude of speedup for finding the ground-state energy in an Edwards–Anderson spin glass, compared to traditional multicanonical sampling. In conclusion, the round-trip times between ground states show a narrower distribution and are significantly shorter compared to the reversible counterpart, suggesting that a lower convergence time with a smaller time variance is feasible.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Intrepid MCMC: Metropolis-Hastings with exploration

In engineering examples, one often encounters the need to sample from unnormalized distributions with complex shapes that may also be implicitly defined through a physical or numerical simulation model, making it computationally expensive to evaluate the associated density function. For such cases, MCMC has proven to be an invaluable tool. Random-walk Metropolis Methods (also known as Metropolis-Hastings (MH)), in particular, are highly popular for their simplicity, flexibility, and ease of implementation. However, most MH algorithms suffer from significant limitations when attempting to sample from distributions with multiple modes (particularly disconnected ones). Here, in this paper, we present Intrepid MCMC - a novel MH scheme that utilizes a simple coordinate transformation to significantly improve the mode-finding ability and convergence rate to the target distribution of random-walk Markov chains while retaining most of the simplicity of the vanilla MH paradigm. Through multiple examples, we showcase the improvement in the performance of Intrepid MCMC over vanilla MH for a wide variety of target distribution shapes. We also provide an analysis of the mixing behavior of the Intrepid Markov chain, as well as the efficiency of our algorithm for increasing dimensions. A thorough discussion is presented on the practical implementation of the Intrepid MCMC algorithm. Finally, its utility is highlighted through a Bayesian parameter inference problem for a two-degree-of-freedom oscillator under free vibration.

97 - MATHEMATICS AND COMPUTING↗

Randomized Adiabatic Quantum Linear Solver Algorithm with Optimal Complexity Scaling and Detailed Running Costs

Solving linear systems of equations is a fundamental problem with a wide variety of applications across many fields of science, and there is increasing effort to develop quantum linear solver algorithms. Subaşı et al. [Phys. Rev. Lett. 122, 060504 (2019)] proposed a randomized algorithm inspired by adiabatic quantum computing, based on a sequence of random Hamiltonian simulation steps, with suboptimal scaling in the condition number 𝜅 of the linear system and the target error 𝜖. Here we go beyond these results in several ways. Firstly, using filtering [Lin and Tong, Quantum 4, 361 (2020)] and Poissonization techniques [Cunningham and Roland, ArXiv:2406.03972 (2024)], the algorithm complexity is improved to the optimal scaling 𝑂⁡(𝜅⁢log (1/𝜖))—an exponential improvement in 𝜖, and a shaving of a log 𝜅 scaling factor in 𝜅. Secondly, the algorithm is further modified to achieve constant factor improvements, which are vital as we progress towards hardware implementations on fault-tolerant devices. We introduce a cheaper randomized walk operator method replacing Hamiltonian simulation—which also removes the need for potentially challenging classical precomputations; randomized routines are sampled over optimized random variables; circuit constructions are improved. We obtain a closed formula rigorously upper bounding the expected number of times one needs to apply a block-encoding of the linear system matrix to output a quantum state encoding the solution to the linear system. The upper bound is 837⁢𝜅 at 𝜖 = 10 −10 for Hermitian matrices.

97 MATHEMATICS AND COMPUTING↗

Delayed Critical and Subcritical Experiments with Polyethylene Moderated Unreflected Thin 15 in. Diameter HEU Metal Plates

The thin ~15 in. diameter highly enriched uranium (HEU) metal plates were assembled to delayed criticality at the Oak Ridge Critical Experiments Facility (ORCEF) in 1969 with various thicknesses of polyethylene (varying from 1/16 to 2$\frac{3}{8}$ inches) between uranium metal plates. The average 235 U enrichment was 93.27 wt. %. These unreflected critical configurations contained 4$\frac{2}{3}$ to 20$\frac{5}{6}$ thin 15 in. diameter HEU metal plates (on loan from Los Alamos National Laboratory [LANL] and shipped to Oak Ridge National Laboratory [ORNL] on June 3, 1969). Depending on the thickness of polyethylene, the enriched uranium masses varying from 28,053 to 135,148 grams. Fractional plate sections consisted of the appropriate number of 60° pie sections. In addition to the measurement at delayed criticality, subcritical measurements were also performed by the inverse kinetic rod drop method. Prompt neutron decay constant measurements were also performed by the Rossi alpha and randomly pulsed neutron method using a time-tagged spontaneous fission californium neutron source; these are briefly reported here. At the time of these measurements in 1969, the thin HEU metal plates were in near-pristine condition with extremely little oxidation, allowing better descriptions of the uranium plates than the use of these plates in a heavily oxidized and deteriorated condition in recent reflected benchmark experiments at the LANL facility at the Nevada Test Site with these same thin highly enriched uranium metal plates. This report documents the experimental information for the measurements performed so that later researchers can perform the required uncertainty and calculational analyses and documentation to use these data for an International Nuclear Criticality Safety Benchmark Evaluation Program (ICSBEP) or a Nuclear Energy Agency (NEA) benchmark. Data from the experiments described should be acceptable for use as criticality safety benchmark experiments for the ICSBEP and the NEA nuclear criticality safety benchmark program once the uncertainty analysis on the measured neutron multiplication factors is completed. Additional data—such as the dimensional inspection reports, uranium isotopic information, and other relevant particulars—should be retrieved from the Y-12 Plant or LANL and incorporated in the final ICSBEP benchmark. Based on previous ICSBEP benchmarks with this enriched uranium metal at ORCEF, the uncertainties in $k_{eff}$ could be as low as ± 0.0002 for some configurations. Other experiments with smaller-diameter than 15 in. diameter HEU metal plates have been benchmarked in HEU-METFAST-001. The prompt neutron time decay measurements could be the basis for an International Reactor Physics Benchmark Program. Preparation of the present report is part of an effort at ORNL to document more than 15 undocumented critical and subcritical experiments enumerated in ORNL/TM-2019/18 and performed by ORNL at ORCEF and other US Department of Energy critical experiments facilities using more than 500 operational days of critical facility time. This work for this report publication was supported by the Nuclear Criticality, Radiation Transport and Safety NCSP Program at ORNL.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Estimating the Adequacy of a Multi-Objective Optimization

Multi-objective optimization methods can be criticized for lacking a statistically valid measure of the quality and representativeness of a solution. This stance is especially relevant to metaheuristic optimization approaches but can also apply to other methods that typically might only report a small representative subset of a Pareto frontier. Here we present a method to address this deficiency based on random sampling of a solution space to determine, with a specified level of confidence, the fraction of the solution space that is surpassed by an optimization. The Superiority of Multi-Objective Optimization to Random Sampling, or SMORS method, can evaluate quality and representativeness using dominance or other measures, e.g., a spacing measure for high-dimensional spaces. SMORS has been tested in a combinatorial optimization context using a genetic algorithm but could be useful for other optimization methods.

42 ENGINEERING↗

A Machine Learning-based Reliability Evaluation Model for Integrated Power-Gas Systems

This article proposes a hybrid machine learning method for the reliability evaluation of integrated power-gas systems (IPGS) under the uncertain component failure probability distributions. The Random Forest (RF) method is designed to select important features to solve the insufficient quantity of data and the curse of dimensionality problems. The Extreme Gradient Boosting (XGBoost) regression algorithm is developed to quantify the relationship between the uncertain parameters and reliability metrics. Moreover, a ten-fold cross-validation method is employed to further improve the accuracy of the regression model. Simulation results on three test systems show that the proposed method can achieve high accuracy for the reliability evaluation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Building thermal dynamics modeling with deep transfer learning using a large residential smart thermostat dataset

Understanding thermal dynamics and obtaining the computational model of residential buildings enable its scaled application in energy retrofits, control optimization and decarbonization. In this paper, we present a deep learning approach to model building thermal dynamics with smart thermostat data collected from residential buildings, with the goal to investigate model generalizability. In the first stage, we developed and compared different Deep Learning architectures including Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM) models and CNN-LSTM to predict indoor air temperature in a multi-step time horizon. In the second stage, we implemented a Transfer Learning (TL) process, which aims to improve the prediction performance on a new set of buildings (targets), exploiting the knowledge of related or similar buildings (sources). Different TL strategies and source model identification methods were investigated. The study showed that the CNN-LSTM performed the best among the architectures compared, with an average Mean Absolute Error (MAE) of 0.26 °C for one-hour-ahead (twelve 5-min future steps) predictions. Furthermore, the results showed that freezing the LSTM layer and fine-tuning the other layers of the CNN-LSTM achieved the best performance among four TL strategies, which further improved the performance with respect to a machine learning approach by 10%, and proving the effectiveness and generalizability of the proposed approach. A comparison of three different source model identification methods showed that randomly selecting source models constrained by similar building characteristics can provide good TL performance while retaining simplicity comparing with other quantitative source identification methods.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Detecting Anomalous Images in Astronomical Datasets

Abstract Environmental and instrumental conditions can cause anomalies in astronomical images, which can potentially bias all kinds of measurements if not excluded. Detection of the anomalous images is usually done by human eyes, which is slow and sometimes not accurate. This is an important issue in weak lensing studies, particularly in the era of large-scale galaxy surveys, in which image qualities are crucial for the success of galaxy shape measurements. In this work we present two automatic methods for detecting anomalous images in astronomical data sets. The anomalous features can be divided into two types: one is associated with the source images, and the other appears on the background. Our first method, called the entropy method, utilizes the randomness of the orientation distribution of the source shapes and the background gradients to quantify the likelihood of an exposure being anomalous. Our second method involves training a neural network (autoencoder) to detect anomalies. We evaluate the effectiveness of the entropy method on the Canada–France–Hawaii Telescope Lensing Survey (CFHTLenS) and Dark Energy Camera Legacy Survey (DECaLS DR3) data. In CFHTLenS, with 1171 exposures, the entropy method outperforms human inspection by detecting 12 of the 13 anomalous exposures found during human inspection and uncovering 10 new ones. In DECaLS DR3, with 17112 exposures, the entropy method detects a significant number of anomalous exposures while keeping a low false-positive rate. We find that although the neural network performs relatively well in detecting source anomalies, its current performance is not as good as the entropy method.

Astronomy & Astrophysics↗

Calabi-Yau CFTs and random matrices

Using numerical methods for finding Ricci-flat metrics, we explore the spectrum of local operators in two-dimensional conformal field theories defined by sigma models on Calabi-Yau targets at large volume. Focusing on the examples of K3 and the quintic, we show that the spectrum, averaged over a region in complex structure moduli space, possesses the same statistical properties as the Gaussian orthogonal ensemble of random matrix theory.

superstring vacua↗

Six Machine-Learning Methods for Predicting Hospital-Stay Duration for Patients with Sepsis: A Comparative Study

Sepsis is a life-threatening medical condition that, if not treated promptly, can result in tissue damage, organ failure, and death. According to the Centers for Disease Control, about 270,000 individuals die of sepsis in the US each year. Further, sepsis expenditures accounted for 13% of total US hospital costs in 2013, totaling more than $24 billion. Our project objectives were to determine if Machine Learning algorithms could reliably predict hospital stay duration for patients with sepsis. The data set we used has been de-identified and is freely available through the BupaR package. The data includes 1050 cases, 15214 events, and 16 types of actions related to sepsis patient care. First, we used process mining to determine how long each patient was in the hospital. Using BupaR’s functions, we created several process model graphs. These process models depict the movement of patients at a hospital and provide duration data for each patent case. Second, we identified outlier data and created two dataset versions: one with and one without outliers. We then applied the following analysis methods: Linear Regression, Random Forest, K-Nearest Neighbors, Neural Networks, XGBoost, and lightGBM. We compared the model validations for the six machine learning models using the same data-splitting method. We found that the XGBoost model had the best prediction accuracy of 73.9 percent for cases with outliers, and 79 percent for cases without outliers. We also found that the lightGBM model had the lowest mean absolute error between prediction and actual duration in days with 3.66 days for the case with outliers, and 2.4 days for the case without outliers. These two models outperformed the other four models. This work will be enhanced in the future by exploring new prediction algorithms and comparing them with the results of this study.

Chen, Lingtao↗

Optimization of particle tracking methods for stochastic media

Random media emerge in several applications involving particle transport, encompassing e.g. photon propagation through Rayleigh-Taylor instabilities in fuel pellets for inertial confinement fusion, or neutron multiplication problems related to the assessment of re-criticality risk following severe accidents with fuel degradation. Reference calculations in such material configurations by means of Monte Carlo transport codes are particularly challenging, since high-density stochastic media might involve several hundreds of thousands of volumes and thus make particle tracking routines extremely cumbersome. In order to cope with these issues, two distinct strategies have been proposed so far: the use of neighbor maps, or the use of delta tracking. In this work we will compare these methods and illustrate their specific merits and drawbacks, as taken both alone and in combination with each other. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

The impact of detection rate changes and correlations on random-coincidence background measurements

Coincidence detection of multiple particles emitted during an experiment can yield a new depth of understanding of the underlying process under study. However, the probability of detecting particles that are generated from the same physical event within a given coincidence time window is generally much lower than that of detecting particles that appear in the same coincidence time window, but were not created from the same physical event, and are therefore detected randomly in coincidence with each other. Thus, accurate and precise methods of measuring this random-coincidence background are essential for a wide variety of fields of science. A method to determine this background directly using the data themselves without any additional experimental run time or fake signals introduced in the data was recently established (O’Donnell, 2016). This method yields a statistical uncertainty on the random-coincidence background that is orders of magnitude smaller than that of the true coincidence data, though the potential for systematic errors of backgrounds from this method was never explored. In this work, we discuss common varieties of correlated and uncorrelated changes in the detection rates of each particle detected in an experiment. Here we demonstrate here that a correlation between particle detection rates from, for example, an incident particle beam that initiates a physical process of interest, creates systematic errors in the random-coincidence background measurement. We also discuss the impact of a variety of other realistic scenarios for rate changes in experiments. Lastly, a method is introduced to correct for errors in the random-coincidence background from any source, yielding an optimization between statistical precision and eliminating potential lingering systematic errors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Surrogate-Based Autotuning for Randomized Sketching Algorithms in Regression Problems

Algorithms from Randomized Numerical Linear Algebra (RandNLA) are known to be effective in handling high-dimensional computational problems, providing high-quality empirical performance as well as strong probabilistic guarantees. However, their practical application is complicated by the fact that the user needs to set various algorithm-specific tuning parameters which are different from those used in traditional NLA. This paper demonstrates how a surrogate-based autotuning approach can be used to address fundamental problems of parameter selection in RandNLA algorithms. In particular, we provide a detailed investigation of surrogate-based autotuning for sketch-and-precondition (SAP)-based randomized least squares methods, which have been one of the great success stories in modern RandNLA. Empirical results show that our surrogate-based autotuning approach can achieve near-optimal performance with much less tuning cost than a random search (up to about 7.6x fewer trials of different parameter configurations). Moreover, while our experiments focus on least squares, our results demonstrate a general-purpose autotuning pipeline applicable to any kind of RandNLA algorithm.

Cho, Younghyun↗