Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Randomized methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

GANISP (GAN-assisted Importance SPlitting)

Genealogical importance splitting marches towards a rare event by iteratively selecting and replicating realizations that are headed towards a rare event. The replication step is made difficult when applied to deterministic systems as the initial conditions of the offspring realizations need to be adjusted. Typically, a random perturbation is applied to the offspring. For some cases, this cloning technique may not be adequate and prevent variance reduction in the probability estimate. A GAN-based replication process is proposed to address this limitation. The perturbations applied to the clones are physically consistent instead of being randomly chosen. The proposed method allows reducing the variance in the probability estimation.

Hassanaly, Malik↗

Transcripts and genomic intervals associated with variation in metabolite abundance in maize leaves under field conditions

Abstract Plants exhibit extensive environment-dependent intraspecific metabolic variation, which likely plays a role in determining variation in whole plant phenotypes. However, much of the work seeking to use natural variation to link genes and transcript’s impacts on plant metabolism has employed data from controlled environments. Here, we generated and analyzed data on the variation in the abundance of 26 metabolites across 660 maize inbred lines under field conditions. We employ these data and previously published transcript and whole plant phenotype data reported for the same field experiment to identify both genomic intervals (through genome-wide association studies (GWAS)) and transcripts (using both transcriptome-wide association studies (TWAS) and an explainable artificial intelligence (AI) approach based on random forest (RF)) associated with variation in metabolite abundance. Both genome-wide association and random forest-based methods identified substantial numbers of significant associations including genes with plausible links to the metabolites they are associated with. In contrast, the transcriptome-wide association identified only six significant associations. In three cases, genetic markers associated with metabolic variation in our study colocalized with markers linked to variation in non-metabolic traits scored in the same experiment. We speculate that the poor performance of transcriptome-wide association studies in identifying transcript-metabolite associations may reflect a high prevalence of non-linear interactions between transcripts and metabolites and/or a bias towards rare transcripts playing a large role in determining intraspecific metabolic variation.

Mathivanan, Ramesh Kanna↗

A Comparison of Time Series Gap-Filling Methods to Impute Solar Radiation Data

Complete solar resource datasets play a critical role at every stage of solar project phases. However, measured or modeled solar resource data come with significant uncertainties and usually suffer from several issues, including but not limited to, data gaps, data quality issue, etc. In order to mitigate these issues an appropriate data imputation method should be implemented to build a complete and reliable temporal (and spatial) database. Being motivated by this, in this study we compare the performances of eight different gap filling methods extensively by creating random and artificial data gaps in (i) hourly irradiance data for one year using a few locations of the National Solar Radiation Database (NSRDB) and (ii) one-minute ground measurement dataset from Surface Radiation Budget Network (SURFRAD) stations.

clearness index↗

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Targeted mutagenesis and high-throughput screening of diversified gene and promoter libraries for isolating gain-of-function mutations

Targeted mutagenesis of a promoter or gene is essential for attaining new functions in microbial and protein engineering efforts. In the burgeoning field of synthetic biology, heterologous genes are expressed in new host organisms. Similarly, natural or designed proteins are mutagenized at targeted positions and screened for gain-of-function mutations. Here, we describe methods to attain complete randomization or controlled mutations in promoters or genes. Combinatorial libraries of one hundred thousands to tens of millions of variants can be created using commercially synthesized oligonucleotides, simply by performing two rounds of polymerase chain reactions. With a suitably engineered reporter in a whole cell, these libraries can be screened rapidly by performing fluorescence-activated cell sorting (FACS). Within a few rounds of positive and negative sorting based on the response from the reporter, the library can rapidly converge to a few optimal or extremely rare variants with desired phenotypes. Library construction, transformation and sequence verification takes 6–9 days and requires only basic molecular biology lab experience. Screening the library by FACS takes 3–5 days and requires training for the specific cytometer used. Further steps after sorting, including colony picking, sequencing, verification, and characterization of individual clones may take longer, depending on number of clones and required experiments.

59 BASIC BIOLOGICAL SCIENCES↗

The PENELOPE Physics Models and Transport Mechanics. Implementation into Geant4

A translation of the PENELOPE physics subroutines to C++, designed as an extension of the GEANT4 toolkit, is presented. The Fortran code system PENELOPE performs Monte Carlo simulation of coupled electron-photon transport in arbitrary materials for a wide energy range, nominally from 50 eV up to 1 GeV. PENELOPE implements the most reliable interaction models that are currently available, limited only by the required generality of the code. In addition, the transport of electrons and positrons is simulated by means of an elaborate class II scheme in which hard interactions (involving deflection angles or energy transfers larger than pre-defined cutoffs) are simulated from the associated restricted differential cross sections. After a brief description of the interaction models adopted for photons and electrons/positrons, we describe the details of the class-II algorithm used for tracking electrons and positrons. The C++ classes are adapted to the specific code structure of GEANT4. They provide a complete description of the interactions and transport mechanics of electrons/positrons and photons in arbitrary materials, which can be activated from the G4ProcessManager to produce simulation results equivalent to those from the original PENELOPE programs. The combined code, named PENG4, benefits from the multi-threading capabilities and advanced geometry and statistical tools of GEANT4.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Comprehensive Investigation of Active Learning Strategies for Conducting Anti-Cancer Drug Screening

It is well-known that cancers of the same histology type can respond differently to a treatment. Thus, computational drug response prediction is of paramount importance for both preclinical drug screening studies and clinical treatment design. To build drug response prediction models, treatment response data need to be generated through screening experiments and used as input to train the prediction models. In this study, we investigate various active learning strategies of selecting experiments to generate response data for the purposes of (1) improving the performance of drug response prediction models built on the data and (2) identifying effective treatments. Here, we focus on constructing drug-specific response prediction models for cancer cell lines. Various approaches have been designed and applied to select cell lines for screening, including a random, greedy, uncertainty, diversity, combination of greedy and uncertainty, sampling-based hybrid, and iteration-based hybrid approach. All of these approaches are evaluated and compared using two criteria: (1) the number of identified hits that are selected experiments validated to be responsive, and (2) the performance of the response prediction model trained on the data of selected experiments. The analysis was conducted for 57 drugs and the results show a significant improvement on identifying hits using active learning approaches compared with the random and greedy sampling method. Active learning approaches also show an improvement on response prediction performance for some of the drugs and analysis runs compared with the greedy sampling method.

60 APPLIED LIFE SCIENCES↗

Hyperuniform and nearly hyperuniform random network materials

This invention is in the field of physical chemistry and relates to novel hyperuniform and nearly hyperuniform random network materials and methods of making said materials. Methods are described for controlling or altering the band gap of a material, and in particular commercially useful materials such as amorphous silicon. These methods can be exploited in the design of semiconductors, transistors, diodes, solar cells and the like.

36 MATERIALS SCIENCE↗

Frontal Slice Approaches for Tensor Linear Systems

Inspired by the row and column action methods for solving large-scale linear systems, in this work, we explore the use of frontal slices for solving tensor linear systems. In particular, this paper presents a novel approach for using frontal slices of a tensor $\mathcal{A}$ to solve tensor linear systems $\mathcal{A} ∗\mathcal{X} = \mathcal{B}$ where ∗ denotes the $t$-product. In addition, we consider variations of this method, including cyclic, block, and randomized approaches, each designed to optimize performance in different operational contexts. Our primary contribution lies in the development and convergence analysis of these methods. Experimental results on synthetically generated and real-world data, including applications such as image and video deblurring, demonstrate the efficacy of our proposed approaches and validate our theoretical findings.

Luo, Hengrui↗

Discrete-Direct Model Calibration and Uncertainty Propagation Method Confirmed on Multi-Parameter Plasticity Model Calibrated to Sparse Random Field Data

A discrete direct (DD) model calibration and uncertainty propagation approach is explained and demonstrated on a 4-parameter Johnson-Cook (J-C) strain-rate dependent material strength model for an aluminum alloy. The methodology's performance is characterized in many trials involving four random realizations of strain-rate dependent material-test data curves per trial, drawn from a large synthetic population. The J-C model is calibrated to particular combinations of the data curves to obtain calibration parameter sets which are then propagated to “Can Crush” structural model predictions to produce samples of predicted response variability. These are processed with appropriate sparse-sample uncertainty quantification (UQ) methods to estimate various statistics of response with an appropriate level of conservatism. This is tested on 16 output quantities (von Mises stresses and equivalent plastic strains) and it is shown that important statistics of the true variabilities of the 16 quantities are bounded with a high success rate that is reasonably predictable and controllable. The DD approach has several advantages over other calibration-UQ approaches like Bayesian inference for capturing and utilizing the information obtained from typically small numbers of replicate experiments in model calibration situations—especially when sparse replicate functional data are involved like force–displacement curves from material tests. The DD methodology is straightforward and efficient for calibration and propagation problems involving aleatory and epistemic uncertainties in calibration experiments, models, and procedures.

42 ENGINEERING↗

Diabetes-specific formula with standard of care improves glycemic control, body composition, and cardiometabolic risk factors in overweight and obese adults with type 2 diabetes: results from a randomized controlled trial

Background and aims Medical nutrition therapy is important for diabetes management. This randomized controlled trial investigated the effects of a diabetes-specific formula (DSF) on glycemic control and cardiometabolic risk factors in adults with type 2 diabetes (T2D). Methods Participants ( n = 235) were randomized to either DSF with standard of care (SOC) (DSF group; n = 117) or SOC only (control group; n = 118). The DSF group consumed one or two DSF servings daily as meal replacement or partial meal replacement. The assessments were done at baseline, on day 45, and on day 90. Results There were significant reductions in glycated hemoglobin (−0.44% vs. –0.26%, p = 0.015, at day 45; −0.50% vs. −0.21%, p = 0.002, at day 90) and fasting blood glucose (−0.14 mmol/L vs. +0.32 mmol/L, p = 0.036, at day 90), as well as twofold greater weight loss (−1.30 kg vs. –0.61 kg, p < 0.001, at day 45; −1.74 kg vs. –0.76 kg, p < 0.001, at day 90) in the DSF group compared with the control group. The decrease in percent body fat and increase in percent fat-free mass at day 90 in the DSF group were almost twice that of the control group (1.44% vs. 0.79%, p = 0.047). In addition, the percent change in visceral adipose tissue at day 90 in the DSF group was several-fold lower than in the control group (−6.52% vs. –0.95%, p < 0.001). The DSF group also showed smaller waist and hip circumferences, and lower diastolic blood pressure than the control group (all overall p ≤ 0.045). Conclusion DSF with SOC yielded significantly greater improvements than only SOC in glycemic control, body composition, and cardiometabolic risk factors in adults with T2D.

Tey, Siew Ling↗

Derivative-free stochastic optimization via adaptive sampling strategies

In this paper, we present a novel derivative-free framework for solving unconstrained stochastic optimization problems. Many problems in fields ranging from simulation optimization to reinforcement learning to quantum computing involve settings where only stochastic function values are obtained via a zeroth-order oracle, which has no available gradient information and necessitates the usage of derivative-free optimization methodologies. Our approach includes estimating gradients using stochastic function evaluations and integrating adaptive sampling techniques to control the accuracy in these stochastic approximations. Our framework encapsulates several gradient estimation techniques, including standard finite-difference, Gaussian smoothing, sphere smoothing, randomized coordinate finite-difference, and randomized subspace finite-difference methods. We provide theoretical convergence guarantees for our framework and analyze the worst-case iteration and sample complexities associated with each gradient estimation method. Finally, we demonstrate the empirical performance of the methods on logistic regression and nonlinear least squares problems.

Adaptive sampling↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

CORRLA-RS

The CORRLA-RS package provides a suite of statistical methods for sampling multidimensional distributions and to conduct sensitivity and correlation analysis of large scale data in the Rust programming language. The software provides a unique solution to multidimensional constrained sampling problems utilizing a combination of parallelized Markov Chain Monte Carlo methods and traditional rejection sampling. The sensitivity and correlation analysis methods are backed by a high performance randomized singular value decomposition implementation which enables datasets larger than the random access memory (RAM) size to be analyzed. Additionally, CORRLA-RS implements the active subspace identification method using a KD-Tree and the randomized singular value decomposition acting in concert.

Gurecky, William [Oak Ridge National Laboratory (O↗

Mixed-State Entanglement from Local Randomized Measurements

In this work, we propose a method for detecting bipartite entanglement in a many-body mixed state based on estimating moments of the partially transposed density matrix. The estimates are obtained by performing local random measurements on the state, followed by postprocessing using the classical shadows framework. Our method can be applied to any quantum system with single-qubit control. We provide a detailed analysis of the required number of experimental runs, and demonstrate the protocol using existing experimental data [Brydges et al., Science 364, 260 (2019)].

1-dimensional spin chains↗

SpotSDC: Revealing the Silent Data Corruption Propagation in High-Performance Computing Systems

We report the trend of rapid technology scaling is expected to make the hardware of high-performance computing (HPC) systems more susceptible to computational errors due to random bit flips. Some bit flips may cause a program to crash or have a minimal effect on the output, but others may lead to silent data corruption (SDC), i.e., undetected yet significant output errors. Classical fault injection analysis methods employ uniform sampling of random bit flips during program execution to derive a statistical resiliency profile. However, summarizing such fault injection result with sufficient detail is difficult, and understanding the behavior of the fault-corrupted program is still a challenge. In this article, we introduce SpotSDC, a visualization system to facilitate the analysis of a program's resilience to SDC. SpotSDC provides multiple perspectives at various levels of detail of the impact on the output relative to where in the source code the flipped bit occurs, which bit is flipped, and when during the execution it happens. SpotSDC also enables users to study the code protection and provide new insights to understand the behavior of a fault-injected program. Based on lessons learned, we demonstrate how what we found can improve the fault injection campaign method.

97 MATHEMATICS AND COMPUTING↗

Accelerating Ab Initio Simulation via Nested Monte Carlo and Machine Learned Reference Potentials

As a corollary of the rapid advances in computing, ab initio simulation is playing an increasingly important role in modeling materials at the atomic scale. Two strategies are possible, ab initio Monte Carlo (AIMC) and molecular dynamics (AIMD) simulation. The former benefits from exact sampling from the correct thermodynamic distribution, while the latter is typically more efficient with its collective all-atom coordinate updates. In this study, using a relatively simple test model comprised of helium and argon, we show that AIMC can be brought up to, and even above, the performance levels of AIMD via a hybrid nested sampling/machine learning (ML) strategy. Here, ML provides an accurate classical reference potential (up to three-body explicit interactions) that can pilot long collective Monte Carlo moves that are accepted or rejected in toto à la nested Monte Carlo (NMC); this is in contrast to the single move nature of a naive implementation. Our proposed method only requires a small up front expense from evaluating the ab initio energies and forces of (100) random configurations for training. Importantly, our method does not totally rely on the trained, assuredly imperfect, interaction. We show that high performance and exact sampling at the desired level of theory can be realized even when the trained interaction has appreciable differences from the ab initio potential. Remarkably, at the highest levels of performance realized via our approach, a pair of statistically uncorrelated atomic configurations can be generated with (1) ab initio calculations.

74 ATOMIC AND MOLECULAR PHYSICS↗

Preliminary Study on TRISO Fuel Cross Section Generation

Cross section self-shielding methodologies for TRISO fuel were assessed to provide accurate multigroup cross sections for a high-fidelity reactor physics code so that the code is able to accurately model and simulate advanced reactors with TRISO fuel. Initially, the two existing methodologies (the SCALE method and the Sanchez-Pomraning method) were studied and implemented to MC2-3 for detailed performance tests. Additionally, a new spatial self-shielding method, named the iterative local spatial self-shielding (ILSS) method, for particulate fuels was developed based on the disadvantage factor and implemented to MC2-3 as well. The new method approximately accounts for the effect of randomly distributed particles on the particle shadowing effect using a homogenized compact region surrounding a particle of interest at the center. The self-shielded cross sections of the particle at the center are determined iteratively since the cross sections of the homogenized compact region are calculated using them. For the energy range above 100 keV where the fuel-to-moderator ratio is more important than the random distribution of particles, a single particle unit-cell model is used by preserving the average amount of moderator per fuel particle in the system. The three self-shielding methods implemented in MC2-3 were tested using numerical benchmark problems made based on fuel compact problems of a prismatic-type very high temperature reactor. Test results indicated that the ILSS method produced slightly better results than the SCALE and Sanchez-Pomraning methods, compared to the Serpent-2 Monte Carlo results obtained with 25 independent random particle configurations. The SCALE and Sanchez-Pomraning methods tend to underestimate the heterogeneity effect by 150 and 100 pcm, respectively, while the new ILSS method overestimates the heterogeneity effect by 70 pcm. In future, the new self-shielding method will be extended to perform pebble calculations and compare results with those from the SCALE and Sanchez-Pomraning methods. Furthermore, the new method will be optimized for practical applications to on-the-fly resonance treatment for lattice or whole-core calculations for advanced reactors.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗