Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random sampling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Autonomous reinforcement learning agent for stretchable kirigami design of 2D materials

Abstract Mechanical behavior of 2D materials such as MoS 2 can be tuned by the ancient art of kirigami. Experiments and atomistic simulations show that 2D materials can be stretched more than 50% by strategic insertion of cuts. However, designing kirigami structures with desired mechanical properties is highly sensitive to the pattern and location of kirigami cuts. We use reinforcement learning (RL) to generate a wide range of highly stretchable MoS 2 kirigami structures. The RL agent is trained by a small fraction (1.45%) of molecular dynamics simulation data, randomly sampled from a search space of over 4 million candidates for MoS 2 kirigami structures with 6 cuts. After training, the RL agent not only proposes 6-cut kirigami structures that have stretchability above 45%, but also gains mechanistic insight to propose highly stretchable (above 40%) kirigami structures consisting of 8 and 10 cuts from a search space of billion candidates as zero-shot predictions.

36 MATERIALS SCIENCE↗

Automated Construction of a Photocatalysis Dataset for Water-Splitting Applications

We present an automatically generated dataset of 15,755 records that were extracted from 47,357 papers. These records contain water-splitting activity in the presence of certain photocatalysts, along with additional information about the chemical reaction conditions under which this activity was recorded. These conditions include any co-catalysts and additives that were present during water splitting, the length of time for which the photocatalytic experiment was conducted, and the type of light source used, including its wavelength. Despite the text extraction of such a wide range of chemical reaction attributes, the dataset afforded good precision (71.2%) and recall (36.3%). These figures-of-merit were calculated based on a random sample of open-access papers from the corpus. Mining such a complex set of attributes required the development of novel techniques in knowledge extraction and interdependency resolution, leveraging inter- and intra-sentence relations, which are also described in this paper. We present a new version (version 2.2) of the chemistry-aware text-mining toolkit ChemDataExtractor, in which these new techniques are included.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Projected U.S. drought extremes through the twenty-first century with vapor pressure deficit

Global warming is expected to enhance drought extremes in the United States throughout the twenty-first century. Projecting these changes can be complex in regions with large variability in atmospheric and soil moisture on small spatial scales. Vapor Pressure Deficit (VPD) is a valuable measure of evaporative demand as moisture moves from the surface into the atmosphere and a dynamic measure of drought. Here, VPD is used to identify short-term drought with the Standardized VPD Drought Index (SVDI); and used to characterize future extreme droughts using grid dependent stationary and non-stationary generalized extreme value (GEV) models, and a random sampling technique is developed to quantify multimodel uncertainties. The GEV analysis was performed with projections using the Weather Research and Forecasting model, downscaled from three Global Climate Models based on the Representative Concentration Pathway 8.5 for present, mid-century and late-century. Results show the VPD based index (SVDI) accurately identifies the timing and magnitude short-term droughts, and extreme VPD is increasing across the United States and by the end of the twenty-first century. The number of days VPD is above 9 kPa increases by 10 days along California’s coastline, 30–40 days in the northwest and Midwest, and 100 days in California’s Central Valley.

54 ENVIRONMENTAL SCIENCES↗

Coupling flux balance analysis with reactive transport modeling through machine learning for rapid and stable simulation of microbial metabolic switching

Integrating genome-scale metabolic networks with reactive transport models (RTMs) provides a detailed description of the dynamic changes in microbial growth and metabolism. Despite promising demonstrations in the past, computational inefficiency has been pointed out as a critical issue to overcome because it requires repeated application of linear programming (LP) to obtain flux balance analysis (FBA) solutions in every time step and spatial grid. To address this challenge, we propose a new simulation method where we train and validate artificial neural networks (ANNs) using randomly sampled FBA solutions and incorporate the resulting surrogate FBA model (represented as algebraic equations) into RTMs as source/sink terms. We demonstrate the efficiency of our method via a case study of Shewanella oneidensis MR-1. During aerobic growth on lactate, S. oneidensis produces metabolic byproducts (such as pyruvate and acetate), which are subsequently consumed as alternative carbon sources when the preferred nutrients are depleted. To effectively simulate these complex dynamics, we used a cybernetic approach that models metabolic switches as the outcome of dynamic competition among multiple growth options. In both zero-dimensional batch and one-dimensional column configurations, the ANN-based surrogate models achieved substantial reduction of computational time by several orders of magnitude compared to the original LP-based FBA models. Moreover, the ANN models produced robust solutions without any special measures to prevent numerical instability. These developments significantly promote our ability to utilize genome-scale networks in complex, multi-physics, and multi-dimensional ecosystem modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Classical and quantum simulations of 1+1-dimensional ${\mathbb{Z}}_{2}$ gauge theory at finite temperature and density

Simulating strongly coupled gauge theories at finite temperature and density is a longstanding challenge in nuclear and high-energy physics with fundamental implications for condensed matter physics. Here, we simulate such systems using minimally entangled typical thermal state (METTS) approaches, which combine classical random sampling with imaginary-time evolution, implementable on either classical or quantum computers, to estimate thermal averages of observables. We study 1+1-dimensional ${\mathbb{Z}}_{2}$ gauge theory coupled to spinless fermionic matter, which maps onto a local quantum spin chain. We benchmark both a classical matrix-product-state implementation of METTS and a recently proposed adaptive variational approach for near-term quantum devices, focusing on the equation of state and measures of fermion confinement. Of particular importance is the choice of basis for METTS sampling, which impacts both the sampling overhead and quantum circuit complexity. Our work sets the stage for future studies of strongly coupled gauge theories using classical and quantum hardware.

Chen, I-Chi [Iowa State Univ., Ames, IA (United St↗

A proposed method for addressing large unphysical uncertainties in mubar

Mubar uncertainties in Section MF 34 MT 2 of ENDF/B-VIII.0 data are sometimes too large. Physical limitations of the bounds of mubar limit the maximum uncertainty to be < 1.0 and usually $\ll$ 1.0. Two ad hoc methods are proposed to artificially constrain the randomly sampled values of mubar to allowable values. The application of NJOY’s mubar covariance matrix to the “sandwich rule” is also discussed and a consistent way to apply sensitivity coefficients is described. Mubar uncertainties for the Jezebel k eff were found to be 156 pcm – a value comparable to other evaluated cross section uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Monte Carlo Perturbation Analysis of Fuel Temperature Variations in the MCNP Model of the Annular Core Research Reactor

The Annular Core Research Reactor (ACRR) Monte Carlo N-Particle (MCNP) model is used by ACRR reactor operators and experiment designers at Sandia National Laboratories for a variety of computational calculations ranging from reactor kinetics parameter estimates and safety analyses to experimental planning. To understand the dominant source of uncertainty within the MCNP model, perturbations in temperature were applied to individual ACRR MCNP fuel rods. Fuel rod temperatures were randomly sampled from a uniform distribution from operational temperatures to quantify temperature-related uncertainty effects. Stochastic mixing was used to blend the cross sections of the desired temperatures using the MCNP continuous and Thermal Neutron Scattering Treatment [S(α,β)] libraries in ENDF/B-VII.1. Furthermore, this uncertainty analysis produced a 640 row × 640 column correlation and covariance matrix of the neutron energy spectra. Positive covariance was produced around the 1-MeV region and the 0.2-eV region. Correlation was found in the thermal and fast energy regions, but no correlation was observed in the slowing-down energy region because interactions in this region are not dominated by fuel.

ACRR↗

Uncertainty Quantification of a Light Water Pulsed-Neutron Die-Away Experiment to Thermal Neutron Scattering Laws

Thermal neutron scattering laws are important nuclear data for many nuclear science and engineering applications. Validation helps to ensure that a thermal neutron scattering law has a high quality and often employs critical benchmarks as integral experiments. Recently, pulsed-neutron die-away benchmarks have been used as an experiment to validate thermal neutron scattering laws. Herein, we evidence how this alternative integral experiment has a high sensitivity to these nuclear data by performing an uncertainty quantification analysis. The analysis randomly sampled the nuclear model parameters associated with hydrogen bound in light water thermal neutron scattering law and sampled other nuclear data that influenced the experiment’s integral parameter (e.g., elastic scattering, absorption in hydrogen and oxygen) from their respective covariance matrices. The thermal neutron scattering law caused an uncertainty in the integral parameter that reached 2.67%, which exceeds by an order of magnitude the uncertainties induced in commonly used thermal solution critical benchmarks. The validation performed here, although limited due to a poor description of the historical experiment, indicated that the ENDF/B-VIII.0 thermal neutron scattering law well predicted the integral parameter. These results motivate further benchmark and validation efforts using pulsed-neutron die-away experiments.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Finding simplicity: unsupervised discovery of features, patterns, and order parameters via shift-invariant variational autoencoders *

Abstract Recent advances in scanning tunneling and transmission electron microscopies (STM and STEM) have allowed routine generation of large volumes of imaging data containing information on the structure and functionality of materials. The experimental data sets contain signatures of long-range phenomena such as physical order parameter fields, polarization, and strain gradients in STEM, or standing electronic waves and carrier-mediated exchange interactions in STM, all superimposed onto scanning system distortions and gradual changes of contrast due to drift and/or mis-tilt effects. Correspondingly, while the human eye can readily identify certain patterns in the images such as lattice periodicities, repeating structural elements, or microstructures, their automatic extraction and classification are highly non-trivial and universal pathways to accomplish such analyses are absent. We pose that the most distinctive elements of the patterns observed in STM and (S)TEM images are similarity and (almost-) periodicity, behaviors stemming directly from the parsimony of elementary atomic structures, superimposed on the gradual changes reflective of order parameter distributions. However, the discovery of these elements via global Fourier methods is non-trivial due to variability and lack of ideal discrete translation symmetry. To address this problem, we explore the shift-invariant variational autoencoders (shift-VAEs) that allow disentangling characteristic repeating features in the images, their variations, and shifts that inevitably occur when randomly sampling the image space. Shift-VAEs balance the uncertainty in the position of the object of interest with the uncertainty in shape reconstruction. This approach is illustrated for model 1D data, and further extended to synthetic and experimental STM and STEM 2D data. We further introduce an approach for training shift-VAEs that allows finding the latent variables that comport to known physical behavior. In this specific case, the condition is that the latent variable maps should be smooth on the length scale of the atomic lattice (as expected for physical order parameters), but other conditions can be imposed. The opportunities and limitations of the shift VAE analysis for pattern discovery are elucidated.

97 MATHEMATICS AND COMPUTING↗

Global Sensitivity Analysis of a Reactive Transport Model for Mineral Scale Formation During Hydraulic Fracturing

Injection of water-based hydraulic fracturing fluid (HFF) into tight shale gas/oil formations can increase formation permeability and enhance production rates, but this process frequently causes mineral scale formation that can occlude pore space and hinder flow. To identify the most important factors that control the formation of mineral scales, we applied a novel global sensitivity analysis method—distance-based generalized sensitivity analysis (DGSA)—to a reactive transport model (RTM) that was previously built and calibrated to simulate precipitation of barite [BaSO4] and iron (hydr)oxide [Fe(OH) 3 ] in shale matrices and on fracture surfaces. Reactive transport simulations were run with model parameters randomly sampled based on assigned uncertainties. Modeling results for barite and Fe(OH)3 formation were clustered using machine-learning algorithms. A list of ranked critical input parameters was obtained after statistical quantification of cumulative distribution functions of input parameters. We found that barite formation is most sensitive to the rate of sulfate ion generation, which is determined by the pyrite dissolution rate coefficient and oxidant availability. In addition, barite formation is sensitive to the initial amounts of barite in HFF and shale, followed by barite thermodynamics/kinetics. For Fe(OH) 3 formation, the ranked factors are Fe(OH)3 precipitation rate coefficients, initial HFF pH, initial Fe(OH) 3 amount in HFF, and oxidant availability. Overall, our results provide insights into managing mineral scale formation during hydraulic fracturing to enhance production. Meanwhile, this study serves as an example of global sensitivity analysis of RTMs using the efficient, straightforward, and open-source DGSA method.

58 GEOSCIENCES↗

Imaging the crust and uppermost mantle structure of Portugal (West Iberia) with seismic ambient noise

SUMMARY We present a new high-resolution 3-D shear wave velocity (Vs) model of the crust and uppermost mantle beneath Portugal, inferred from ambient seismic noise tomography. We use broad-band seismic data from a dense temporary deployment covering the entire Portuguese mainland between 2010 and 2012 in the scope of the WILAS project. Vertical component data are processed using phase correlation and phase weighted stack to obtain empirical Green functions (EGFs) for 2016 station pairs. Further, we use a random sampling and subset stacking strategy to measure robust Rayleigh-wave group velocities in the period range 7–30 s and associated uncertainties. The tomographic inversion is performed in two steps: First, we determine group-velocity lateral variations for each period. Next, we invert them at each grid point using a new trans-dimensional inversion scheme to obtain the 3-D shear wave velocity model. The final 3-D model extends from the upper crust (5 km) down to the uppermost mantle (60 km) and has a lateral resolution of ∼50 km. In the upper and middle crusts, the Vs anomaly pattern matches the tectonic units of the Variscan Massif and Alpine basins. The transition between the Lusitanian Basin and the Ossa Morena Zone is marked by a contrast between moderate- and high-velocity anomalies, in addition to two arched earthquake lineations. Some faults, namely, the Manteigas–Vilariça–Bragança fault and the Porto–Tomar–Ferreira do Alentejo fault, have a clear signature from the upper crust down to the uppermost mantle (60 km). Our 3-D shear wave velocity model offers new insights into the continuation of the main tectonic units at depth and contributes to better understanding the seismicity of Portugal.

Silveira, Graça (ORCID:0000000221102554)↗

Emulation of seismic-phase traveltimes with machine learning

SUMMARY We present a machine learning (ML) method for emulating seismic-phase traveltimes that are computed using a global-scale 3-D earth model and physics-based ray tracing. Accurate traveltime predictions based on 3-D earth models are known to reduce the bias of event location estimates, increase our ability to assign phase labels to seismic detections and associate detections to events. However, practical use of 3-D models is challenged by slow computational speed and the unwieldiness of pre-computed lookup tables that are often large and have prescribed computational grids. In this work, we train a ML emulator using pre-computed traveltimes, resulting in a compact and computationally fast way to approximate traveltimes that are based on a 3-D earth model. Our model is trained using approximately 850 million P-wave traveltimes that are based on the global LLNL-G3D-JPS model, which was developed for more accurate event location. The training-set consists of traveltimes between 10 393 global seismic stations and randomly sampled event locations that provide a prescribed, distance-dependent geographic sample density for each station. Prediction accuracy is dependent on event-station distance and whether the station was included in the training set. For stations included in the training set the mean absolute deviation (MAD) of the difference between traveltimes computed using ray tracing through the 3-D model and the ML emulator for local, regional, and teleseismic distances are 0.090, 0.125 and 0.121 s, respectively. For tested station locations not included in the training set, MAD values for the three distance ranges increase to 0.173, 0.219 and 0.210 s, respectively. Empirical traveltime residuals for a global reference data are indistinguishable when ML emulation or the 3-D model is used to compute traveltimes. This result holds regardless of whether the recording station is used in ML training or not.

58 GEOSCIENCES↗

Is Terzan 5 the remnant of a building block of the Galactic bulge? Evidence from APOGEE

ABSTRACT It has been proposed that the globular cluster-like system Terzan 5 is the surviving remnant of a primordial building block of the Milky Way bulge, mainly due to the age/metallicity spread and the distribution of its stars in the α–Fe plane. We employ Sloan Digital Sky Survey data from the Apache Point Observatory Galactic Evolution Experiment to test this hypothesis. Adopting a random sampling technique, we contrast the abundances of 10 elements in Terzan 5 stars with those of their bulge field counterparts with comparable atmospheric parameters, finding that they differ at statistically significant levels. Abundances between the two groups differ by more than 1σ in Ca, Mn, C, O, and Al, and more than 2σ in Si and Mg. Terzan 5 stars have lower [α/Fe] and higher [Mn/Fe] than their bulge counterparts. Given those differences, we conclude that Terzan 5 is not the remnant of a major building block of the bulge. We also estimate the stellar mass of the Terzan 5 progenitor based on predictions by the Evolution and Assembly of GaLaxies and their Environments suite of cosmological numerical simulations, concluding that it may have been as low as ∼3 × 108 M⊙ so that it was likely unable to significantly influence the mean chemistry of the bulge/inner disc, which is significantly more massive (∼1010 M⊙). We briefly discuss existing scenarios for the nature of Terzan 5 and propose an observational test that may help elucidate its origin.

79 ASTRONOMY AND ASTROPHYSICS↗

The Adaptive Potential of the Middle Domain of Yeast Hsp90

Abstract The distribution of fitness effects (DFEs) of new mutations across different environments quantifies the potential for adaptation in a given environment and its cost in others. So far, results regarding the cost of adaptation across environments have been mixed, and most studies have sampled random mutations across different genes. Here, we quantify systematically how costs of adaptation vary along a large stretch of protein sequence by studying the distribution of fitness effects of the same ≈2,300 amino-acid changing mutations obtained from deep mutational scanning of 119 amino acids in the middle domain of the heat shock protein Hsp90 in five environments. This region is known to be important for client binding, stabilization of the Hsp90 dimer, stabilization of the N-terminal-Middle and Middle-C-terminal interdomains, and regulation of ATPase–chaperone activity. Interestingly, we find that fitness correlates well across diverse stressful environments, with the exception of one environment, diamide. Consistent with this result, we find little cost of adaptation; on average only one in seven beneficial mutations is deleterious in another environment. We identify a hotspot of beneficial mutations in a region of the protein that is located within an allosteric center. The identified protein regions that are enriched in beneficial, deleterious, and costly mutations coincide with residues that are involved in the stabilization of Hsp90 interdomains and stabilization of client-binding interfaces, or residues that are involved in ATPase–chaperone activity of Hsp90. Thus, our study yields information regarding the role and adaptive potential of a protein sequence that complements and extends known structural information.

Cote-Hammarlof, Pamela A.↗

Composite Qdrift-product formulas for quantum and classical simulations in real and imaginary time

Recent study has shown that it can be advantageous to implement a composite channel that partitions the Hamiltonian H for a given simulation problem into subsets A and B such that H = A + B , where the terms in A are simulated with a Trotter-Suzuki channel and the B terms are randomly sampled via the Qdrift algorithm. Here we extend Qdrift and composite product formulas to imaginary time, formulating candidate classical algorithms for quantum Monte Carlo calculations. We upper bound the induced Schatten- 1 → 1 norm on both imaginary-time Qdrift and composite channels. Another recent result demonstrated that simulations of lattice Hamiltonians containing geometrically local interactions can be improved using a Lieb-Robinson argument to decompose H into subsets that contain only terms supported on that subset of the lattice. Here, we provide a quantum algorithm by unifying this result with the composite approach into “local composite channels” and we upper bound the diamond distance. We provide exact numerical simulations of algorithmic cost by counting the number of gates of the form e − i H j t and e − H j β to meet a certain error tolerance ε . In doing so, we optimize the partitioning into sets A and B using gradient boosted tree models from machine learning. These numerical studies are important given that product formulas have been historically known to outperform analytic upper bounds. We show constant factor advantages for a variety of interesting Hamiltonians, the maximum of which is a ≈ 20 -fold speedup that occurs in the simulation of Jellium. Published by the American Physical Society 2024

Pocrnic, Matthew (ORCID:0000000203089376)↗

Aspects of propagator sparsening in lattice QCD

In lattice field theory, field sparsening aims to replace quantum fields, or objects constructed from them, with approximations that preserve the appropriate symmetries and maintain many aspects of the physics that the fields determine. For example, an effective sparsening of a quark propagator provides an efficient map from a quark propagator on a fine lattice geometry to a quark propagator defined on a coarser geometry in order to reduce storage and computational costs of subsequent calculational stages while maintaining long-distance correlations and corresponding low-energy physical information. Previous studies have focused on decimating lattice sites or randomly sampling lattice sites to reduce the size of the propagator and subsequent costs of Wick contractions. Here, we extend the study of sparsening to incorporate covariant averaging of spatial sites and examine the effects on two-point and three-point correlation functions involving various hadrons. We find that sparsening is most effective in reproducing the unsparsened versions of these correlation functions when weighted covariant-averaging is sequentially applied many times.

Lattice QCD↗

Mitigating Catastrophic Forgetting in Deep Learning in a Streaming Setting Using Historical Summary

Recent advancements in scientific equipment and the adaptation of electronics and the Internet of Things (IoT) in our everyday lives resulted in large and complex data production at a high rate. Making meaningful and timely knowledge discovery at a modest cost from this big data is difficult for computing power and storage limitations. Training deep learning models incrementally in a streaming setting can help us with overcoming these limitations. However, in a well-known phenomenon named catastrophic forgetting, incrementally trained models increasingly perform poorly on the past data. To mitigate catastrophic forgetting in training in a streaming setting, we propose constructing a historical summary over time and use the summary with newly arrived data during incremental training. We propose various data summarization techniques such as random sampling, micro clustering, coreset computation, and Auto Encoders to counteract catastrophic forgetting. We built a pipeline for incremental training with a historical summary for training deep learning models for streaming data. We demonstrate the effectiveness of historical summary in mitigating catastrophic forgetting using three case studies involving three different deep learning applications: an Artificial Neural Network (ANN) for classification task on MNIST dataset, a language model (RNN-LM) on the WikiText2 dataset, and a Convolutional Neural Network (CNN), ResNet50 to classify the ImageNet dataset. Through the training of the models, we observe that catastrophic forgetting is evident in ANN and CNN but not in an RNN. For the first task, our method recovers up to 47.9% lost accuracy due to catastrophic forgetting. For the third task, the historical summary recovers classification accuracy by up to 25%. For the second task, though there is not proof of catastrophic forgetting, the training performance (PPL) improves by up to 26% with historical summary.

Dash, Sajal↗