Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Randomized methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

A Computational Information Criterion for Particle-Tracking with Sparse or Noisy Data

Traditional probabilistic methods for the simulation of advection-diffusion equations (ADEs) often overlook the entropic contribution of the discretization, e.g., the number of particles, within associated numerical methods. Many times, the gain in accuracy of a highly discretized numerical model is outweighed by its associated computational costs or the noise within the data. Herein, we address the question of how many particles are needed in a simulation to best approximate and estimate parameters in one-dimensional advective-diffusive transport. To do so, we use the well-known Akaike Information Criterion (AIC) and a recently-developed correction called the Computational Information Criterion (COMIC) to guide the model selection process. Random-walk and mass-transfer particle tracking methods are employed to solve the model equations at various levels of discretization. Numerical results demonstrate that the COMIC provides an optimal number of particles that can describe a more efficient model in terms of parameter estimation and model prediction compared to the model selected by the AIC even when the data is sparse or noisy, the sampling volume is not uniform throughout the physical domain, or the error distribution of the data is non-IID Gaussian.

97 MATHEMATICS AND COMPUTING↗

Late Breaking Results: COPPER: Computation Obfuscation by Producing Permutations for Encoding Randomly

Deployed embedded devices face security risks due to increased ease of physical access to the devices by unauthorized users. Capable adversaries can intercept a device to recover the data in memory, including results of performed sensitive computations. Device owners require data confidentiality on their physically insecure devices. To satisfy this goal we implement a novel method, COPPER (Computation Obfuscation by Producing Permutations for Encoding Randomly), to create data which never exists on the device digitally in plaintext format and which is subsequently used for computation. In this paper we utilize COPPER to calculate a moving average computation on encoded data.

embedded systems↗

RandONets: Shallow networks with random projections for learning linear and nonlinear operators

Deep neural networks have been extensively used for the solution of both the forward and the inverse problem for dynamical systems. However, their implementation necessitates optimizing a high-dimensional space of parameters and hyperparameters. This fact, along with the requirement of substantial computational resources, pose a barrier to achieving high numerical accuracy, but also interpretability. Here, to address the above challenges, we present Random Projection-based Operator Networks (RandONets): shallow networks with random projections and tailor-made numerical analysis methods that learn accurately and fast linear and nonlinear operators. Building on previous works, we prove that RandOnets are universal approximators of linear and nonlinear operators. Due to their simplicity, RandONets provide a one-step transformation of the input space, facilitating interpretability. For the evaluation of their performance, we focus on operators of PDEs. We show, that RandONets outperform by several orders of magnitude, both in terms of numerical approximation accuracy and computational cost, the “vanilla” DeepONets. Hence, we believe that our method will trigger further developments in the field of scientific machine learning, for the development of new ‘’light”schemes that will provide high accuracy while reducing dramatically the computational cost. A MATLAB toolbox for RandONets, including demos, is available on GitHub at https://github.com/GianlucaFabiani/RandONets.

Interpretable machine learning↗

Static Subspace Approximation for Random Phase Approximation Correlation Energies: Applications to Materials for Catalysis and Electrochemistry

Modeling complex materials using high-fidelity, ab initio methods at low cost is a fundamental goal for quantum chemical software packages. The GW approximation and random phase approximation (RPA) provide a unified description of both electronic structure and total energies using the same physics in a many-body perturbative approach that can be more accurate than generalized-gradient density functional theory (DFT) methods. However, GW/RPA implementations have historically been limited to either specific materials classes or application toward small chemical systems. Here, the static subspace approximation allows for reduced cost full-frequency GW/RPA calculations and has previously been benchmarked thoroughly for GW calculations. Here, we describe our approach to including partial occupations of electronic orbitals in full-frequency GW and RPA calculations for the study of electrocatalysts. We benchmarked RPA total energy calculations using the subspace approximation across a diverse test suite of materials for a variety of computational parameters. The benchmarking quantifies the impact of different extrapolation procedures for representing the static polarizability at infinite screened cutoff, and shows that using screened cutoffs above 20-25 Ryd result in diminishing accuracy returns for predicting RPA total energies. Additionally, for moderately sized electrocatalytic models, 2-3 times fewer computational resources are used to compute RPA total energies by representing the static polarizability with 20-30% of the static subspace basis, with an error of approximately 0.01 eV or better in RPA adsorption energy calculations. Finally, we show that for these electrochemical models RPA can shift DFT adsorption energy shifts by up to 0.5 eV and that GW can frequently shift DFT eigenvalues of surface and adsorbate states by approximately 0.5-1 eV.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A scalable variational method for estimating the latent infection-rate field of an outbreak

In this paper, we explore whether the infection-rate of a disease can serve as a robust monitoring variable in epidemiological surveillance algorithms. The infection-rate is dependent on population mixing patterns that do not vary erratically day-to-day; in contrast, daily case-counts used in contemporary surveillance algorithms are corrupted by reporting errors. The technical challenge lies in estimating the latent infection-rate from case-counts. Here we devise a Bayesian method to estimate the infection-rate across multiple adjoining areal units, and then use it, via an anomaly detector, to discern a change in epidemiological dynamics. We extend an existing model for estimating the infection-rate in an areal unit by incorporating a Markov random field model, so that we may estimate infection-rates across multiple areal units, while preserving spatial correlations observed in the epidemiological dynamics. To carry out the high-dimensional Bayesian inverse problem, we develop an implementation of mean-field variational inference specific to the infection model and integrate it with the random field model to incorporate correlations across counties. The method is tested on estimating the COVID-19 infection-rates across all 33 counties in New Mexico using data from the summer of 2020, and then employing them to detect the arrival of the Fall 2020 COVID-19 wave. We perform the detection using a temporal algorithm that is applied county-by-county. We also show how the infection-rate field can be used to cluster counties with similar epidemiological dynamics.

60 APPLIED LIFE SCIENCES↗

Estimating Subhourly Inverter Clipping Loss From Satellite-Derived Irradiance Data

Photovoltaic system production simulations are conventionally run using hourly weather datasets. Hourly simulations are sufficiently accurate to predict the majority of long-term system behavior but cannot resolve high-frequency effects like inverter clipping caused by short-duration irradiance variability. Direct modeling of this subhourly clipping error is only possible for the few locations with high-resolution irradiance datasets. This paper describes a method of predicting the magnitude of this error using a machine learning regressor ensemble model, comprised of a random forest and an XGBoost model, and 30-minute satellite irradiance data. The method predicts a correction for each 30-minute interval with the potential to roll up into 60-minute corrections to match an hourly energy model. The model is trained and validated at locations where the error can be directly simulated from 1-minute ground data. The validation shows low bias at most ground station locations. The model is also applied to gridded satellite irradiance to produce a heatmap of the estimated clipping error across the United States. Finally, the relative importance of each predictor satellite variable is retrieved from the model and discussed.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Differentially Private Synthesis and Sharing of Network Data Via Bayesian Exponential Random Graph Models

Abstract Network data often contain sensitive relational information. One approach to protecting sensitive information while offering flexibility for network analysis is to share synthesized networks based on the information in originally observed networks. We employ differential privacy (DP) and exponential random graph models (ERGMs) and propose the DP-ERGM method to synthesize network data. We apply DP-ERGM to two real-world networks. We then compare the utility of synthesized networks generated by DP-ERGM, the DyadWise Randomized Response (DWRR) approach, and the Synthesis through Conditional distribution of Edge given nodal Attribute (SCEA) approach. In general, the results suggest that DP-ERGM preserves the original information significantly better than two other approaches in network structural statistics and inference for ERGMs and latent space models. Furthermore, DP-ERGM satisfies node DP through modeling the global network structure with ERGM, a stronger notion of privacy than the edge DP under which DWRR and SCEA operate.

graph synthesis↗

Circulating levels of micronutrients and risk of osteomyelitis: a Mendelian randomization study

Background Few observational studies have investigated the effect of micronutrients on osteomyelitis, and these findings are limited by confounding and conflicting results. Therefore, we conducted Mendelian randomization (MR) analyses to evaluate the association between blood levels of eight micronutrients (copper, selenium, zinc, vitamin B12, vitamin C, and vitamin D, vitamin B6, vitamin E) and the risk of osteomyelitis. Methods We performed the two-sample and multivariable Mendelian randomization (MVMR) to investigate causation, where instrument variables for the predictor (micronutrients) were derived from the summary data of micronutrients from independent cohorts of European ancestry. The outcome instrumental variables were used from the summary data of European-ancestry individuals ( n = 486,484). The threshold of statistical significance was set at p < 0.00625. Results We found a significant causal association that elevated zinc heightens the risk of developing osteomyelitis in European ancestry individuals OR = 1.23 [95% confidence interval (CI) [1.07, 1.43]; p = 4.26E-03]. Similarly, vitamin B6 showed a similar significant causal effect on osteomyelitis as a risk factor OR = 2.78 (95% CI [1.34, 5.76]; p = 6.04E-03; in the secondary analysis). Post-hoc analysis suggested this result (vitamin B6). However, the multivariable Mendelian randomization (MVMR) provides evidence against the causal association between zinc and osteomyelitis OR = 0.98(95% CI [−0.11, 0.07]; p = 7.20E-1). After searching in PhenoScanner, no SNP with confounding factors was found in the analysis of vitamin B6. There was no evidence of a reverse causal impact of osteomyelitis on zinc and vitamin B6. Conclusion This study supported a strong causal association between vitamin B6 and osteomyelitis while reporting a dubious causal association between zinc and osteomyelitis.

Zhang, Xu↗

Entanglement features of random neural network quantum states

Restricted Boltzmann machines (RBMs) are a class of neural networks that have been successfully employed as a variational ansatz for quantum many-body wave functions. Here, we develop an analytic method to study quantum many-body spin states encoded by random RBMs with independent and identically distributed complex Gaussian weights. By mapping the computation of ensemble-averaged quantities to statistical mechanics models, we are able to investigate the parameter space of the RBM ensemble in the thermodynamic limit. We discover qualitatively distinct wave functions by varying RBM parameters, which correspond to distinct phases in the equivalent statistical mechanics model. Notably, there is a regime in which the typical RBM states have near-maximal entanglement entropy in the thermodynamic limit, similar to that of Haar-random states. However, these states generically exhibit nonergodic behavior in the Ising basis, and do not form quantum state designs, making them distinguishable from Haar-random states.

36 MATERIALS SCIENCE↗

Likelihood-Based Particle Identification in the Short-Baseline Near Detector

Accurate particle identification is crucial in any high-energy physics experiment, allowing scientists to understand the unique interactions and mechanisms at play in a detector. In this project, I develop and study a new particle identification (PID) algorithm for the Short-Baseline Near Detector, a likelihood-based approach, different from out current $\chi^2$ method. A likelihood estimation offers a more physically motivated strategy for PID. The distribution random energy losses of charged particles traveling through a medium are described by the Vavilov probability density function. By using this model, we can account for random energy losses and construct likelihood functions specific to each particle type, potentially enabling a more accurate method for PID.

Vanderwaal, Sophia [U. Alabama, Huntsville] (ORCID↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Coupling flux balance analysis with reactive transport modeling through machine learning for rapid and stable simulation of microbial metabolic switching

Integrating genome-scale metabolic networks with reactive transport models (RTMs) provides a detailed description of the dynamic changes in microbial growth and metabolism. Despite promising demonstrations in the past, computational inefficiency has been pointed out as a critical issue to overcome because it requires repeated application of linear programming (LP) to obtain flux balance analysis (FBA) solutions in every time step and spatial grid. To address this challenge, we propose a new simulation method where we train and validate artificial neural networks (ANNs) using randomly sampled FBA solutions and incorporate the resulting surrogate FBA model (represented as algebraic equations) into RTMs as source/sink terms. We demonstrate the efficiency of our method via a case study of Shewanella oneidensis MR-1. During aerobic growth on lactate, S. oneidensis produces metabolic byproducts (such as pyruvate and acetate), which are subsequently consumed as alternative carbon sources when the preferred nutrients are depleted. To effectively simulate these complex dynamics, we used a cybernetic approach that models metabolic switches as the outcome of dynamic competition among multiple growth options. In both zero-dimensional batch and one-dimensional column configurations, the ANN-based surrogate models achieved substantial reduction of computational time by several orders of magnitude compared to the original LP-based FBA models. Moreover, the ANN models produced robust solutions without any special measures to prevent numerical instability. These developments significantly promote our ability to utilize genome-scale networks in complex, multi-physics, and multi-dimensional ecosystem modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Using intrusive approaches as a step towards accounting for stochasticity in wind turbine design

Current wind turbine design methods require tens of thousands of time-domain simulations and use different random seeds to account for the stochasticity of the environmental conditions. The account of stochasticity is nonintrusive because the sampling method calls a deterministic model multiple times without changing its underlying equations. In this work, we investigate and demonstrate using simple proof of concepts how intrusive approaches can be used to directly account for stochasticity in the equations representing a mechanical system. Our long term goal is to apply such methodology to the design of wind turbines without requiring an excessive number of simulations. Intrusive methods manipulate stochastic variables directly to provide the probability density functions (PDFs) of the states and outputs at any time as functions of the PDFs of the inputs. We illustrate how different methods can be used with a reduced-order model of a wind turbine with one degree of freedom and for linear and nonlinear models. We discuss how the methods can be extended and what it will take to apply them to a level of fidelity similar to current state-of-the-art wind turbine design tools.

17 WIND ENERGY↗

Two datasets are better than one: method of double moments for 3D reconstruction in cryo-EM

Cryo-electron microscopy is a powerful imaging technique for reconstructing three-dimensional molecular structures from noisy tomographic projection images of randomly oriented particles. We introduce a new data fusion framework, termed the method of double moments, which reconstructs molecular structures from two instances of the second-order moment of projection images obtained under distinct orientation distributions: one uniform, the other non-uniform and unknown. We prove that these moments generically uniquely determine the underlying structure, up to a global rotation and reflection, and we develop a convex-relaxation-based algorithm that achieves accurate recovery using only second-order statistics. Our results demonstrate the advantage of collecting and modeling multiple datasets under different experimental conditions, illustrating that leveraging dataset diversity can substantially enhance reconstruction quality in computational imaging tasks.

Kam’s method↗

Generation of random geological models using multi-randomization for machine learning

Generating high-fidelity geological models is essential for advancing machine learning (ML) methods in automated seismic interpretation. For instance, seismic images paired with corresponding fault labels are foundational for ML-based fault detection from seismic migration sections. While several open-access datasets of random geological models exist, open-source tools specifically designed to produce large volumes of such models for ML applications remain scarce. To address this gap, we present RGM (Random Geological Model), an open-source software package for efficiently generating 2D and 3D synthetic geological models tailored for ML workflows. RGM supports the creation of diverse model components, including medium property distributions (P-/S-wave velocities and density), seismic reflectivity images (i.e., synthetic migration sections), relative geological time, and discrete fault attributes such as probability, dip, strike, rake, and displacement. It also accommodates the creation of complex geological features such as salt bodies and unconformities. The model generation algorithm employs a multi-randomization strategy, yielding an effectively infinite-dimensional model space that encompasses a wide range of geological scenarios and associated seismic features. Furthermore, RGM incorporates a method to generate synthetic elastic migration images using analytical elastic reflection coefficients combined with frequency-dependent scaling. This functionality enables the creation of training datasets for ML models that leverage elastic seismic images. RGM is implemented in modern object-oriented Fortran, allowing users to flexibly control statistical parameters governing model variability. We demonstrate the capability, performance, and geological realism of the package through comprehensive 2D and 3D examples.

58 GEOSCIENCES↗

Machine Learning Analysis of Impact of Western US Fires on Central US Hailstorms

Fires, including wildfires, harm air quality and essential public services like transportation, communication, and utilities. These fires can also influence atmospheric conditions, including temperature and aerosols, potentially affecting severe convective storms. Here, we investigate the remote impacts of fires in the western United States (WUS) on the occurrence of large hail (size: $\geqslant$ 2.54 cm) in the central US (CUS) over the 20-year period of 2001–20 using the machine learning (ML), Random Forest (RF), and Extreme Gradient Boosting (XGB) methods. The developed RF and XGB models demonstrate high accuracy (> 90%) and F1 scores of up to 0.78 in predicting large hail occurrences when WUS fires and CUS hailstorms coincide, particularly in four states (Wyoming, South Dakota, Nebraska, and Kansas). The key contributing variables identified from both ML models include the meteorological variables in the fire region (temperature and moisture), the westerly wind over the plume transport path, and the fire features (i.e., the maximum fire power and burned area). Importantly, the results confirm a linkage between WUS fires and severe weather in the CUS, corroborating the findings of our previous modeling study conducted on case simulations with a detailed physics model.

54 ENVIRONMENTAL SCIENCES↗

The role of surface chemistry in the synthesis of supported CuPd bimetallic/intermetallic catalysts for selective hydrogenation reactions

Understanding and controlling the structure of supported bimetallic/intermetallic catalysts are crucial for many heterogeneous catalytic reactions. However, commonly used synthesis techniques such as coimpregnation and coadsorption often yield nonuniform alloying and phase segregation, largely owing to the lack of interactions between different metal precursors. Here we show that by adopting sequential adsorption of complex metal cations and anions with the assistance of a proper ligand, the interaction between different metal precursors is enhanced, thus resulting in uniformly alloyed intermetallic catalysts. Importantly, the supported CuPd intermetallic catalyst exhibits greatly improved selectivity towards monoolefins in the semihydrogenation of acetylene and butadiene compared to those random CuPd alloys synthesized via coimpregnation and coadsorption methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning models for estimating contamination across different curbside collection strategies

Contaminated recyclables, which are frequently discarded as waste, pose a significant challenge to the implementation of a circular economy. These contaminated recyclables impede the circulation of resources, resulting in higher processing costs at material recovery facilities (MRFs). Over the past few decades, machine learning (ML) models such as linear regression (LR), support vector machine (SVM), and random forest (RF) have evolved to provide new methods for predicting inbound contamination rates in addition to traditional statistical models. In this study, we applied ML models to predict inbound contamination rates using demographic features from 15 counties in the U.S. with different curbside collection strategies. In general, we found that ML models outperformed linear mixed models. Specifically, SVM models had the highest performance (R 2 = 0.75; mean absolute error (MAE) = 0.06), which may be due to their ability to model nonlinear relationships between features and inbound contamination rates. Further, the key predictor was population, with poverty rate being positively correlated and median age negatively correlated with inbound contamination rates. To improve the management of contamination and enhance the implementation of a circular economy, better models are needed to understand and estimate inbound contamination rates as well as identify critical factors in the present and future.

54 ENVIRONMENTAL SCIENCES↗