Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Randomized methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

ML-Based Power System Stability Assessment Considering Network Topology Changes: WECC 20,000+ Bus System Case Study

Modern power grids are fast-changing and thus require real-time monitoring and online stability assessment. With the rapid development of machine learning (ML) techniques, using data-driven models to provide fast and accurate estimations of power system stability marginal information, such as frequency nadir for frequency stability and critical clearing time (CCT) for transient stability, have become possible. However, despite the numerous research on ML-based methods for frequency nadir and CCT prediction, there is limited work on the impact of different network topology changes. Furthermore, most previous studies only focused on small or synthetic systems, and there is a lack of research on actual large power system models. In this paper, the above issues are addressed by studying the actual U.S. Western Electricity Coordinating Council (WECC) system model with more than 20,000 buses. Massive simulations are conducted in PowerWorld Simulator to study the impact of various topology change scenarios on both frequency stability and transient stability. System operating information is extracted from the success dispatch cases of various network topologies to generate a comprehensive dataset for ML-based models. Two ML methods, random forest (RF) and multilayer perceptron (MLP) neural network, are trained and tested for both frequency nadir prediction and CCT prediction. Test results have proven the models are capable of online stability assessment for large power networks such as the WECC system with sufficient accuracy.

critical clearing time↗

Evaluating the performance of random forest and iterative random forest based methods when applied to gene expression data

Gene-to-gene networks, such as Gene Regulatory Networks (GRN) and Predictive Expression Networks (PEN) capture relationships between genes and are beneficial for use in downstream biological analyses. There exists multiple network inference tools to produce these gene-to-gene networks from matrices of gene expression data. Random Forest-Leave One Out Prediction (RF-LOOP) is a method that has been shown to be efficient at producing these gene-to-gene networks, frequently known as GEne Network Inference with Ensemble of trees (GENIE3). Random Forest can be replaced in this process by iterative Random Forest (iRF), which performs variable selection and boosting. Here we validate that iterative Random Forest-Leave One Out Prediction (iRF-LOOP) produces higher quality networks than GENIE3 (RF-LOOP). We use both synthetic and empirical networks from the Dialogue for Reverse Engineering Assessment and Methods (DREAM) Challenges by Sage Bionetworks, as well as two additional empirical networks created from Arabidopsis thaliana and Populus trichocarpa expression data.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluating the Performance of Random Forest and Iterative Random Forest Based Methods when Applied to Gene Expression Data

Gene-to-gene networks, such as Gene Regulatory Networks (GRN) and Predictive Expression Networks (PEN) capture relationships between genes and are beneficial for use in downstream biological analyses. There exists multiple network inference tools to produce these gene-to-gene networks from matrices of gene expression data. Random Forest-Leave One Out Prediction (RF-LOOP) is a method that has been shown to be efficient at producing these gene-to-gene networks, frequently known as GEne Network Inference with Ensemble of trees (GENIE3). Here we validate that iterative Random Forest-Leave One Out Prediction (iRF-LOOP) produces higher quality networks than GENIE3. We use both synthetic and empirical networks from the Dialogue for Reverse Engineering Assessment and Methods (DREAM) Challenges by Sage Bionetworks, as well as two additional empirical networks created from Arabidopsis thaliana and Populus trichocarpa expression data.

iRF-Loop, expression network, Populus Trichocarpa↗

Randomized Federated Learning Methods for Nonsmooth, Nonconvex, and Hierarchical Optimization (Final Technical Report)

This final technical report summarizes the outcomes of a DOE-funded project on federated scientific machine learning (FL) under nonsmooth, nonconvex, and hierarchical optimization settings. The project develops new mathematical models, algorithms, and theoretical guarantees for decentralized stochastic, bilevel, and minimax optimization problems arising in DOE mission-relevant applications. A unified framework of randomized and zeroth-order federated optimization methods is introduced, providing provable convergence, communication efficiency, and sample-complexity guarantees. The report documents algorithmic design, theoretical analysis, and empirical validation of the proposed federated learning methods. The project also contributes to workforce development through graduate training and dissemination of results via publications and seminars.

97 MATHEMATICS AND COMPUTING↗

Randomized Algorithms for Symmetric Nonnegative Matrix Factorization

Symmetric Nonnegative Matrix Factorization (SymNMF) is a technique in data analysis and machine learning that approximates a matrix with a product of a nonnegative, low-rank matrix and it transpose. To design faster and more scalable algorithms for SymNMF we develop two randomized algorithms for its computation. The first method uses randomized matrix sketching to compute an initial low-rank approximation to the input matrix and proceeds to uses this as a low-rank input to rapidly compute a SymNMF. The second methods uses randomized leverage score sampling to approximately solve constrained least squares problems. Many successful methods for SymNMF rely on (approximately) solving sequences of constrained least squares problems. Here, we prove theoretically that leverage score sampling can approximately solve constrained least squares problems to e-accuracy. Finally we demonstrate both methods work in practice by applying them to graph clustering tasks on large real world data sets. These experiments show that our methods approximately maintain solution quality and achieve significant speed ups for both large dense and large sparse problems.

97 MATHEMATICS AND COMPUTING↗

Devices and methods for increasing the speed and efficiency at which a computer is capable of modeling a plurality of random walkers using a particle method

A method for increasing a speed or energy efficiency at which a computer is capable of modeling a plurality of random walkers. The method includes defining a virtual space in which a plurality of virtual random walkers will move among different locations in the virtual space. The method also includes either assigning a corresponding set of ringed neurons in a spiking neural network to a corresponding virtual random walker, or assigning a corresponding set of ringed neurons to a point in the virtual space. Movement of a given virtual random walker is tracked by decoding differences between states of individual neurons in a corresponding given set of ringed neurons. A virtual random walk of the plurality of virtual random walkers is executed using the spiking neural network.

Aimone, James Bradley↗

Devices and methods for increasing the speed and efficiency at which a computer is capable of modeling a plurality of random walkers using a density method

A method for increasing a speed or energy efficiency at which a computer is capable of modeling a plurality of random walkers. The method includes defining a virtual space in which a plurality of virtual random walkers will move among different locations in the virtual space, wherein the virtual space comprises a plurality of vertices and wherein the different locations are ones of the plurality of vertices. A corresponding set of neurons in a spiking neural network is assigned to a corresponding vertex such that there is a correspondence between sets of neurons and the plurality of vertices, wherein a spiking neural network comprising a plurality of sets of spiking neurons is established. A virtual random walk of the plurality of virtual random walkers is executed using the spiking neural network, wherein executing includes tracking how many virtual random walkers are at each vertex at a given time increment.

Aimone, James Bradley↗

The Global LAnd Surface Satellite (GLASS) evapotranspiration product Version 5.0: Algorithm development and preliminary validation

An accurate estimation of spatially and temporally continuous global terrestrial evapotranspiration (ET) is essential in the assessment of surface energy, water and carbon cycles. The Global LAnd Surface Satellite (GLASS) ET product Version 4.0 (v4.0) based on the Bayesian model averaging (BMA) method was generated to estimate global terrestrial ET. However, certain uncertainty for the GLASS ET product v4.0 limits its application. In this study, we introduced the deep neural networks (DNN) merging framework to improve terrestrial ET estimation for GLASS ET product Version 5.0 (v5.0) generation by integrating five satellite-derived ET products [Moderate Resolution Imaging Spectroradiometer (MODIS) ET product (MOD16), Shuttleworth–Wallace dual-source ET product (SW), Priestley–Taylor-based ET product (PT-JPL), modified satellite-based Priestley–Taylor ET product (MS-PT) and simple hybrid ET product (SIM)]. We compared the performance of DNN method against other merging methods, including GLASS ET algorithm v4.0 (BMA), the gradient boosting regression tree (GBRT) method and the random forest (RF) method, based on 195 global eddy covariance (EC) flux towers covering observations from 2000 through 2015. Validations indicated that the DNN had the highest accuracy among four merging methods across different land cover types, yielding the highest average determination coefficients (R 2 , 0.62), root-mean-squared-error (RMSE, 24.1 W/m 2 ) and Kling–Gupta efficiency (KGE, 0.77) with a of 99% confidence interval. Compared with GLASS ET algorithm v4.0, the DNN improved on the R 2 by approximately 7% (p < 0.01) and the KGE by 10%. Based on the DNN, we then generated 8-day GLASS ET product v5.0 globally with a 1 km spatial resolution from 2001 to 2015 driven by GLASS vegetation and surface net radiation (R n ) datasets and Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA2) datasets. Finally, this global terrestrial ET product provides a valuable dataset for monitoring regional and global water resources and environmental changes.

54 ENVIRONMENTAL SCIENCES↗

Application of the locally self-consistent embedding approach to the Anderson model with non-uniform random distributions

Highlights: • Typical Medium Theory (TMT) for the Anderson Localization. • Locally Self-Consistent Multiple Scattering Method (LSMS) for Random Disordered Systems. • Linear Scaling Computational Method for Random Disordered Systems. We apply the recently developed embedding scheme for the locally self-consistent method to random disorder electrons systems. The method is based on the locally self-consistent multiple scattering theory and the typical medium theory. The locally self-consistent multiple scattering theory divides a system into many small designated local interaction zones. The subsystem within each local interaction zone is embedded in a self-consistent field from the typical medium theory. This approximation allows the study of random systems with large numbers of sites. We present results for the three dimensional Anderson model with different random disorder potential distributions. Using the typical density of states as an indicator of Anderson localization, we find that the method can capture the localization for commonly studied disorder potentials. These include the uniform distribution, the Gaussian distribution, and even the unbounded Cauchy distribution.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A non‐intrusive domain‐decomposition model reduction method for linear steady‐state partial differential equations with random coefficients

Abstract Domain decomposition methods have been proved to be an effective strategy to reduce the dimension of parametric partial differential equations (PDEs). However, existing domain decomposition methods for parametric PDEs are usually intrusive, which means domain decomposition based solvers need to be implemented from scratch for each target parametric PDE. To address this issue, we develop a new non‐intrusive domain‐decomposition model reduction method for linear steady‐state PDEs with random‐field coefficients. As a variant of our previous work by Mu and Zhang, the new method only needs access to the final linear system, that is, the global stiffness matrix and the right hand side, of a deterministic PDE solver, in order to build a domain‐decomposition‐based reduced model without intrusive implementation from scratch. The key idea is to remove the interface condition between sub‐domains and rely on the correlation between columns of the linear system to couple the sub‐domains. The non‐intrusive feature enables the applicability of the proposed method to a broader class of uncertainty quantification problems, where many legacy codes/solvers can be fully reused by our method. Two numerical examples including diffusion equations with random diffusivity and convection‐dominated transport with random velocity, are provided to demonstrate the effectiveness and efficiency of our method.

Zhang, Guannan↗

Multivariate Testing of Sampling Techniques to Address Class Imbalance in Building Use Type Classification

This study addresses the challenges inherent in building use type classification, particularly focusing on the issue of class imbalance in the training datasets for machine learning classifiers. We comprehensively analyze the efficacy of various class-balancing sampling techniques. Employing Monte Carlo simulations and Bayesian optimization, we evaluated the performance of multiple sampling methods, including Random Oversampling, Random Undersampling, SMOTE, Borderline-SMOTE, and ADASYN, across a dataset encompassing nine southeastern coastal states of the United States. Our findings reveal that simple random over- and undersampling techniques outperform more sophisticated methods. Additionally, we show inherent value in creating an imbalance in training data to effectively train a machine learning classifier for distinguishing between residential and nonresidential buildings. This study provides valuable guidance for future research on building use type classification research and lays essential groundwork for developing attribute-rich building stock datasets.

Adams, Daniel↗

A randomized sketching trust-region secant method for low-memory dynamic optimization

The numerical solution of dynamic optimization problems is often limited by the memory required to store the state trajectory, which is used to evaluate the objective function and its derivatives. Recently, [R. Muthukumar et al., SIAM Journal on Optimization 31(2), pp. 1242–1275 (2021)] introduced a trust-region method for dynamic optimization that employs randomized sketching to compress the state trajectory, resulting in inexact derivative computations. By adaptively learning the sketch rank, the trust-region algorithm achieves rigorous convergence guarantees. Here, we extend this approach to use secant Hessian approximations. Due to the randomness introduced by the sketch, the traditional secant update formulae can produce poor Hessian approximations. In particular, the difference of two gradients, computed from two different sketches, may be inconsistent. To overcome this, we employ a sketched approximation of the Hessian application, in lieu of computing the gradient difference. We numerically demonstrate the improved stability of this approach on an example from PDE-constrained optimization.

dynamic optimization↗

Adaptive Data-Driven Deep-Learning Surrogate Model for Frontal Polymerization in Dicyclopentadiene

Frontal polymerization (FP) is a self-sustaining curing process that enables rapid and energy-efficient manufacturing of thermoset polymers and composites. Computational methods conventionally used to simulate the FP process are time-consuming, and repeating simulations are required for sensitivity analysis, uncertainty quantification, or optimization of the manufacturing process. Here, in this work, we develop an adaptive surrogate deep-learning model for FP of dicyclopentadiene (DCPD), which predicts the evolution of temperature and degree of cure orders of magnitude faster than the finite-element method (FEM). The adaptive algorithm provides a strategy to select training samples efficiently and save computational costs by reducing the redundancy of FEM-based training samples. The adaptive algorithm calculates the residual error of the FP governing equations using automatic differentiation of the deep neural network. A probability density function expressed in terms of the residual error is used to select training samples from the Sobol sequence space. The temperature and degree of cure evolution of each training sample are obtained by a 2D FEM simulation. The adaptive method is more efficient and has a better prediction accuracy than the random sampling method. With the well-trained surrogate neural network, the FP characteristics (front speed, shape, and temperature) can be extracted quickly from the predicted temperature and degree-of-cure fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Stochastic Quasi-Newton Method in the Absence of Common Random Numbers

We present Q-SASS, a quasi-Newton method for unconstrained stochastic optimization that does not rely on common random numbers. Most existing quasi-Newton approaches leverage common random numbers to construct second-order updates. However, motivated by challenges in variational quantum algorithms—where such coordination is not possible—we consider the setting in which function values and gradients are accessible only through noisy probabilistic zeroth- and first-order oracles, and no common random numbers can be exploited. We derive high-probability tail bounds on the iteration complexity of our algorithm for nonconvex, convex, and strongly convex (more generally, those satisfying the PL condition) objective functions. Finally, we demonstrate the empirical benefits of our quasi-Newton updating scheme on both synthetic and quantum chemistry problems.

Complexity bound↗

A Randomization-Based, Zero-Trust Cyberattack Detection Method for Hierarchical Systems

This paper demonstrates a novel randomization-based approach for verifying power system control signals with application to detecting cyberattacks. We consider fully connected hierarchical systems containing multiple local agents and a global "trust" agent. The global agent uses a time-varying randomized assignment scheme to identify corrupt network links based on principles of zero trust and majority rule. To evaluate the performance of this detection approach, we implement our algorithm in MATLAB and run it against nearly 43 million unique attack scenarios spanning a range of system sizes. For each scenario, the algorithm determines whether the identified corruptions satisfy a set of validity constraints reflecting network topology and uses that result to say whether the recovered state value for one or more local agents is malicious. We compare the algorithm's determination to the true state of the system to assess performance and find that classification accuracy converges to 100% as system size increases, suggesting that the validity constraints become more difficult to satisfy for larger systems. We further explore the scenarios that evade detection to understand practical implications for employing this detection approach.

cybersecurity↗

Fast yaw optimization for wind plant wake steering using Boolean yaw angles

Abstract. In wind plants, turbines can be yawed into the wind to steer their wakes away from downstream turbines and achieve an overall increase in plant power. Mathematical optimization is typically used to determine the best yaw angles at which to operate the turbines in a plant. In this paper, we present a new heuristic to rapidly determine the yaw angles in a wind plant. In this method, we define the turbine yaw angles as Boolean – either yawed at a predefined angle or nonyawed – as opposed to the typical methods of defining yaw angles as continuous or with fine discretizations. We then optimize which turbines should be yawed with an algorithm that sweeps through the turbines from the most upstream to the most downstream. We demonstrate that our new Boolean optimization method can find turbine yaw angles that perform well compared to a traditionally used gradient-based optimizer for which the yaw angles are defined as continuous. There is less than 0.6 % difference in the optimized power between the two optimization methods for randomly placed turbine layouts and less than a 0.6 % difference in the optimal annual energy production between the two optimization methods for a real wind farm. Additionally, we show that our new method is much more computationally efficient than the traditional method. For plants with nonzero optimal yaw angles, our new method is generally able to solve for the turbine yaw angles 50–150 times faster, and in some extreme cases up to 500 times faster, than the traditional method.

17 WIND ENERGY↗