Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

An inexact semismooth Newton method with application to adaptive randomized sketching for dynamic optimization

In many applications, one can only access the inexact gradients and inexact hessian times vector products. Thus it is essential to consider algorithms that can handle such inexact quantities with a guaranteed convergence to solution. An inexact adaptive and provably convergent semismooth Newton method is considered to solve constrained optimization problems. In particular, dynamic optimization problems, which are known to be highly expensive, are the focus. A memory efficient semismooth Newton algorithm is introduced for these problems. The source of efficiency and inexactness is the randomized matrix sketching. Further, applications to optimization problems constrained by partial differential equations are also considered.

97 MATHEMATICS AND COMPUTING↗

Adding GPU Support to the Markov Chain Monte Carlo Code Catmip

In geophysics, we are confronted with many under-determined inverse problems. For example, all of our observations of earthquakes are made at the Earth’s surface. So, when we try to infer how slip during an earthquake evolves in space and time, we find that there are many potential slip histories that are consistent with our limited observations and our understanding of earthquake physics. One way to approach these problems is with Bayesian analysis which allows us to infer the ensemble of all potential slip models that satisfy the observations and our prior knowledge of earthquake physics. In Bayesian analysis, our prior knowledge is known as the prior probability density function or prior PDF, the fit to the data is known as the data likelihood, and the target PDF that satisfies both the prior PDF and data likelihood is known as the posterior PDF. However, simulating the posterior PDF typically requires using Markov Chain Monte Carlo (MCMC) to draw tens of billions of random realizations of earthquake slip models, which may not be computationally feasible. To make this and similar geophysical inversions computationally tractable, we developed the Cascading Adaptive Transitional Metropolis In Parallel (CATMIP) algorithm. CATMIP is an efficient parallel Markov Chain Monte Carlo (MCMC) sampler that is used for model fitting and uncertainty quantification in geophysics. Example use cases are earthquake rupture modeling, determining mineral composition on Mars, reconstructing the history of ocean salinity, and historical earthquake relocation. CATMIP employs many parallel instances of the Metropolis algorithm for sampling in a transitioning framework. Transitioning is a process in which a set of random samples at equilibrium with a known probability density function (PDF) are used as seeds for the Markov chains to sample successive target PDFs that incrementally move the distribution from the starting seeds to the final desired PDF that describes the relative plausibility of potential values for the model parameters. The algorithm is implemented as a Master-Worker model employing MPI for communication. The worker processes are loosely coupled with global parameters periodically optimized by the master process. This provides a very high amount of parallelism with little communication between updates. During the presentation we will discuss the history of the algorithm and elaborate the earthquake rupture modeling use case for the CATMIP package. Our first step toward GPU optimization was to optimize the code for the CPU. CPU profiling revealed that most of the compute time is spent in calls to level 2 BLAS routines and calls to GSL random number generators. We revised the algorithm to employ level 3 BLAS routines instead. In our presentation we will describe how this was accomplished. Adding GPU support to CATMIP consisted mostly of replacing the calls to GSL with calls to GPU vendor-provided library routines. A small number of loops were directly implemented in CUDA. In the presentation will provide implementation details. Finally, we will discuss methods for profiling and opportunities for further optimizing GPU execution. By creating a code with the flexibility to run on either a CPU or GPU architecture, CATMIP can be used on systems ranging from large CPU-based HPC environments to single servers with GPU acceleration and everything in between.

HECC↗

Topology of large-scale structure. IV - Topology in two dimensions

In a recent series of papers, an algorithm was developed for quantitatively measuring the topology of the large-scale structure of the universe and this algorithm was applied to numerical models and to three-dimensional observational data sets. In this paper, it is shown that topological information can be derived from a two-dimensional cross section of a density field, and analytic expressions are given for a Gaussian random field. The application of a two-dimensional numerical algorithm for measuring topology to cross sections of three-dimensional models is demonstrated.

Melott, Adrian L.↗

Technical note: Uncertainties in eddy covariance CO 2 fluxes in a semiarid sagebrush ecosystem caused by gap-filling approaches

Abstract. Gap-filling eddy covariance CO2 fluxes is challenging at dryland sites due to small CO2 fluxes. Here, four machine learning (ML) algorithms including artificial neural network (ANN), k-nearest neighbors (KNNs), random forest (RF), and support vector machine (SVM) are employed and evaluated for gap-filling CO2 fluxes over a semiarid sagebrush ecosystem with different lengths of artificial gaps. The ANN and RF algorithms outperform the KNN and SVM in filling gaps ranging from hours to days, with the RF being more time efficient than the ANN. Performances of the ANN and RF are largely degraded for extremely long gaps of 2 months. In addition, our results suggest that there is no need to fill the daytime and nighttime net ecosystem exchange (NEE) gaps separately when using the ANN and RF. With the ANN and RF, the gap-filling-induced uncertainties in the annual NEE at this site are estimated to be within 16 g C m−2, whereas the uncertainties by the KNN and SVM can be as large as 27 g C m−2. To better fill extremely long gaps of a few months, we test a two-layer gap-filling framework based on the RF. With this framework, the model performance is improved significantly, especially for the nighttime data. Therefore, this approach provides an alternative in filling extremely long gaps to characterize annual carbon budgets and interannual variability in dryland ecosystems.

Yao, Jingyu↗

Photometric redshift estimation of BASS DR3 quasars by machine learning

ABSTRACT Correlating Beijing–Arizona Sky Survey (BASS) data release 3 (DR3) catalogue with the ALLWISE data base, the data from optical and infrared information are obtained. The quasars from Sloan Digital Sky Survey are taken as training and test samples while those from LAMOST are considered as external test sample. We propose two schemes to construct the redshift estimation models with XGBoost, CatBoost, and Random Forest. One scheme (namely one-step model) is to predict photometric redshifts directly based on the optimal models created by these three algorithms; the other scheme (namely two-step model) is to first classify the data into low- and high-redshift data sets, and then predict photometric redshifts of these two data sets separately. For one-step model, the performance of these three algorithms on photometric redshift estimation is compared with different training samples, and CatBoost is superior to XGBoost and Random Forest. For two-step model, the performances of these three algorithms on the classification of low and high redshift subsamples are compared, and CatBoost still shows the best performance. Therefore, CatBoost is regarded as the core algorithm of classification and regression in two-step model. In contrast to one-step model, two-step model is optimal when predicting photometric redshift of quasars, especially for high-redshift quasars. Finally, the two models are applied to predict photometric redshifts of all quasar candidates of BASS DR3. The number of high-redshift quasar candidates is 3938 (redshift ≥3.5) and 121 (redshift ≥4.5) by two-step model. The predicted result will be helpful for quasar research and follow-up observation of high-redshift quasars.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-Driven Security Assessment of Power Grids Based on Machine Learning Approach: Preprint

Data-driven security assessment provides key indicators on power system stability using simulations on scheduling models, as opposed to dynamic simulations that are more time-consuming. This paper investigates data-driven security assessment of power grids based on machine learning. Multivariate random forest regression is used as the machine learning algorithm due to its high robustness to the input data. Three stability issues are analyzed using the proposed machine learning tool, including transient stability, frequency stability and small signal stability. The estimation values from machine learning tool are compared with those from dynamic simulations. Results show that the proposed machine learning tool can effectively predict the stability margins for the three stability metrics.

14 SOLAR ENERGY↗

Fouling modeling and prediction approach for heat exchangers using deep learning

In this article, we develop a generalized and scalable statistical model for accurate prediction of fouling resistance using commonly measured parameters of industrial heat exchangers. This prediction model is based on deep learning where a scalable algorithmic architecture learns non-linear functional relationships between a set of target and predictor variables from large number of training samples. Here, the efficacy of this modeling approach is demonstrated for predicting fouling in an analytically modeled cross-flow heat exchanger, designed for waste heat recovery from flue-gas using room temperature water. The performance results of the trained models demonstrate that the mean absolute prediction errors are under 10 –4 KW –1 for flue-gas side, water side and overall fouling resistances. The coefficients of determination (R 2 ), which characterize the goodness of fit between the predictions and observed data, are over 99%. Even under varying levels of measurement noise in the inputs, we demonstrate that predictions over an ensemble of multiple neural networks achieves better accuracy and robustness to noise. We find that the proposed deep-learning fouling prediction framework learns to follow heat exchanger flow and heat transfer physics, which we confirm using locally interpretable model agnostic explanations around randomly selected operating points. Overall, we provide a robust algorithmic framework for fouling prediction that can be generalized and scaled to various types of industrial heat exchangers.

42 ENGINEERING↗

Determining the Number of Clusters in a Data Set Without Graphical Interpretation

Cluster analysis is a data mining technique that is meant ot simplify the process of classifying data points. The basic clustering process requires an input of data points and the number of clusters wanted. The clustering algorithm will then pick starting C points for the clusters, which can be either random spatial points or random data points. It then assigns each data point to the nearest C point where "nearest usually means Euclidean distance, but some algorithms use another criterion. The next step is determining whether the clustering arrangement this found is within a certain tolerance. If it falls within this tolerance, the process ends. Otherwise the C points are adjusted based on how many data points are in each cluster, and the steps repeat until the algorithm converges,

Aguirre, Nathan S.↗

Comparison of Cloud Detection Algorithms for Sentinel-2 Imagery

Accurate, automated cloud and cloud shadow detection is a key component of the processing needed to prepare optical satellite imagery for scientific analysis. Many existing cloud detection algorithms rely on temperature information to identify clouds, making detection difficult for imagers that lack a thermal band, like Sentinel-2. To get maximum benefit from Sentinel-2 products it is critical to understand which algorithms best identify clouds and their shadows in images. We examined the relative performance of five different cloud-masking algorithms (Sen2Cor, MAJA, LaSRC, Fmask and Tmask) in 6 Sentinel-2 scenes (28 total images) distributed across the Eastern Hemisphere. Expanding on these comparisons, we tested ensemble approaches to improve results. We tested three ensemble approaches to cloud and shadow classification based on the outputs of the five initial algorithms using the cloud masks in: (1) a majority prediction model; (2) a random forests model; and (3) a conditional logic model. Accuracy assessments show a trade-off between omission and commission errors in cloud detection for individual algorithms across all sites, and some algorithms are better at detecting either clouds or cloud shadows. No single algorithm outperforms the others for both clouds and shadows. Aggregating the results from multiple algorithms produces fewer undetected clouds and higher overall accuracy than any single algorithm, with as high as 2.7% improvement over the top-performing algorithm, suggesting an ensemble approach may be the most useful for processing of Sentinel-2 data.

Sentinel-2↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Estimating Species-Specific Leaf Area Index and Basal Area Using Optical and SAR Remote Sensing Data in Acadian Mixed Spruce-Fir Forests, USA

This study combined Sentinel-1 synthetic aperture radar (SAR), Sentinel-2 multispectral, and site variable datasets to model leaf area index (LAI) and basal area per ha (BAPH) of two economically important tree species in Northeast, USA; red spruce (Picea rubens Sarg.; RS), and balsam fir (Abies balsamea (L.) Mill.; BF). We used Random Forest (RF), and Multi-Layer Perceptron (MLP) algorithms for LAI and BAPH modeling. The results showed that RF outperformed MLP by reducing the normalized root mean square error (nRMSE) by 0.01 and 0.06 for LAI and BAPH, respectively. The final variables selected for modeling of both LAI and BAPH indicated the superiority of Sentinel-2 variables over the Sentinel-1 SAR with minor contributions of site variables (mainly elevation). The red-edge spectral vegetation indices played a significant role in both LAI and BAPH estimation. We attained the lowest nRMSEs of 0.12, and 0.16 for the final LAI model of RS, and BF, respectively using Sentinel-2 and site variables. The lowest nRMSE for both RS and BF BAPH models was 0.12. As RS and BF are the primary host species for a cyclically occurring and most destructive pest of the region, eastern spruce budworm (Choristoneura fumiferana; SBW), these estimations will be useful to evaluate SBW dynamics in the region.

Forest inventory↗

Computational methods for structural load and resistance modeling

An automated capability for computing structural reliability considering uncertainties in both load and resistance variables is presented. The computations are carried out using an automated Advanced Mean Value iteration algorithm (AMV +) with performance functions involving load and resistance variables obtained by both explicit and implicit methods. A complete description of the procedures used is given as well as several illustrative examples, verified by Monte Carlo Analysis. In particular, the computational methods described in the paper are shown to be quite accurate and efficient for a material nonlinear structure considering material damage as a function of several primitive random variables. The results show clearly the effectiveness of the algorithms for computing the reliability of large-scale structural systems with a maximum number of resolutions.

Thacker, B. H.↗

Post-extreme-event restoration using linear topological constraints and DER scheduling to enhance distribution system resilience

In this paper, a post-extreme-event restoration (PEER) algorithm is proposed to improve distribution system resilience. Linear topological constraints are proposed to ensure radial topology after N-k contingencies, possibly in multiple islands. The approach is made comprehensive by considering dispatchable distributed energy resources (DERs), non-dispatchable DERs, and demand responses, as well as on-load tap changers (OLTCs) and shunt capacitors. The goal is to minimize the accumulative expense caused by load reduction payment or penalty, as well as DER operation cost. As a result, the overall system will survive longer with higher resilience during an extreme event. To verify the effectiveness of the PEER algorithm, we proposed a resilience evaluation algorithm using Monte Carlo simulation (MCS) with reduced scenarios. This is based on a probabilistic model for generating random scenarios which consider the uncertainty of line faults and solar irradiance. Combined with the proposed PEER algorithm, this reduced-scenario MCS can evaluate the expected energy not served (EENS) which is an essential index for distribution system resilience. Case studies of the IEEE 33-bus and 123-bus test systems validate the proposed algorithm in reducing EENS.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Design and simulation of stratified probability digital receiver with application to the multipath communication

One approach to the problem of simplifying complex nonlinear filtering algorithms is through using stratified probability approximations where the continuous probability density functions of certain random variables are represented by discrete mass approximations. This technique is developed in this paper and used to simplify the filtering algorithms developed for the optimum receiver for signals corrupted by both additive and multiplicative noise.

Deal, J. H.↗

Finding diverse ways to improve algebraic connectivity through multi-start optimization

The algebraic connectivity, also known as the Fiedler value, is a spectral measure of network connectivity that can be increased through edge addition. We present an algorithm for producing many diverse ways to add a fixed number of edges to a network to achieve a near optimal Fiedler value. Previous Fielder value optimization algorithms (i.e. the greedy algorithm) output only one solution. Obtaining a single solution is rarely good enough for real-world network redesign problems, as practical constraints (political, physical or financial) may prevent implementation. Our algorithm takes a multi-start optimization approach, adding a random initial edge and then applies a greedy heuristic to improve the Fiedler value. The random choice moves us to a new region of the search space, enabling discovery of diverse solutions. Additionally, we present a Determinantal Point Process framework for quantifying diversity. We then apply a Markov chain Monte Carlo technique to sift through the large number of output solutions and locate a smaller, more manageable collection of highly diverse solutions that can be presented to network redesign engineers. We demonstrate the effectiveness of our algorithm on real-world graphs with varied structures.

97 MATHEMATICS AND COMPUTING↗

Automatic microseismic event picking via unsupervised machine learning

SUMMARY Effective and efficient arrival picking plays an important role in microseismic and earthquake data processing and imaging. Widely used short-term-average long-term-average ratio (STA/LTA) based arrival picking algorithms suffer from the sensitivity to moderate-to-strong random ambient noise. To make the state-of-the-art arrival picking approaches effective, microseismic data need to be first pre-processed, for example, removing sufficient amount of noise, and second analysed by arrival pickers. To conquer the noise issue in arrival picking for weak microseismic or earthquake event, I leverage the machine learning techniques to help recognizing seismic waveforms in microseismic or earthquake data. Because of the dependency of supervised machine learning algorithm on large volume of well-designed training data, I utilize an unsupervised machine learning algorithm to help cluster the time samples into two groups, that is, waveform points and non-waveform points. The fuzzy clustering algorithm has been demonstrated to be effective for such purpose. A group of synthetic, real microseismic and earthquake data sets with different levels of complexity show that the proposed method is much more robust than the state-of-the-art STA/LTA method in picking microseismic events, even in the case of moderately strong background noise.

Chen, Yangkang↗

Genetic Algorithm for Optimization of Neural Networks for Bayesian Inference of Model Uncertainty

The objective of this work was to develop a genetic optimization algorithm that can design a neural network capable of producing uncertainty estimates along with predictions. This algorithm is necessary because the inclusion of uncertainty modeling in a neural network greatly complicates the network’s design space, making the development of a converging model extremely difficult and time consuming. The genetic algorithm presented in this work uses a number of value ranges for various configurable neural network parameters to create a randomly generated population of network architectures. The initially generated population is then evolved over the course of several generations, with the best performing models breeding to produce novel network configurations. Mutations are randomly applied to the network designs to facilitate the development of adaptations beneficial to the task being performed. An experiment was conducted to validate the proposed algorithm, in which the genetic optimizer was tasked with producing a neural network capable of predicting the sound pressure level (SPL) resulting from jet-surface interaction (JSI) noise. The data used for this task was generated at the NASA Glenn Research Center in the Aero-Acoustic Propulsion Laboratory. Starting with an initial population size of 35 randomly generated networks, and evolved over the course of 10 generations, the genetic algorithm produced a design able to predict SPL as a result of JSI noise within 0.272 dB, on average.

Genetic algorithm↗

A Stochastic Quasi-Newton Method in the Absence of Common Random Numbers

We present Q-SASS, a quasi-Newton method for unconstrained stochastic optimization that does not rely on common random numbers. Most existing quasi-Newton approaches leverage common random numbers to construct second-order updates. However, motivated by challenges in variational quantum algorithms—where such coordination is not possible—we consider the setting in which function values and gradients are accessible only through noisy probabilistic zeroth- and first-order oracles, and no common random numbers can be exploited. We derive high-probability tail bounds on the iteration complexity of our algorithm for nonconvex, convex, and strongly convex (more generally, those satisfying the PL condition) objective functions. Finally, we demonstrate the empirical benefits of our quasi-Newton updating scheme on both synthetic and quantum chemistry problems.

Complexity bound↗