Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “randomized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Using soil library hyperspectral reflectance and machine learning to predict soil organic carbon: Assessing potential of airborne and spaceborne optical soil sensing

Soil organic carbon (SOC) is a key variable to determine soil functioning, ecosystem services, and global carbon cycles. Spectroscopy, particularly optical hyperspectral reflectance coupled with machine learning, can provide rapid, efficient, and cost-effective quantification of SOC. However, how to exploit soil hyperspectral reflectance to predict SOC concentration, and the potential performance of airborne and satellite data for predicting surface SOC at large scales remain relatively underknown. Here, this study utilized a continental-scale soil laboratory spectral library (37,540 full-pedon 350–2500 nm reflectance spectra with SOC concentration of 0–780 g·kg –1 across the US) to thoroughly evaluate seven machine learning algorithms including Partial-Least Squares Regression (PLSR), Random Forest (RF), K-Nearest Neighbors (KNN), Ridge, Artificial Neural Networks (ANN), Convolutional Neural Networks (CNN), and Long Short-Term Memory (LSTM) along with four preprocessed spectra, i.e. original, vector normalization, continuum removal, and first-order derivative, to quantify SOC concentration. Furthermore, by using the coupled soil-vegetation-atmosphere radiative transfer model, we simulated twelve airborne and spaceborne hyper/multi-spectral remote sensing data from surface bare soil laboratory spectra to evaluate their potential for estimating SOC concentration of surface bare soils. Results show that LSTM achieved best predictive performance of quantifying SOC concentration for the whole data sets (R 2 = 0.96, RMSE = 30.81 g·kg –1 ), mineral soils (SOC ≤ 120 g·kg –1 , R 2 = 0.71, RMSE = 10.60 g·kg –1 ), and organic soils (SOC > 120 g·kg –1 , R 2 = 0.78, RMSE = 62.31 g·kg –1 ). Spectral data preprocessing, particularly the first-order derivative, improved the performance of PLSR, RF, Ridge, KNN, and ANN, but not LSTM or CNN. We found that the SOC models of mineral and organic soils should be distinguished given their distinct spectral signatures. Finally, we identified that the shortwave infrared is vital for airborne and spaceborne hyperspectral sensors to monitor surface SOC. This study highlights the high accuracy of LSTM with hyperspectral/multispectral data to mitigate a certain level of noise (soil moisture <0.4 m 3 ·m –3 , green leaf area < 0.3 m 2 ·m –2 , plant residue <0.4 m 2 ·m –2 ) for quantifying surface SOC concentration. Forthcoming satellite hyperspectral missions like Surface Biology and Geology (SBG) have a high potential for future global soil carbon monitoring, while high-resolution satellite multispectral fusion data can be an alternative.

54 ENVIRONMENTAL SCIENCES↗

Measuring photometric redshifts for high-redshift radio source surveys

With the advent of deep, all-sky radio surveys, the need for ancillary data to make the most of the new, high-quality radio data from surveys like the Evolutionary Map of the Universe (EMU), GaLactic and Extragalactic All-sky Murchison Widefield Array survey eXtended, Very Large Array Sky Survey, and LOFAR Two-metre Sky Survey is growing rapidly. Radio surveys produce significant numbers of Active Galactic Nuclei (AGNs) and have a significantly higher average redshift when compared with optical and infrared all-sky surveys. Thus, traditional methods of estimating redshift are challenged, with spectroscopic surveys not reaching the redshift depth of radio surveys, and AGNs making it difficult for template fitting methods to accurately model the source. Machine Learning (ML) methods have been used, but efforts have typically been directed towards optically selected samples, or samples at significantly lower redshift than expected from upcoming radio surveys. This work compiles and homogenises a radio-selected dataset from both the northern hemisphere (making use of Sloan Digital Sky Survey optical photometry) and southern hemisphere (making use of Dark Energy Survey optical photometry). We then test commonly used ML algorithms such as k-Nearest Neighbours (kNN), Random Forest, ANNz, and GPz on this monolithic radio-selected sample. We show that kNN has the lowest percentage of catastrophic outliers, providing the best match for the majority of science cases in the EMU survey. We note that the wider redshift range of the combined dataset used allows for estimation of sources up to z = 3 before random scatter begins to dominate. When binning the data into redshift bins and treating the problem as a classification problem, we are able to correctly identify ≈ 76% of the highest redshift sources—sources at redshift z > 2.51 —as being in either the highest bin (z > 2.51) or second highest (z = 2.25).

79 ASTRONOMY AND ASTROPHYSICS↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Simultaneous quantification of uranium( VI ), samarium, nitric acid, and temperature with combined ensemble learning, laser fluorescence, and Raman scattering for real-time monitoring

In this work, laser-induced fluorescence spectroscopy (LIFS), Raman spectroscopy, and a stacked regression ensemble was developed for near real-time quantification of uranium(VI) (1–100 μg mL –1 ), samarium (0–200 μg mL –1 ) and nitric acid (0.1–4 M) with varying temperature (20 °C–45 °C). LIFS applications range from fundamental lab-scale studies to real-time process monitoring at industrial levels, such as nuclear reprocessing applications, provided the phenomena affecting the fluorescence spectrum are accounted for (e.g., absorption, quenching, complexation). Multiple chemometric models were examined and compared to a more traditional multivariate regression approach called partial least squares (PLS). Results obtained on synthetic samples selected using D-optimal experimental design indicated that a stacked regression method, which included ridge regression, random forest, PLS, and an eXtreme gradient boost algorithm, successfully measured uranium(VI) concentrations directly in nitric acid without measuring luminescence lifetimes or standard addition. The top model resulted in percent root-mean-square error of prediction values of 5.2, 1.9, 3.0, and 2.3% for U(VI), Sm 3+ , HNO 3 , and temperature, respectively. The approach may be useful for quantifying fluorescent fission products (e.g., Sm 3+ ) to provide information on burnup of irradiated nuclear fuel. This novel framework reinforces the applicability of LIFS for real-time applications in nuclear fuel cycle applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Dynamic Mode Decomposition of Random Pressure Fields over Bluff Bodies

Fluctuating surface pressures on a bluff body exposed to a boundary layer flow generally are characterized as a spatiotemporally varying random field. In this paper, a dynamic mode decomposition (DMD) was applied to extract dominant features embedded in these random pressure fields. Utilizing an unsupervised machine learning algorithm, spatial modes and their temporal variations were grouped into different clusters at scales, e.g., macro, meso, and micro. A proper orthogonal decomposition (POD) of the experimental data was carried out to observe commonalities and distinctive perspectives each decomposition offers. Here, a comprehensive examination of the DMD/POD for their convergence criteria, data sufficiency, and modal components analysis was conducted. The physical interpretation of the spatiotemporal pressure field based on these decomposition schemes was discussed. At different scales, the DMD modes can capture the evolution of aerodynamic features, e.g., convection of vortices (or vortex tubes) and other structures. The distribution of energy among these three broad scales also reflects an energy cascade in pressure fluctuations akin to turbulence.

97 MATHEMATICS AND COMPUTING↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

Monitoring Operational States of a Nuclear Reactor Using Seismoacoustic Signatures and Machine Learning

Monitoring nuclear reactors is an important safety and security task with growing requirements. We explore the possibility of using seismic and acoustic data for inferring the power level of an operating reactor. Continuous data recorded at a single seismoacoustic station that is located about 50 m away from a research reactor was visualized and analyzed. The data show a clear correlation between seismoacoustic features and reactor main operational states. We designed a workflow that includes two machine learning (ML) models to classify the reactor operational states (OFF, transition, and ON) and estimate reactor power levels (10%, 30%, 50%, 70%, and 90%). We applied and compared five ML algorithms for the reactor OFF-transition-ON and four approaches for the power level classification. We also compared the performance of ML models trained with seismic-only, acoustic-only, and both types of data. Five-fold cross validations were implemented to assure a thorough evaluation of the model performances. Additionally, the results show the extreme boosting gradient algorithm worked best for the first model, whereas random forests performed best for the second model. Combining seismic and acoustic data leads to better performance than using a single type of data. Seismic data contributed more than acoustic data for both models. We reached an accuracy of 0.98 for reactor OFF and ON. The accuracies for the transition state and power levels are less optimal with a minimum accuracy of 0.66. However, our results suggest seismic and acoustic data contain useful information about the transition state as well as power levels. Seismic and acoustic data could be integrated with other observations to improve monitoring performance.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Automatic Waveform Quality Control for Surface Waves Using Machine Learning

Surface-wave seismograms are widely used by researchers to study Earth’s interior and earthquakes. To extract information reliably and robustly from a suite of surface waveforms, the signals require quality control screening to reduce artifacts from signal complexity and noise. This process has usually been completed by human experts labeling each waveform visually, which is time consuming and tedious for large data sets. We explore automated approaches to improve the efficiency of waveform quality control processing by investigating logistic regression, support vector machines, K-nearest neighbors, random forests (RF), and artificial neural networks (ANN) algorithms. To speed up signal quality assessment, we trained these five machine learning (ML) methods using nearly 400,000 human-labeled waveforms. The ANN and RF models outperformed other algorithms and achieved a test accuracy of 92%. We evaluated these two best-performing models using seismic events from geographic regions not used for training. The results show that the two trained models agree with labels from human analysts but required only 0.4% of the time. Although the original (human) quality assignments assessed general waveform signal-to-noise, the ANN or RF labels can help facilitate detailed waveform analysis. Our investigations demonstrate the capability of the automated processing using these two ML models to reduce outliers in surface-wave-related measurements without human quality control screening.

58 GEOSCIENCES↗

Fast tensor disentangling algorithm

Many recent tensor network algorithms apply unitary operators to parts of a tensor network in order to reduce entanglement. However, many of the previously used iterative algorithms to minimize entanglement can be slow. We introduce an approximate, fast, and simple algorithm to optimize disentangling unitary tensors. Our algorithm is asymptotically faster than previous iterative algorithms and often results in a residual entanglement entropy that is within 10 to 40% of the minimum. For certain input tensors, our algorithm returns an optimal solution. When disentangling order-4 tensors with equal bond dimensions, our algorithm achieves an entanglement spectrum where nearly half of the singular values are zero. We further validate our algorithm by showing that it can efficiently disentangle random 1D states of qubits.

Slagle, Kevin↗

A Comparison of Machine Learning Methods to Forecast Tropospheric Ozone Levels in Delhi

Ground-level ozone is a pollutant that is harmful to urban populations, particularly in developing countries where it is present in significant quantities. It greatly increases the risk of heart and lung diseases and harms agricultural crops. This study hypothesized that, as a secondary pollutant, ground-level ozone is amenable to 24 h forecasting based on measurements of weather conditions and primary pollutants such as nitrogen oxides and volatile organic compounds. We developed software to analyze hourly records of 12 air pollutants and 5 weather variables over the course of one year in Delhi, India. To determine the best predictive model, eight machine learning algorithms were tuned, trained, tested, and compared using cross-validation with hourly data for a full year. The algorithms, ranked by R2 values, were XGBoost (0.61), Random Forest (0.61), K-Nearest Neighbor Regression (0.55), Support Vector Regression (0.48), Decision Trees (0.43), AdaBoost (0.39), and linear regression (0.39). When trained by separate seasons across five years, the predictive capabilities of all models increased, with a maximum R 2 of 0.75 during winter. Bidirectional Long Short-Term Memory was the least accurate model for annual training, but had some of the best predictions for seasonal training. Out of five air quality index categories, the XGBoost model was able to predict the correct category 24 h in advance 90% of the time when trained with full-year data. Separated by season, winter is considerably more predictable (97.3%), followed by post-monsoon (92.8%), monsoon (90.3%), and summer (88.9%). These results show the importance of training machine learning methods with season-specific data sets and comparing a large number of methods for specific applications.

54 ENVIRONMENTAL SCIENCES↗

An inexact semismooth Newton method with application to adaptive randomized sketching for dynamic optimization

In many applications, one can only access the inexact gradients and inexact hessian times vector products. Thus it is essential to consider algorithms that can handle such inexact quantities with a guaranteed convergence to solution. An inexact adaptive and provably convergent semismooth Newton method is considered to solve constrained optimization problems. In particular, dynamic optimization problems, which are known to be highly expensive, are the focus. A memory efficient semismooth Newton algorithm is introduced for these problems. The source of efficiency and inexactness is the randomized matrix sketching. Further, applications to optimization problems constrained by partial differential equations are also considered.

97 MATHEMATICS AND COMPUTING↗

Technical note: Uncertainties in eddy covariance CO 2 fluxes in a semiarid sagebrush ecosystem caused by gap-filling approaches

Abstract. Gap-filling eddy covariance CO2 fluxes is challenging at dryland sites due to small CO2 fluxes. Here, four machine learning (ML) algorithms including artificial neural network (ANN), k-nearest neighbors (KNNs), random forest (RF), and support vector machine (SVM) are employed and evaluated for gap-filling CO2 fluxes over a semiarid sagebrush ecosystem with different lengths of artificial gaps. The ANN and RF algorithms outperform the KNN and SVM in filling gaps ranging from hours to days, with the RF being more time efficient than the ANN. Performances of the ANN and RF are largely degraded for extremely long gaps of 2 months. In addition, our results suggest that there is no need to fill the daytime and nighttime net ecosystem exchange (NEE) gaps separately when using the ANN and RF. With the ANN and RF, the gap-filling-induced uncertainties in the annual NEE at this site are estimated to be within 16 g C m−2, whereas the uncertainties by the KNN and SVM can be as large as 27 g C m−2. To better fill extremely long gaps of a few months, we test a two-layer gap-filling framework based on the RF. With this framework, the model performance is improved significantly, especially for the nighttime data. Therefore, this approach provides an alternative in filling extremely long gaps to characterize annual carbon budgets and interannual variability in dryland ecosystems.

Yao, Jingyu↗

Photometric redshift estimation of BASS DR3 quasars by machine learning

ABSTRACT Correlating Beijing–Arizona Sky Survey (BASS) data release 3 (DR3) catalogue with the ALLWISE data base, the data from optical and infrared information are obtained. The quasars from Sloan Digital Sky Survey are taken as training and test samples while those from LAMOST are considered as external test sample. We propose two schemes to construct the redshift estimation models with XGBoost, CatBoost, and Random Forest. One scheme (namely one-step model) is to predict photometric redshifts directly based on the optimal models created by these three algorithms; the other scheme (namely two-step model) is to first classify the data into low- and high-redshift data sets, and then predict photometric redshifts of these two data sets separately. For one-step model, the performance of these three algorithms on photometric redshift estimation is compared with different training samples, and CatBoost is superior to XGBoost and Random Forest. For two-step model, the performances of these three algorithms on the classification of low and high redshift subsamples are compared, and CatBoost still shows the best performance. Therefore, CatBoost is regarded as the core algorithm of classification and regression in two-step model. In contrast to one-step model, two-step model is optimal when predicting photometric redshift of quasars, especially for high-redshift quasars. Finally, the two models are applied to predict photometric redshifts of all quasar candidates of BASS DR3. The number of high-redshift quasar candidates is 3938 (redshift ≥3.5) and 121 (redshift ≥4.5) by two-step model. The predicted result will be helpful for quasar research and follow-up observation of high-redshift quasars.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-Driven Security Assessment of Power Grids Based on Machine Learning Approach: Preprint

Data-driven security assessment provides key indicators on power system stability using simulations on scheduling models, as opposed to dynamic simulations that are more time-consuming. This paper investigates data-driven security assessment of power grids based on machine learning. Multivariate random forest regression is used as the machine learning algorithm due to its high robustness to the input data. Three stability issues are analyzed using the proposed machine learning tool, including transient stability, frequency stability and small signal stability. The estimation values from machine learning tool are compared with those from dynamic simulations. Results show that the proposed machine learning tool can effectively predict the stability margins for the three stability metrics.

14 SOLAR ENERGY↗

Fouling modeling and prediction approach for heat exchangers using deep learning

In this article, we develop a generalized and scalable statistical model for accurate prediction of fouling resistance using commonly measured parameters of industrial heat exchangers. This prediction model is based on deep learning where a scalable algorithmic architecture learns non-linear functional relationships between a set of target and predictor variables from large number of training samples. Here, the efficacy of this modeling approach is demonstrated for predicting fouling in an analytically modeled cross-flow heat exchanger, designed for waste heat recovery from flue-gas using room temperature water. The performance results of the trained models demonstrate that the mean absolute prediction errors are under 10 –4 KW –1 for flue-gas side, water side and overall fouling resistances. The coefficients of determination (R 2 ), which characterize the goodness of fit between the predictions and observed data, are over 99%. Even under varying levels of measurement noise in the inputs, we demonstrate that predictions over an ensemble of multiple neural networks achieves better accuracy and robustness to noise. We find that the proposed deep-learning fouling prediction framework learns to follow heat exchanger flow and heat transfer physics, which we confirm using locally interpretable model agnostic explanations around randomly selected operating points. Overall, we provide a robust algorithmic framework for fouling prediction that can be generalized and scaled to various types of industrial heat exchangers.

42 ENGINEERING↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Post-extreme-event restoration using linear topological constraints and DER scheduling to enhance distribution system resilience

In this paper, a post-extreme-event restoration (PEER) algorithm is proposed to improve distribution system resilience. Linear topological constraints are proposed to ensure radial topology after N-k contingencies, possibly in multiple islands. The approach is made comprehensive by considering dispatchable distributed energy resources (DERs), non-dispatchable DERs, and demand responses, as well as on-load tap changers (OLTCs) and shunt capacitors. The goal is to minimize the accumulative expense caused by load reduction payment or penalty, as well as DER operation cost. As a result, the overall system will survive longer with higher resilience during an extreme event. To verify the effectiveness of the PEER algorithm, we proposed a resilience evaluation algorithm using Monte Carlo simulation (MCS) with reduced scenarios. This is based on a probabilistic model for generating random scenarios which consider the uncertainty of line faults and solar irradiance. Combined with the proposed PEER algorithm, this reduced-scenario MCS can evaluate the expected energy not served (EENS) which is an essential index for distribution system resilience. Case studies of the IEEE 33-bus and 123-bus test systems validate the proposed algorithm in reducing EENS.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Finding diverse ways to improve algebraic connectivity through multi-start optimization

The algebraic connectivity, also known as the Fiedler value, is a spectral measure of network connectivity that can be increased through edge addition. We present an algorithm for producing many diverse ways to add a fixed number of edges to a network to achieve a near optimal Fiedler value. Previous Fielder value optimization algorithms (i.e. the greedy algorithm) output only one solution. Obtaining a single solution is rarely good enough for real-world network redesign problems, as practical constraints (political, physical or financial) may prevent implementation. Our algorithm takes a multi-start optimization approach, adding a random initial edge and then applies a greedy heuristic to improve the Fiedler value. The random choice moves us to a new region of the search space, enabling discovery of diverse solutions. Additionally, we present a Determinantal Point Process framework for quantifying diversity. We then apply a Markov chain Monte Carlo technique to sift through the large number of output solutions and locate a smaller, more manageable collection of highly diverse solutions that can be presented to network redesign engineers. We demonstrate the effectiveness of our algorithm on real-world graphs with varied structures.

97 MATHEMATICS AND COMPUTING↗