Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Support Vector Machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Analysis and prediction of intersection traffic violations using automated enforcement system data

We report that the automated enforcement system (AES) is an effective way of supplementing traditional traffic enforcement, and the traffic violation data from AES can also be effectively used for safety research. In this study, traffic violation data were used to analyze the influencing factors associated with traffic violations and to predict the probability of violations at intersections. The potential factors influencing violations include 24 independent factors related to time, space, traffic and weather. Results from a logistic model showed that the midday period, weekends, residential districts, collector roads, congested traffic conditions, high traffic flow, lower wind speed and low temperature would increase the probability of traffic violations. The probability of violations was predicted by the random forest algorithm, which was proven to be the best traffic violation prediction model among logistic regression, Gaussian naive Bayes, and support vector machine. Moreover, the proximity weighted synthetic oversampling technique (ProWSyn) method was applied to reduce the impact of the imbalance ratio (IR) and improve the model’s prediction performance. The receiver operating characteristics (ROC) curves and Precision-Recall (PR) curves illustrated that the random forest algorithm using oversampling data had the best classifier prediction performance than undersampling data. The area under curve (AUC) and out-of-bag (OOB) error with IR = 1 reached 0.914 and 0.0787, which showed the better performance of the random forest algorithm using ProWSyn in dealing with imbalanced traffic violation data.

42 ENGINEERING↗

Time series anomaly detection in power electronics signals with recurrent and ConvLSTM autoencoders

The anomalies in the high voltage converter modulator (HVCM) remain a major down time for the spallation neutron source facility, that delivers the most intense neutron beam in the world for scientific materials research. In this work, we propose neural network architectures based on Recurrent AutoEncoders (RAE) to detect anomalies ahead of time in the power signals coming from the HVCM. Bi-directional gated recurrent unit, bi-directional long-short term memory (LSTM), and convolutional LSTM (ConvLSTM) are developed, trained, and tested using real experimental signals from the HVCM module. The results show a good performance of the proposed RAE models, achieving precision up to 91%, recall up to 88%, false omission rate as low as 20% (i.e. 80% of the anomalies were detected), and area under the ROC curve up to 0.9. The three RAE models provide very comparable performance, with LSTM showing slightly better performance than GRU and ConvLSTM. The RAE models are benchmarked against other anomaly detection methods, including isolation forest, support vector machine, local outlier factor, feedforward and convolutional autoencoders, and others; showing a better performance. Here, the results of this study demonstrate the promising potential of RAE in anomaly detection for real-world power systems, and for increasing the reliability of the HVCM modules in the spallation neutron source.

42 ENGINEERING↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Applying NIR and MIR spectroscopy for C and soil property prediction in northern cold-region ecosystems. Which approach works better?

Here, developing reliable predictions of soil attributes is necessary to understand northern cold-region climate-soil feedback. Calibration models using near-infrared (NIR) and mid-infrared (MIR) spectroscopy were developed to predict eight commonly measured soil properties for 119 soil samples representing a range of vegetation types, parent materials, and soil types spanning >23° of latitude from southeast Alaska to the Canadian high Arctic. In order to obtain a more accurate prediction, this study compared the performance of linear and non-linear calibration techniques, including lasso regression (Lasso), support vector machine (SVM), random forest (RF) and classic partial least squares (PLS) to predict different soil properties of these soils. Comparing the four models, we noticed that their performance was quite similar for MIR overall, while NIR achieved better results with a PLS model for our dataset. PLS coupled with MIR showed a better performance for soil parameters, such as total organic carbon (TOC), total nitrogen (TN), cation exchange capacity (CEC) and clay (R-squared of 0.9, 0.81, 0.80, and 0.84) when compared with NIR (R-squared of 0.85, 0.72, 0.81 and 0.68). However, using either MIR or NIR spectroscopy, PLS predictions for bulk density (BD) and sand content were not accurate. The variable importance analysis based on the PLS model successfully estimated the relative contribution of wavelengths influencing soil property predictions most. Overall, TOC, TN, CEC and clay mineral predictions are closely related to the occurrence of specific spectral bands in the MIR region. For example, wavelengths at 2978 and 1761 cm -1 for TOC and TN, as well as at 3064 cm -1 for CEC, were selected as the most influential predictor variables. We demonstrated that MIR spectroscopy is a powerful tool for more extensive monitoring in soils of the northern cold climate region; however, NIR could be utilized for rapid estimates when the highest accuracy is not essential.

54 ENVIRONMENTAL SCIENCES↗

Cause identification of electromagnetic transient events using spatiotemporal feature learning

This paper presents a spatiotemporal feature learning method for cause identification of electromagnetic transient events in power grids. The proposed method is formulated based on the availability of time-synchronized high-frequency measurements and using the convolutional neural network as the spatiotemporal feature representation along with softmax function for the classification. Despite the existing threshold-based, or energy-based events analysis methods, such as support vector machine autoencoder, and tapered multi-layer perceptron neural network, the proposed feature learning is carried out with respect to both time and space. The effectiveness of the proposed feature learning and the subsequent cause identification is validated through the Electromagnetic Transients Program (EMTP) simulation of different events such as line energization, capacitor bank energization, lightning, fault, and high-impedance fault in the IEEE 30-bus, and the real-time digital simulation of the Western System Coordinating Council (WSCC) 9-bus system.

42 ENGINEERING↗

Predicting melt pool depth and grain length using multiple signatures from in-situ single camera two-wavelength imaging pyrometry for laser powder bed fusion

In laser powder bed fusion (LPBF), the in-situ process signatures are known to have a direct correlation with the microstructural properties of the solidified melt pool (MP). It is known that the MP cooling and heating rates, and laser processing parameters can critically determine the grain structure and thereby affect the part properties. The objective of this work is to study the feasibility of using in-process, high-speed imaging pyrometry for evaluating the solidified MP properties “below” the surface, such as depth and microstructural properties. To accomplish this, we employ an in-house single camera-based two-wavelength imaging pyrometry (STWIP) system for monitoring the printing of single-scan tracks with Inconel 718 on a commercial LPBF printer (EOS M290). Further, the lab designed STWIP system is a coaxial high-speed (>10,000 fps) imaging system capable of monitoring MP temperature, morphology, and intensity profiles. The temperature measurements from STWIP are emissivity independent. The STWIP measured MP signatures of the printed tracks are correlated with the ex-situ microscopy characterized MP depth and the average grain lengths. From the data analysis, using support vector machine (SVM)-based regression models, we found that the MP temperature signatures are crucial for an accurate prediction of MP depth and the grain length, thus validating the novelty and necessity of the developed in-situ monitoring methods and analysis.

36 MATERIALS SCIENCE↗

BigNeuron: a resource to benchmark and predict performance of algorithms for automated tracing of neurons in light microscopy datasets

BigNeuron is an open community bench-testing platform with the goal of setting open standards for accurate and fast automatic neuron tracing. We gathered a diverse set of image volumes across several species that is representative of the data obtained in many neuroscience laboratories interested in neuron tracing. Here, we report generated gold standard manual annotations for a subset of the available imaging datasets and quantified tracing quality for 35 automatic tracing algorithms. The goal of generating such a hand-curated diverse dataset is to advance the development of tracing algorithms and enable generalizable benchmarking. Together with image quality features, we pooled the data in an interactive web application that enables users and developers to perform principal component analysis, t-distributed stochastic neighbor embedding, correlation and clustering, visualization of imaging and tracing data, and benchmarking of automatic tracing algorithms in user-defined data subsets. The image quality metrics explain most of the variance in the data, followed by neuromorphological features related to neuron size. Furthermore, we observed that diverse algorithms can provide complementary information to obtain accurate results and developed a method to iteratively combine methods and generate consensus reconstructions. The consensus trees obtained provide estimates of the neuron structure ground truth that typically outperform single algorithms in noisy datasets. However, specific algorithms may outperform the consensus tree strategy in specific imaging conditions. Finally, to aid users in predicting the most accurate automatic tracing results without manual annotations for comparison, we used support vector machine regression to predict reconstruction quality given an image volume and a set of automatic tracings.

97 MATHEMATICS AND COMPUTING↗

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Reinforcement Learning for Anomaly Detection in Nuclear Power Plant Operation and Maintenance

In nuclear power plants (NPPs), timely identification of sensor and human errors is critical to ensure safe and efficient plant operations. Anomaly detection models can be employed for this task. However, traditional anomaly detection approaches may have high dependency on labeled datasets and struggle with adaptability in complex, dynamic environments. Reinforcement learning (RL) has demonstrated significant potential in fault diagnosis and anomaly detection; however, its application to anomaly detection in NPPs remains a relatively underexplored research direction. Hence, to address this gap, in this study, we present a novel physics-informed reinforcement learning model, PIRL-AD: Physics-Informed Reinforcement Learning for Anomaly Detection, that integrates domain knowledge from calorimetric equations into the RL framework for enhanced sensor and human error anomaly detection. We evaluate the performance of PIRL-AD against a non-physics informed RL benchmark and a support vector machine (SVM) on data collected from a forced flow loop testbed. Experimental results suggest that PIRL-AD outperforms other baselines on a range of anomalous datasets that include both sensor and human-induced anomalies across key performance metrics, statistically outperforming the RL and SVM benchmarks with respect to geometric mean (respectively, 92.96% vs. 91.06% vs. 83.01%) and F1-score (respectively, 89.23% vs. 86.98% vs. 77.01%). Furthermore, the findings suggest the potential of physics-integrated reinforcement learning models for enhanced anomaly detection performance in NPPs.

Reinforcement learning↗

Wind turbine gearbox fault prognosis using high-frequency SCADA data

Condition-based maintenance using routinely collected Supervisory Control and Data Acquisition (SCADA) data is a promising strategy to reduce downtime and costs associated with wind farm operations and maintenance. New approaches are continuously being developed to improve the condition monitoring for wind turbines. Development of normal behaviour models is a popular approach in studies using SCADA data. This paper first presents a data-driven framework to apply normal behaviour models using an artificial neural network approach for wind turbine gearbox prognostics. A one-class support vector machine classifier, combining different error parameters, is used to analyse the normal behaviour model error to develop a robust threshold to distinguish anomalous wind turbine operation. A detailed sensitivity study is then conducted to evaluate the potential of using high-frequency SCADA data for wind turbine gearbox prognostics. The results based on operational data from one wind turbine show that, compared to the conventionally used 10-min averaged SCADA data, the use of high-frequency data is valuable as it leads to improved prognostic predictions. High-frequency data provides more insights into the dynamics of the condition of the wind turbine components and can aid in earlier detection of faults.

17 WIND ENERGY↗

Understanding effects of printhead geometry in aerosol jet printing

Aerosol jet printing offers a versatile, high-resolution digital patterning capability broadly relevant to flexible and printed electronic systems. Despite its promise and numerous demonstrations, the theoretical principles driving process outputs have not been thoroughly explored. In this study, a custom-built, modular printing system is developed to provide a head-to-head comparison of two print nozzle geometries to better understand the technology. Print resolution data from a range of process parameters are analyzed using a support vector machine framework. The linear deposition rate is identified as a key variable, which can confound careful studies of printing performance. Taking this into account, a clear difference is observed between the printheads, corresponding to a difference in resolution of 57% ± 11% under typical conditions. Models to understand differences in aerodynamic and mass transport effects identify enhanced drying within the NanoJet printhead as a likely cause of this difference. Overall, this study provides improved understanding of the aerosol jet printing process, including valuable insight to inform process optimization, robust data analysis, ink formulation, and printer geometric design.

42 ENGINEERING↗

Shifts in hydroclimatology of US megaregions in response to climate change

Most of the population and economic growth in the United States occurs in megaregions as the clustered metropolitan areas, whereas climate change may amplify negative impacts on water and natural resources. This study assesses shifts in regional hydroclimatology of fourteen US megaregions in response to climate change over the 21st century. Hydroclimatic projections were simulated using the Variable Infiltration Capacity (VIC) model driven by three downscaled climate models from the Multivariate Adaptive Constructed Analogs (MACA) dataset to cover driest to wettest future conditions in the conterminous United States (CONUS). Shifts in the regional hydroclimatolgy and basin characteristics of US megaregions were represented as a combination of changes in the aridity and evaporative indices using the Budyko framework and Fu's equation. Changes in the climate types of US megaregions were estimated using the Fine Gaussian Support Vector Machine (SVM) method. The results indicate that Los Angeles, San Diego, and San Francisco are more likely to experience less arid conditions with some shifts from Continental to Temperate climate type while the hydroclimatology of Houston may become drier with some shifts from Temperate to Continental climate type. Additionally, water yield is likely to decrease in Seattle. Change in the hydroclimatology of Denver and Phoenix highly depends on the selected climate model. However, the basin characteristics of Phoenix have the highest sensitivity to climate change. Overall, the hydroclimatic conditions of Los Angeles, San Diego, Phoenix, Denver, and Houston have the highest sensitivity to climate change. Understanding of future shifts in hydroclimatology of megaregions can help decision-makers to attenuate negative consequences by implementing appropriate adaptation strategies, particularly in the water-scare megaregions.

54 ENVIRONMENTAL SCIENCES↗

Improving qubit readout with hidden Markov models

We demonstrate the application of pattern recognition algorithms via hidden Markov models (HMM) for qubit readout. This scheme provides a state-path trajectory approach capable of detecting qubit-state transitions and makes for a robust classification scheme with higher starting-state assignment fidelity than when compared to a multivariate Gaussian or a support vector machine scheme. Therefore, the method also eliminates the qubit-dependent readout time optimization requirement in current schemes. Using a HMM state discriminator we estimate fidelities reaching the ideal limit. Unsupervised learning gives access to transition matrix, priors, and IQ distributions, providing a toolbox for studying qubit-state dynamics during strong projective readout.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Cyber-physical security framework for Photovoltaic Farms

With the evolution of PV converters, a growing number of vulnerabilities in PV farms are exposing to cyber threats. To mitigate the influence of cyber-attack on PV farms, it is necessary to study attacks' impact and propose detection methods. To meet this requirement, a cyber-physical security framework is proposed for PV farms. Data integrity attacks (DIAs) are studied on different control loops. As μPMU is gaining in popularity, a lower sampling rate of μPMU data is applied to develop a detection algorithm. We have evaluated two data-driven methods, which are support vector machine (SVM) and long short-term memory (LSTM). Lastly, the data-driven methods verify the feasibility of μPMU data in attack detection.

Attack Impact Analysis↗

Data-driven Cyberattack Detection for Photovoltaic (PV) Systems through Analyzing Micro-PMU Data

With increasing exposure to software-based sensing and control, Photovoltaic (PV) systems are facing higher risks of cyber attacks. Here, to ensure the system stability and minimize potential economic losses, it is imperative to monitor operating states and detect attacks at the early stage. To meet this demand, Micro-Phasor Measurement Units (μPMU) are increasingly popular in monitoring distribution networks. However, due to the relatively low sampling rate, μPMU has not yet been used to detect and classify cyber-attacks in power electronics enabled smart grid. To our knowledge, this is one of the first attempts to use μPMU to detect cyber attacks that degrade the performance of power electronics systems. We propose to apply data-driven methods on micro-PMU data to implement attack detection. We have evaluated data-driven methods, including decision tree (DT), K-nearest neighbor (KNN), support vector machine (SVM), artificial neural network (ANN), long short-term memory (LSTM) and convolutional neural network (CNN). The proposed CNN model achieves the required performances with the highest 99.23% accuracy and 0.9963 F 1 score.

14 SOLAR ENERGY↗

Robust Medium-Voltage Distribution System State Estimation using Multi-Source Data

Due to the lack of sufficient online measurements for distribution system observability, pseudo-measurements from short-term load or distributed renewable energy resources (DERs) forecasting are used. However, the accuracy of them is low and thus significantly limits the performance of distribution system state estimation (DSSE). In this paper, a robust DSSE that integrates multi-source measurement data is proposed. Specifically, the historical low-voltage (LV) side smart meters are used to forecast load and DERs injections via the support vector machine (SVM) with optimally tuned parameters. By contrast, the online smart meters at LV side are utilized to derive equivalent power injections at the MV/LV transformers, yielding more accurate pseudo-measurements compared to the forecasted injections. Furthermore, to deal with bad data caused by communication loss, instrumental errors and cyber attacks, robust DSSE that relies on generalized maximum-likelihood (GM)-estimation criterion is developed. The projection statistics are developed to adjust the weights of each measurement, leading to better balance between pseudo- and real-time measurements. Numerical results conducted on modified IEEE 33-bus system with DG integration demonstrate the effectiveness and robustness of the proposed method.

distribution system state estimation↗

Large-Scale Classification of Urban Structural Units From Remote Sensing Imagery

Remote sensing in combination with deep learning has become instrumental for efficiently and accurately classifying land-use and land-cover across large geographic areas. These technologies have also been successful in characterizing urban environments in terms of their structural units, structure types, or morphological regions. In these approaches, an urban area is partitioned into regions that exhibit homogeneous physical characteristics. However, existing approaches are typically limited to a single city, use inconsistent typologies, and lack scalability and generalization capacity. In this article, we propose an urban structural units categorization scheme and demonstrate its utility by applying it to 13 cities. Inspired by the lack of scalability and generalization capacity in urban structural units mapping, we extend the reach of deep learning and conduct a set of classification experiments in all 13 cities. These experiments offer insights into the strengths and limitations of deep neural networks for classifying urban structural units over diverse geographic regions and on heterogeneous collections of satellite imagery. The efficacy of the proposed deep learning approach is compared to a baseline method of multiscale image features and support vector machines. Our validation on five cities shows that better performance is achieved with deep neural networks. Additionally, we evaluate the impact of input size, model depth, and spatial pyramid pooling to assess the generalization capacity of deep neural networks.

47 OTHER INSTRUMENTATION↗