Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Support vector machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Root identification in minirhizotron imagery with multiple instance learning

In this study, multiple instance learning (MIL) algorithms to automatically perform root detection and segmentation in minirhizotron imagery using only image-level labels are proposed. Root and soil characteristics vary from location to location, and thus, supervised machine learning approaches that are trained with local data provide the best ability to identify and segment roots in minirhizotron imagery. However, labeling roots for training data (or otherwise) is an extremely tedious and time-consuming task. This paper aims to address this problem by labeling data at the image level (rather than the individual root or root pixel level) and train algorithms to perform individual root pixel level segmentation using MIL strategies. Three MIL methods (multiple instance adaptive cosine coherence estimator, multiple instance support vector machine, multiple instance learning with randomized trees) were applied to root detection and compared to non-MIL approaches. The results show that MIL methods improve root segmentation in challenging minirhizotron imagery and reduce the labeling burden. In our results, multiple instance support vector machine outperformed other methods. The multiple instance adaptive cosine coherence estimator algorithm was a close second with an added advantage that it learned an interpretable root signature which identified the traits used to distinguish roots from soil and did not require parameter selection.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive Model for Workload in Remote Operators During sUAS Contingency Scenarios

The increase in automated capabilities of small Uncrewed Aerial Systems (sUAS) has enabled the human operators to manage larger numbers of vehicles simultaneously. As this happens, the operational paradigm shifts to an m:N configuration where multiple operators (m) are managing multiple vehicles (N) together. However, many questions about how operators will interact with each other and share interaction across the vehicle pool are yet unanswered. Therefore, stakeholders from government and industry have partnered to develop ground control station concepts for such operations. The work presented in this paper aims to identify factors that contribute to operator workload. A supervised machine learning-based method built using Support Vector Machines and K-fold cross-validation was used to create workload prediction models for various NASA TLX subscales by leveraging features related to interactions and their relative timings during m:N operations. Results show that the models yielded fairly high predictive accuracies ranging from ~60-75%.

workload prediction↗

nu-Anomica: A Fast Support Vector Based Novelty Detection Technique

In this paper we propose nu-Anomica, a novel anomaly detection technique that can be trained on huge data sets with much reduced running time compared to the benchmark one-class Support Vector Machines algorithm. In -Anomica, the idea is to train the machine such that it can provide a close approximation to the exact decision plane using fewer training points and without losing much of the generalization performance of the classical approach. We have tested the proposed algorithm on a variety of continuous data sets under different conditions. We show that under all test conditions the developed procedure closely preserves the accuracy of standard one-class Support Vector Machines while reducing both the training time and the test time by 5 - 20 times.

Das, Santanu↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

Construction of Women’s All-Around Speed Skating Event Performance Prediction Model and Competition Strategy Analysis Based on Machine Learning Algorithms

Introduction Accurately predicting the competitive performance of elite athletes is an essential prerequisite for formulating competitive strategies. Women’s all-around speed skating event consists of four individual subevents, and the competition system is complex and challenging to make accurate predictions on their performance. Objective The present study aims to explore the feasibility and effectiveness of machine learning algorithms for predicting the performance of women’s all-around speed skating event and provide effective training and competition strategies. Methods The data, consisting of 16 seasons of world-class women’s all-around speed skating competition results, used in the present study came from the International Skating Union (ISU). According to the competition rules, distinct features are filtered using lasso regression, and a 5,000 m race model and a medal model are built using a fivefold cross-validation method. Results The results showed that the support vector machine model was the most stable among the 5,000 m race and the medal models, with the highest AUC (0.86, 0.81, respectively). Furthermore, 3,000 m points are the main characteristic factors that decide whether an athlete can qualify for the final. The 11th lap of the 5,000 m, the second lap of the 500 m, and the fourth lap of the 1,500 m are the main characteristic factors that affect the athlete’s ability to win medals. Conclusion Compared with logistic regression, random forest, K-nearest neighbor, naive Bayes, neural network, support vector machine is a more viable algorithm to establish the performance prediction model of women’s all-around speed skating event; excellent performance in the 3,000 m event can facilitate athletes to advance to the final, and athletes with outstanding performance in the 500 m event are more likely competitive for medals.

Liu, Meng↗

An Exploratory Approach Using Regression and Machine Learning in the Analysis of Mass Absorption Cross Section of Black Carbon Aerosols: Model Development and Evaluation

Mass absorption cross-section of black carbon (MAC BC ) describes the absorptive cross-section per unit mass of black carbon, and is, thus, an essential parameter to estimate the radiative forcing of black carbon. Many studies have sought to estimate MAC BC from a theoretical perspective, but these studies require the knowledge of a set of aerosol properties, which are difficult and/or labor-intensive to measure. We therefore investigate the ability of seven data analytical approaches (including different multivariate regressions, support vector machine, and neural networks) in predicting MAC BC for both ambient and biomass burning measurements. Our model utilizes multi-wavelength light absorption and scattering as well as the aerosol size distributions as input variables to predict MAC BC across different wavelengths. We assessed the applicability of the proposed approaches in estimating MAC BC using different statistical metrics (such as coefficient of determination (R 2 ), mean square error (MSE), fractional error, and fractional bias). Overall, the approaches used in this study can estimate MAC BC appropriately, but the prediction performance varies across approaches and atmospheric environments. Based on an uncertainty evaluation of our models and the empirical and theoretical approaches to predict MAC BC , we preliminarily put forth support vector machine (SVM) as a recommended data analytical technique for use. We provide an operational tool built with the approaches presented in this paper to facilitate this procedure for future users.

54 ENVIRONMENTAL SCIENCES↗

A Decision-Making Machine Learning Approach in Hermite Spectral Approximations of Partial Differential Equations

The accuracy and effectiveness of Hermite spectral methods for the numerical discretization of partial differential equations on unbounded domains are strongly affected by the amplitude of the Gaussian weight function employed to describe the approximation space. This is particularly true if the problem is under-resolved, i.e., there are no enough degrees of freedom. The issue becomes even more crucial when the equation under study is time-dependent, forcing in this way the choice of Hermite functions where the corresponding weight depends on time. In order to adapt dynamically the approximation space, it is here proposed an automatic decision-making process that relies on machine learning techniques, such as deep neural networks and support vector machines. The algorithm is numerically tested with success on a simple 1D problem, but the main goal is its exportability in the context of more serious applications. Here we also show at the end an application in the framework of plasma physics.

97 MATHEMATICS AND COMPUTING↗

Assessing decision boundaries under uncertainty

In order to make design decisions, engineers may seek to identify regions of the design domain that are acceptable in a computationally efficient manner. A design is typically considered acceptable if its reliability with respect to parametric uncertainty exceeds the designer’s desired level of confidence. Despite major advancements in reliability estimation and in design classification via decision boundary estimation, the current literature still lacks a design classification strategy that incorporates parametric uncertainty and desired design confidence. To address this gap, this paper offers a novel interpretation of the acceptance region by defining the decision boundary as the hypersurface which isolates the designs that exceed a user-defined level of confidence given parametric uncertainty. This work addresses the construction of this novel decision boundary using computationally efficient algorithms that were developed for reliability analysis and decision boundary estimation. The approach proposed in this paper is verified on two physical examples from structural and thermal analysis using Support Vector Machines and Efficient Global Optimization-based contour estimation.

97 MATHEMATICS AND COMPUTING↗

Physics-Infused AI/ML Based Digital-Twin Framework for Flow-Induced-Vibration Damage Prediction in a Nuclear Reactor Heat Exchanger

This report summarizes some of the ongoing work related to the development of an expert-elicitation-digital-twin framework for real time damage state prediction in heat exchanger components of a nuclear reactor. The framework is targeted towards predicting damage associated with coupled low cycle fatigue (associated with regular heat-up, cool-down and power operation transients) and high cycle fatigue (associated with flow induced vibration transients). The overall framework will be based on a NoSQL based database, physics-infused-geometry-dependent virtual-sensor data, different AI/ML techniques-based data-driven-predictive-model applications (Apps) and real-time plant sensor measurements available through few existing sensors. Towards this overall goal, this report updates some of the ongoing work, such as on implementation of a NoSQL Database (such as MongoDB), FE based heat transfer analysis of a heat exchanger (e.g. of a PWR steam generator) for generating geometry-dependent virtual sensor data and evaluation of various AI/ML models such as based on multivariate linear regression, ensembled decision-tree based Random-Forest and Gradient-Boosting regression and high-dimensional-kernel-function-transformation based Support-Vector-Machine regression models. The AI/ML models were evaluated for predicting multi-time-series thermal states at thousands of 3D point-clouds

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Modified Sequence-to-point HVAC Load Disaggregation Algorithm

This paper presents a modified sequence-to-point (S2P) algorithm for disaggregating the heat, ventilation, and air conditioning (HVAC) load from the total building electricity consumption. The original S2P model is convolutional neural network (CNN) based, which uses load profiles as inputs. We propose three modifications. First, the input convolution layer is changed from 1D to 2D so that normalized temperature profiles are also used inputs to the S2P model. Second, a drop-out layer is added to improve adaptability and generalizability so that the model trained in one area can be transferred to other geographical areas without labelled HVAC data. Third, a fine-tuning process is proposed for areas with a small amount of labelled HVAC data so that the pre-trained S2P model can be fine-tuned to achieve higher disaggregation accuracy (i.e., better transferability) in other areas. The model is first trained and tested using smart meter and sub-metered HVAC data collected in Austin, Texas. Then, the trained model is tested on two other areas: Boulder, Colorado and San Diego, California. Simulation results show that the proposed modified S2P algorithm outperforms the original S2P model and the support-vector machine based approach in accuracy, adaptability, and transferability.

Ye, Kai↗

A Chlorophyll-a Algorithm for Landsat-8 Based on Mixture Density Networks

Retrieval of aquatic biogeochemical variables, such as the near-surface concentration of chlorophyll-a (Chla) in inland and coastal waters via remote observations, has long been regarded as a challenging task. This manuscript applies Mixture Density Networks (MDN) that use the visible spectral bands available by the Operational Land Imager (OLI) aboard Landsat-8 to estimate Chla. We utilize a database of co-located in situ radiometric and Chla measurements (N = 4,354), referred to as Type A data, to train and test an MDN model (MDN(A)). This algorithm’s performance, having been proven for other satellite missions, is further evaluated against other widely used machine learning models (e.g., support vector machines), as well as other domain-specific solutions (OC3), and shown to offer significant advancements in the field. Our performance assessment using a held-out test data set suggests that a 49% (median) accuracy with near-zero bias can be achieved via the MDN(A) model, offering improvements of 20 to 100% in retrievals with respect to other models. The sensitivity of the MDN(A) model and benchmarking methods to uncertainties from atmospheric correction (AC) methods, is further quantified through a semi-global matchup dataset (N = 3,337), referred to as Type B data. To tackle the increased uncertainties, alternative MDN models (MDN(B)) are developed through various features of the Type B data (e.g., Rayleigh-corrected reflectance spectra ρ(s)). Using held-out data, along with spatial and temporal analyses, we demonstrate that these alternative models show promise in enhancing the retrieval accuracy adversely influenced by the AC process. Results lend support for the adoption of MDN(B) models for regional and potentially global processing of OLI imagery, until a more robust AC method is developed. Index Terms—Chlorophyll-a, coastal water, inland water, Landsat-8, machine learning, ocean color, aquatic remote sensing.

Brandon Smith↗

Hybrid NN/SVM Computational System for Optimizing Designs

A computational method and system based on a hybrid of an artificial neural network (NN) and a support vector machine (SVM) (see figure) has been conceived as a means of maximizing or minimizing an objective function, optionally subject to one or more constraints. Such maximization or minimization could be performed, for example, to optimize solve a data-regression or data-classification problem or to optimize a design associated with a response function. A response function can be considered as a subset of a response surface, which is a surface in a vector space of design and performance parameters. A typical example of a design problem that the method and system can be used to solve is that of an airfoil, for which a response function could be the spatial distribution of pressure over the airfoil. In this example, the response surface would describe the pressure distribution as a function of the operating conditions and the geometric parameters of the airfoil. The use of NNs to analyze physical objects in order to optimize their responses under specified physical conditions is well known. NN analysis is suitable for multidimensional interpolation of data that lack structure and enables the representation and optimization of a succession of numerical solutions of increasing complexity or increasing fidelity to the real world. NN analysis is especially useful in helping to satisfy multiple design objectives. Feedforward NNs can be used to make estimates based on nonlinear mathematical models. One difficulty associated with use of a feedforward NN arises from the need for nonlinear optimization to determine connection weights among input, intermediate, and output variables. It can be very expensive to train an NN in cases in which it is necessary to model large amounts of information. Less widely known (in comparison with NNs) are support vector machines (SVMs), which were originally applied in statistical learning theory. In terms that are necessarily oversimplified to fit the scope of this article, an SVM can be characterized as an algorithm that (1) effects a nonlinear mapping of input vectors into a higher-dimensional feature space and (2) involves a dual formulation of governing equations and constraints. One advantageous feature of the SVM approach is that an objective function (which one seeks to minimize to obtain coefficients that define an SVM mathematical model) is convex, so that unlike in the cases of many NN models, any local minimum of an SVM model is also a global minimum.

Rai, Man Mohan↗

A novel probabilistic regression model for electrical peak demand estimate of commercial and manufacturing buildings

Due to the high cost of electricity in commercial and industrial sectors, demand forecast models have gained increasing attention. However, there are two unresolved issues: (1) Models are not adaptable when exposed to previously unknown data (2) The value of regression methods vs. state-of-the-art machine learning models has not been made apparent before. This study’s goal is to develop probabilistic demand estimation models. Herein, we propose a probabilistic Bayesian regression framework that can not only estimate future demands with high accuracy but also be updated once new information is available. By applying the proposed algorithm to two real-world case studies (commercial and manufacturing), we show a 40.3% and 30.8% improvement in terms of mean absolute error for the two cases. Moreover, the proposed technique outperforms powerful machine learning approaches, including support vector machine by 10.39%, random forest by 6.17%, and multilayer perceptron by 9.14% in terms of mean absolute percentage error.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Delineating Variabilities of Groundwater Level Prediction Across the Agriculturally Intensive Transboundary Aquifers of South Asia

We report the groundwater depletion in South Asia’s Himalayan, transboundary Indus-Ganges-Brahmaputra-Meghna (IGBM) rivers basin is among the highest globally. Given the high irrigation demand and population, groundwater sustainability requires an improved understanding of groundwater systems for the accurate prediction of groundwater levels (GWLs). However, the prediction of groundwater system behaviors is a significant challenge since it is dominated by spatiotemporal and subsurface depth-dependent drivers. Earlier studies that address the challenges are mainly based on the short spatial and temporal extent and/or do not separate the renewable (i.e., shallow) vs nonrenewable (i.e., deeper) groundwater signals. Here, we first identified the variable importance of spatial and depth-dependent drivers on GWL in the IGBM basin. Our results indicate a greater influence of anthropogenic factors (i.e., widespread pumping and increased population) in most parts of the IGBM basin, except in the precipitation-dominated basin of the Brahmaputra. Our next purpose was to delineate a multifactorial approach for GWL prediction using the two most used machine learning models (i.e., support vector machine and feed-forward neural network) in the literature. In general, the machine learning model outputs show a good match in comparison to the GWL from the observation wells (n = 2303 distributed across India and Bangladesh) with some limitations in areas with increased groundwater irrigation. We separately compared the results from shallow (<35 m) and deep (>35 m) observation wells, emphasizing the significance of deep groundwater pumping. Our approach highlights the importance of spatiotemporal to multidepth factors in GWL prediction and can be adopted in other parts of the globe to predict GWLs.

54 ENVIRONMENTAL SCIENCES↗

Differentiation and classification of bacterial endotoxins based on surface enhanced Raman scattering and advanced machine learning

Bacterial endotoxin, a major component of the Gram-negative bacterial outer membrane leaflet, is a lipopolysaccharide shed from bacteria during their growth and infection and can be utilized as a biomarker for bacterial detection. Here, the surface enhanced Raman scattering (SERS) spectra of eleven bacterial endotoxins with an average detection amount of 8.75 pg per measurement have been obtained based on silver nanorod array substrates, and the characteristic SERS peaks have been identified. With appropriate spectral pre-processing procedures, different classical machine learning algorithms, including support vector machine, k-nearest neighbor, random forest, etc., and a modified deep learning algorithm, RamanNet, have been applied to differentiate and classify these endotoxins. It has been found that most conventional machine learning algorithms can attain a differentiation accuracy of >99%, while RamanNet can achieve 100% accuracy. Such an approach has the potential for precise classification of endotoxins and could be used for rapid medical diagnoses and therapeutic decisions for pathogenic infections.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Classification of four-qubit entangled states via machine learning

We apply the support vector machine (SVM) algorithm to derive a set of entanglement witnesses (EW) to identify entanglement patterns in families of four-qubit states. The effectiveness of SVM for practical EW implementations stems from the coarse-grained description of families of equivalent entangled quantum states. The equivalence criteria in our work is based on the stochastic local operations and classical communication classification and the description of the four-qubit entangled Werner states. We numerically verify that the SVM approach provides an effective tool to address the entanglement witness problem when the coarse-grained description of a given family state is available. Here, we also discuss and demonstrate the efficiency of nonlinear kernel SVM methods as applied to four-qubit entangled state classification.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Cyber-Attack Identification of Synchrophasor Data Via VMD and Multifusion SVM

A large amount of synchrophasor data in the wide area measurement system (WAMS) needs to be collected and transmitted to the phasor data concentrator, thereby increasing the possibility of being attacked by hackers. The attacked data are therefore hidden into the normal synchrophasor data so that the synchrophasor data based application will be affected. To remedy this problem, an identification framework is proposed to detect the data cyber-attack in WAMS utilizing variational mode decomposition (VMD) and multifusion support vector machine (MSVM). First, VMD is used to transform the attacked data into multiple modal components. Thereafter, a novel MSVM is employed to classify the deterministic features using the proposed linear combined multikernel (LCM). Further, this LCM can fuse multiple types of features, including the time, frequency, and statistical domains of the synchrophasor data. Utilizing the actual data from FNET/GridEye, different experiments are conducted under multiple attack strengths and types. The results demonstrate that the identification framework has higher precision and robustness compared with other conventional classifiers.

97 MATHEMATICS AND COMPUTING↗