Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “support vector machine (SVM)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

133 records · Page 8

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

Automated Knowledge Discovery From Simulators

A computational method, SimLearn, has been devised to facilitate efficient knowledge discovery from simulators. Simulators are complex computer programs used in science and engineering to model diverse phenomena such as fluid flow, gravitational interactions, coupled mechanical systems, and nuclear, chemical, and biological processes. SimLearn uses active-learning techniques to efficiently address the "landscape characterization problem." In particular, SimLearn tries to determine which regions in "input space" lead to a given output from the simulator, where "input space" refers to an abstraction of all the variables going into the simulator, e.g., initial conditions, parameters, and interaction equations. Landscape characterization can be viewed as an attempt to invert the forward mapping of the simulator and recover the inputs that produce a particular output. Given that a single simulation run can take days or weeks to complete even on a large computing cluster, SimLearn attempts to reduce costs by reducing the number of simulations needed to effect discoveries. Unlike conventional data-mining methods that are applied to static predefined datasets, SimLearn involves an iterative process in which a most informative dataset is constructed dynamically by using the simulator as an oracle. On each iteration, the algorithm models the knowledge it has gained through previous simulation trials and then chooses which simulation trials to run next. Running these trials through the simulator produces new data in the form of input-output pairs. The overall process is embodied in an algorithm that combines support vector machines (SVMs) with active learning. SVMs use learning from examples (the examples are the input-output pairs generated by running the simulator) and a principle called maximum margin to derive predictors that generalize well to new inputs. In SimLearn, the SVM plays the role of modeling the knowledge that has been gained through previous simulation trials. Active learning is used to determine which new input points would be most informative if their output were known. The selected input points are run through the simulator to generate new information that can be used to refine the SVM. The process is then repeated. SimLearn carefully balances exploration (semi-randomly searching around the input space) versus exploitation (using the current state of knowledge to conduct a tightly focused search). During each iteration, SimLearn uses not one, but an ensemble of SVMs. Each SVM in the ensemble is characterized by different hyper-parameters that control various aspects of the learned predictor - for example, whether the predictor is constrained to be very smooth (nearby points in input space lead to similar output predictions) or whether the predictor is allowed to be "bumpy." The various SVMs will have different preferences about which input points they would like to run through the simulator next. SimLearn includes a formal mechanism for balancing the ensemble SVM preferences so that a single choice can be made for the next set of trials.

Burl, Michael↗

Support Vector Machines for Classification of Direct Energy Deposition Standoff Distance for Improved Process Control

A critical factor in the implementation of direct energy deposition is the ability to maintain the standoff distance between the nozzle and the build surface, as this influences powder capture efficiency and overall part quality. Due to process-related variations, layer height may vary, causing unintended variation in standoff distance and poor build quality. While prior work has utilized contact probing to qualify standoff distance during processing, in situ methods for qualification of standoff distance are of major interest. The present work seeks to understand efficacy of image-based methods for classifying standoff distance variation in real-time using support vector machines (SVMs). It was hypothesized that the size of the melt pool and the amount of spatter will have significant correlations with deviations in the standoff distance; thus, SVMs were used on a dataset that is comprised of morphological features of melt pool size and image entropy. The SVM model was used to classify melt pool images into categories according to standoff distance variation from nominal. K-folds cross validation was used to find the optimal hyperparameters for the SVM model. To understand the impact of the selected features on the classification performance and inference speed, multiple models were trained with differing numbers of included features. Results for classification score, inference time, and image preprocessing/feature extraction from these data are reported. The present results show that the SVM model was able to predict the standoff distance classification with an accuracy of 97 percent and a speed of 0.122 s per image, making it a viable solution for real-time control of standoff distance.

Klesmith, Zoe↗

Cyber Spoofing Detection for Grid Distributed Synchrophasor Using Dynamic dual-Kernel SVM

Cyber spoofing with distributed synchrophasor adversely affects the decision-making and situational awareness of the power grid. To detect the spoofing trail, this letter proposes a composite signature-based cyber spoofing detection methodology. The intrinsic principal modes are first extracted from the distributed synchrophasor data. Then, multiple signatures of different intrinsic model components are derived to quantify the spoofing. Thereafter, the dynamic dual-kernel support vector machine is proposed to identify cyber spoofing using multiple signatures. Multiple experimental results using six spoofing methods have verified the validity of the methodology.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Cyber-Attack Identification of Synchrophasor Data Via VMD and Multifusion SVM

A large amount of synchrophasor data in the wide area measurement system (WAMS) needs to be collected and transmitted to the phasor data concentrator, thereby increasing the possibility of being attacked by hackers. The attacked data are therefore hidden into the normal synchrophasor data so that the synchrophasor data based application will be affected. To remedy this problem, an identification framework is proposed to detect the data cyber-attack in WAMS utilizing variational mode decomposition (VMD) and multifusion support vector machine (MSVM). First, VMD is used to transform the attacked data into multiple modal components. Thereafter, a novel MSVM is employed to classify the deterministic features using the proposed linear combined multikernel (LCM). Further, this LCM can fuse multiple types of features, including the time, frequency, and statistical domains of the synchrophasor data. Utilizing the actual data from FNET/GridEye, different experiments are conducted under multiple attack strengths and types. The results demonstrate that the identification framework has higher precision and robustness compared with other conventional classifiers.

97 MATHEMATICS AND COMPUTING↗

Unsupervised Detection of SOC Spoofing in OCPP 2.0.1 EV Charging Communication Protocol Using One-Class SVM

The electric vehicles (EVs) market keeps growing globally; thus, it is critical to secure the EV charging communication protocols in order to guarantee reliable and fair charging operations among the customers. The Open Charge Point Protocol (OCPP) 2.0.1 supports the communication between the Electric Vehicle Supply Equipment (EVSE) and Charging Station Management Systems (CSMSs); therefore, it becomes vulnerable to several types of attacks, which aim to jeopardize smart charging, billing, and energy management. Specifically, OCPP 2.0.1 allows the self-reporting of the State of Charge (SOC) values, which makes it vulnerable to spoofing-based cyberattacks, which target manipulating the scheduling priorities, distorting the load forecasts, and extending the charging sessions in an unfair manner. In this paper, we try to address this type of attack by providing a comprehensive analysis of the SOC spoofing attacks and introducing a novel unsupervised detection framework based on the One-Class Support Vector Machine (OCSVM) algorithm. Specifically, two types of attack scenarios are analyzed (i.e., priority manipulation and session extension) by deriving engineered features that capture the nonlinear relationships under normal charging behavior. Detailed simulation-based results are derived by utilizing the DESL-EPFL Level 3 EV charging dataset. Our results demonstrate high F1-score and recall in identifying spoofed SOC values and that the proposed OCSVM model demonstrates superior performance compared to alternative clustering and deep-learning based detectors.

EV charging↗

Analysis of an optical imaging system prototype for autonomously monitoring zooplankton in an aquaculture facility

Traditional approaches to biomonitoring in aquatic systems, such as sample collection, sorting, and identification, require significant time and effort, thereby limiting the spatiotemporal resolution of sample collection. Additionally, collection and preservation of samples for subsequent taxonomic identification and enumeration leads to mortality of organisms. Recent advances in technologies that utilize optical imaging and machine learning have provided new opportunities to expedite biomonitoring and lead to significant cost savings. These technologies can be advantageous to scientists or managers that conduct routine biomonitoring to inform operations, as in the case of aquaculture facilities. The Small Aquatic Organism optical imaging system (SAO) is a high-throughput optical imaging and classification prototype system that relies on computer vision and machine learning (Support Vector Machines, or SVMs) to autonomously identify and enumerate aquatic organisms. The SAO provides a more sustainable method of collecting large volumes of data and has the benefit of being used in situ. In this study, we tested the performance of the SAO in providing comparable results to manual zooplankton community monitoring in ten ponds at an aquaculture facility. We performed a side-by-side study comparing the sampling methods of plankton tow nets, where major zooplankton taxonomic classes were manually identified and enumerated, to sampling with the SAO. Vouchered samples were used to develop a training library for the SAO, where classes consisted of water boatman and zooplankton groups: cladocerans, copepod adults, copepod nauplii, and rotifers. SAO imagery was manually classified and compared with predicted results for validation. Accuracy for the SVM classifier of the SAO was 37.4 %. Convolutional Neural Networks (CNN) and Random Forest classifiers were also applied to SAO imagery and image features for comparison. The best CNN model and our Random Forest model had accuracies of 80.4 % and 46.6 % respectively. Challenges faced included the small size of copepod nauplii and rotifers and the limited resolution of the imaging camera, although there are tradeoffs between imaging resolution and the sample processing rate. Furthermore, our comparison shows that advancement in both optical imaging and ML are needed in order for the SAO prototype to yield comparable results to manual community monitoring in an aquaculture facility.

54 ENVIRONMENTAL SCIENCES↗