Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “support vector machine (SVM)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

An Accurate Vegetation and Non-Vegetation Differentiation Approach Based on Land Cover Classification

Accurate vegetation detection is important for many applications, such as crop yield estimation, land cover land use monitoring, urban growth monitoring, drought monitoring, etc. Popular conventional approaches to vegetation detection incorporate the normalized difference vegetation index (NDVI), which uses the red and near infrared (NIR) bands, and enhanced vegetation index (EVI), which uses red, NIR, and the blue bands. Although NDVI and EVI are efficient, their accuracies still have room for further improvement. In this paper, we propose a new approach to vegetation detection based on land cover classification. That is, we first perform an accurate classification of 15 or more land cover types. The land covers such as grass, shrub, and trees are then grouped into vegetation and other land cover types such as roads, buildings, etc. are grouped into non-vegetation. Similar to NDVI and EVI, only RGB and NIR bands are needed in our proposed approach. If Laser imaging, Detection, and Ranging (LiDAR) data are available, our approach can also incorporate LiDAR in the detection process. Results using a well-known dataset demonstrated that the proposed approach is feasible and achieves more accurate vegetation detection than both NDVI and EVI. In particular, a Support Vector Machine (SVM) approach performed 6% better than NDVI and 50% better than EVI in terms of overall accuracy (OA).

54 ENVIRONMENTAL SCIENCES↗

Using Downwelling Far- and Thermal-Infrared Hyperspectral Radiance for Cloud Phase Classification in the Antarctic

The cloud phase is one of the most important parameters of clouds. In this paper, we propose a method for cloud phase classification that synergistically utilizes the far- and thermal-infrared bands based on the Atmospheric Emitted Radiance Interferometer (AERI) at the Atmospheric Radiation Measurement West Antarctic Radiation Experiment (AWARE) observatory in 2016. The possible features in the far- and thermal-infrared bands are analyzed based on the differences in the simulated cloud brightness temperature (BT) spectra with different cloud phases. Using the support vector machine (SVM) algorithm, four features are determined to identify the cloud phase, which include the BT at 900 cm -1 , the slope of the fitted function of BT in the 900–1000 cm -1 interval, the BT difference (BTD) between 512 cm -1 and 726 cm -1 , and the BTD between 550 cm -1 and 726 cm -1 . Here, the performance of the proposed method is evaluated with Shupe’s and Turner’s method. The monthly average accuracy of the proposed method, the method without the two far-infrared features, and Turner’s method are about 76%, 36%, and 49%, respectively, which infer the good performance of the proposed method and also indicate that the far-infrared band features can effectively enhance cloud phase classification. It is notable that, compared to Shupe’s method, the accuracy for the proposed method is only 61% during the Antarctic summer, which results from the definitions of cloud phase and radiative effect. In addition, the accuracy is only 44% for Turner’s method in seasons with a low frequency of mixed clouds due to the significant effect of water vapor.

54 ENVIRONMENTAL SCIENCES↗

Hierarchical Tactile Sensation Integration from Prosthetic Fingertips Enables Multi-Texture Surface Recognition

Multifunctional flexible tactile sensors could be useful to improve the control of prosthetic hands. To that end, highly stretchable liquid metal tactile sensors (LMS) were designed, manufactured via photolithography, and incorporated into the fingertips of a prosthetic hand. Three novel contributions were made with the LMS. First, individual fingertips were used to distinguish between different speeds of sliding contact with different surfaces. Second, differences in surface textures were reliably detected during sliding contact. Third, the capacity for hierarchical tactile sensor integration was demonstrated by using four LMS signals simultaneously to distinguish between ten complex multi-textured surfaces. Four different machine learning algorithms were compared for their successful classification capabilities: K-nearest neighbor (KNN), support vector machine (SVM), random forest (RF), and neural network (NN). The time-frequency features of the LMSs were extracted to train and test the machine learning algorithms. The NN generally performed the best at the speed and texture detection with a single finger and had a 99.2 ± 0.8% accuracy to distinguish between ten different multi-textured surfaces using four LMSs from four fingers simultaneously. The capability for hierarchical multi-finger tactile sensation integration could be useful to provide a higher level of intelligence for artificial hands.

Abd, Moaed A. (ORCID:0000000284954244)↗

IoT Intrusion Detection Taxonomy, Reference Architecture, and Analyses

This paper surveys the deep learning (DL) approaches for intrusion-detection systems (IDSs) in Internet of Things (IoT) and the associated datasets toward identifying gaps, weaknesses, and a neutral reference architecture. A comparative study of IDSs is provided, with a review of anomaly-based IDSs on DL approaches, which include supervised, unsupervised, and hybrid methods. All techniques in these three categories have essentially been used in IoT environments. To date, only a few have been used in the anomaly-based IDS for IoT. For each of these anomaly-based IDSs, the implementation of the four categories of feature(s) extraction, classification, prediction, and regression were evaluated. We studied important performance metrics and benchmark detection rates, including the requisite efficiency of the various methods. Four machine learning algorithms were evaluated for classification purposes: Logistic Regression (LR), Support Vector Machine (SVM), Decision Tree (DT), and an Artificial Neural Network (ANN). Therefore, we compared each via the Receiver Operating Characteristic (ROC) curve. The study model exhibits promising outcomes for all classes of attacks. The scope of our analysis examines attacks targeting the IoT ecosystem using empirically based, simulation-generated datasets (namely the Bot-IoT and the IoTID20 datasets).

97 MATHEMATICS AND COMPUTING↗

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering↗

Measuring Young Stars in Space and Time. II. The Pre-main-sequence Stellar Content of N44

The Hubble Space Telescope survey Measuring Young Stars in Space and Time (MYSST) entails some of the deepest photometric observations of extragalactic star formation, capturing even the lowest-mass stars of the active star-forming complex N44 in the Large Magellanic Cloud. We employ the new MYSST stellar catalog to identify and characterize the content of young pre-main-sequence (PMS) stars across N44 and analyze the PMS clustering structure. To distinguish PMS stars from more evolved line of sight contaminants, a non-trivial task due to several effects that alter photometry, we utilize a machine-learning classification approach. This consists of training a support vector machine (SVM) and a random forest (RF) on a carefully selected subset of the MYSST data and categorize all observed stars as PMS or non-PMS. Combining SVM and RF predictions to retrieve the most robust set of PMS sources, we find ∼26,700 candidates with a PMS probability above 95% across N44. Employing a clustering approach based on a nearest neighbor surface density estimate, we identify 16 prominent PMS structures at 1σ significance above the mean density with sub-clusters persisting up to and beyond 3σ significance. The most active star-forming center, located at the western edge of N44's bubble, is a subcluster with an effective radius of ∼5.6 pc entailing more than 1100 PMS candidates. Furthermore, we confirm that almost all identified clusters coincide with known H ii regions and are close to or harbor massive young O stars or YSOs previously discovered by MUSE and Spitzer observations.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Technical note: Uncertainties in eddy covariance CO 2 fluxes in a semiarid sagebrush ecosystem caused by gap-filling approaches

Abstract. Gap-filling eddy covariance CO2 fluxes is challenging at dryland sites due to small CO2 fluxes. Here, four machine learning (ML) algorithms including artificial neural network (ANN), k-nearest neighbors (KNNs), random forest (RF), and support vector machine (SVM) are employed and evaluated for gap-filling CO2 fluxes over a semiarid sagebrush ecosystem with different lengths of artificial gaps. The ANN and RF algorithms outperform the KNN and SVM in filling gaps ranging from hours to days, with the RF being more time efficient than the ANN. Performances of the ANN and RF are largely degraded for extremely long gaps of 2 months. In addition, our results suggest that there is no need to fill the daytime and nighttime net ecosystem exchange (NEE) gaps separately when using the ANN and RF. With the ANN and RF, the gap-filling-induced uncertainties in the annual NEE at this site are estimated to be within 16 g C m−2, whereas the uncertainties by the KNN and SVM can be as large as 27 g C m−2. To better fill extremely long gaps of a few months, we test a two-layer gap-filling framework based on the RF. With this framework, the model performance is improved significantly, especially for the nighttime data. Therefore, this approach provides an alternative in filling extremely long gaps to characterize annual carbon budgets and interannual variability in dryland ecosystems.

Yao, Jingyu↗

Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry

Advancing automated classification of atmospheric aerosols from Single-Particle Mass Spectrometry (SPMS) data remains challenging due to overlapping ion signatures, compositional diversity, and limited labeled data. This study evaluates supervised and semi-supervised learning frameworks to enhance aerosol identification by jointly leveraging labeled and unlabeled spectra. Four models were compared: a supervised Support Vector Machine (SVM), a self-training SVM, a stacked autoencoder classifier, and a stacked autoencoder trained using a temporal-ensembling Mean Teacher approach. All models achieved high and stable accuracies (90.0 %–91.1 %), surpassing previous results on the same dataset (87 %) and matching the performance of state-of-the-art deep learning methods. Despite small global metric differences (≤ 1 %), semi-supervised variants yielded up to 5 %–10 % improvements for compositionally rare particle types – such as soot (0.77 % of spectra, F1-score: 0.93–0.97) and hazelnut pollen (0.98 % of spectra, F1-score: 0.97–1.00) – equating to roughly ∼ 187 additional correctly classified spectra. These gains are scientifically significant, as such rare particles exert disproportionate influence on radiative absorption and ice nucleation processes; their improved detection reduces modeled uncertainties in aerosol absorption optical depth and mixed-phase cloud ice nucleation rates. The models' residual misclassifications (≈ 9 %) largely arise from true spectral overlap among chemically adjacent species (e.g., Na- vs. K-feldspar, coated vs. uncoated feldspars), reflecting physical compositional continuity rather than algorithmic error. Collectively, these findings demonstrate that leveraging unlabeled data to learn robust spectral representations and refine classification enhances both fidelity and interpretability, bridging data-driven analysis with aerosol–climate process understanding.

54 ENVIRONMENTAL SCIENCES↗

Integrating Predictions for Improving Defect Classification Accuracy in NDT-based Assessment of Concrete - 20229

There is an increasing need to create predictive models for defect classification in concrete using the output of non-destructive testing (NDT) techniques. Recent advancement of machine-learning algorithms has offered several techniques for developing classification models for different types of data sets. However, the performance of these algorithms is very uncertain, mainly when applied to small and noisy datasets. For example, when human access is limited (e.g., nuclear facility), robot-based NDT is preferred. But compared to manual tests with humans present on site, the data sets are small, and more noise can exist. Therefore, it is imperative to develop new approaches to ensure a consistently high classification accuracy for inadequate data sets. This study explores the classification performance on NDT dataset using classifiers from different machine-learning algorithms, namely k-Nearest Neighbor (kNN), Decision Tree, Naive Bayes, Logistic Regression, and Support Vector Machine (SVM). The authors further integrated the predictions from these classifiers using proposed methods. The integration strategy combines the output of the classifiers based on two different measures, accuracy, and performance (ACC and PERF), using equations such as sum, average, and square-root-of-sums-of-squares (SRSS). Our results reveal varying classification accuracies across individual classifiers with different misclassifications across the test data set. The integration strategy provided significant improvement in the classification accuracy compared to the individual classifiers. Furthermore, the results indicate minimal variation across the integration methods as compared to the variation across the individual classifiers. To conclude, prediction integration offers a unique approach for combining the output of multiple classifiers to create redundancies with the potential of achieving high classification performance and improved reliability in predictive models for defect detection in concrete. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Decentralized Microgrid Protection Through Relative Fault Direction Classification: Preprint

Protection in inverter-based resources (IBRs) dominated microgrids generally face significant challenges due to the low fault current and inconsistent fault behaviors from IBRs. Recently, machine learning-based approaches have attracted considerable attention to address these challenges. This paper introduces a novel decentralized protection strategy for microgrids. The proposed method decomposes the protection challenge into several distributed learning tasks, enabling individual relays to autonomously determine the direction of faults using a binary classification framework based on support vector machine (SVM) algorithms. Following the distributed fault direction estimation, classifier outcomes are shared among neighboring relays, facilitating a local decision-making process to ascertain the presence of faults within the neighborhood. Finally, a tripping signal is generated based on the classifier results of each relay to operate the circuit breaker. To test and validate this approach, a 100% renewable microgrid model is simulated in MATLAB/Simulink. In the numerical analysis, the application of SVM classifiers in our approach yields impressive results: an average relay classification accuracy of 98%, and a 96% accuracy in circuit breaker control. These findings highlight the potential of machine-learning-based approaches in enhancing the efficiency and reliability of microgrid protection systems.

decentralized algorithm↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Appendices for Geothermal Exploration Artificial Intelligence Report

The Geothermal Exploration Artificial Intelligence looks to use machine learning to spot geothermal identifiers from land maps. This is done to remotely detect geothermal sites for the purpose of energy uses. Such uses include enhanced geothermal system (EGS) applications, especially regarding finding locations for viable EGS sites. This submission includes the appendices and reports formerly attached to the Geothermal Exploration Artificial Intelligence Quarterly and Final Reports. The appendices below include methodologies, results, and some data regarding what was used to train the Geothermal Exploration AI. The methodology reports explain how specific anomaly detection modes were selected for use with the Geo Exploration AI. This also includes how the detection mode is useful for finding geothermal sites. Some methodology reports also include small amounts of code. Results from these reports explain the accuracy of methods used for the selected sites (Brady Desert Peak and Salton Sea). Data from these detection modes can be found in some of the reports, such as the Mineral Markers Maps, but most of the raw data is included the DOE Database which includes Brady, Desert Peak, and Salton Sea Geothermal Sites.

15 GEOTHERMAL ENERGY↗

Brady Geodatabase for Geothermal Exploration Artificial Intelligence

These files contain the geodatabases related to Brady's Geothermal Field. It includes all input and output files for the Geothermal Exploration Artificial Intelligence. Input and output files are sorted into three categories: raw data, pre-processed data, and analysis (post-processed data). In each of these categories there are six additional types of raster catalogs which are titled Radar, SWIR, Thermal, Geophysics, Geology, and Wells. These inputs and outputs were used with the Geothermal Exploration Artificial Intelligence to identify indicators of blind geothermal systems at the Brady Hot Springs Geothermal Site. The included zip file is a geodatabase to be used with ArcGIS and the tar file is an inclusive database that encompasses the inputs and outputs for the Brady Hot Springs Geothermal Site.

15 GEOTHERMAL ENERGY↗

Desert Peak Geodatabase for Geothermal Exploration Artificial Intelligence

These files contain the geodatabases related to the Desert Peak Geothermal Field. It includes all input and output files used in the project. The files include data categories of raw data, pre-processed data, and analysis (post-processed data). In each of these categories there are six additional types of raster catalogs including Radar, SWIR, Thermal, Geophysics, Geology, and Wells. The files for the Desert Peak Geothermal Site are used with the Geothermal Exploration Artificial Intelligence to identify indicators of blind geothermal systems. The included zip file is a geodatabase to be used with ArcGIS and the tar file is an inclusive database that encompasses the inputs and outputs for the Desert Peak Geothermal Field.

15 GEOTHERMAL ENERGY↗

Salton Sea Geodatabase for Geothermal Exploration Artificial Intelligence

These files contain the geodatabases related to Salton Sea Geothermal Field. It includes all input and output files used with the Geothermal Exploration Artificial Intelligence. Input and output files are sorted into three categories: raw data, pre-processed data, and analysis (post-processed data). In each of these categories there are six additional types of raster catalogs which are titled Radar, SWIR, Thermal, Geophysics, Geology, and Wells. The files are used with the Geothermal Exploration Artificial Intelligence for the Salton Sea Geothermal Site to identify indicators of blind geothermal systems. The included zip file is a geodatabase to be used with ArcGIS and the tar file is an inclusive database that encompasses the inputs and outputs for the Salton Sea Geothermal Site.

15 GEOTHERMAL ENERGY↗

Data supporting manuscript from L. Sheneman, G. Stephanopoulos, A.E. Vasdekis titled "Deep learning classification of lipid droplets in quantitative phase images" as currently under review at PLOS ONE. This includes: 1) raw and binary labeled Quantitative Phase Images (QPI) of Y. lipolytica cells used in the analyses described within the manuscript. 2) various derived data including classifier scores, etc.

Data supporting manuscript from L. Sheneman, G. Stephanopoulos, A.E. Vasdekis titled "Deep learning classification of lipid droplets in quantitative phase images" as currently under review at PLOS ONE. This includes: 1) raw and binary labeled Quantitative Phase Images (QPI) of Y. lipolytica cells used in the analyses described within the manuscript. 2) various derived data including classifier scores, etc.

ANN↗

Estimating mass-absorption cross-section of ambient black carbon aerosols: theoretical, empirical, and machine learning models

The mass-absorption cross-section of black carbon (MAC BC ) is an essential parameter to link the atmospheric concentration of black carbon (BC) with its radiative forcing. When a direct calculation of MAC BC based on observations of aerosol light absorption and BC mass concentration is impossible, we rely on modeling and simulations to estimate MAC BC , but currently, there is no consensus model that can be relied on for accurate predictions across all atmospheric environments when BC particles have different coating thicknesses. Here, we applied five MAC BC prediction models (including three light scattering theories, an empirical model based on observations of particle mass concentrations, and a machine learning model developed in our previous work) to aerosols from three Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) field campaigns. While many studies have found that increasing the complexity of the models helps to constrain biases of the estimated MAC BC , our effort is to evaluate the models based on the criteria of simplicity and accuracy. We find that our machine learning model (support vector machine for regression, SVM) generally performs well across all DOE ARM field campaign data, while the accuracy of core-shell Mie theory depends on the bias correction algorithm applied to filter-based light absorption data. Generally, the empirical model for internally-mixed particles that we considered tends to over-predict MAC BC , while Mie theory for externally-mixed particles tends to under-predict MAC BC . An examination of the influence of coating material on BC cores suggests that the performance of our current SVM model is degraded when the BC is thickly-coated (e.g., it has undergone aging and mixing with other materials in the atmosphere).

54 ENVIRONMENTAL SCIENCES↗