Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Support vector machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Linac_Gen: integrating machine learning and particle-in-cell methods for enhanced beam dynamics at Fermilab

Here, we introduce Linac_Gen, a tool developed at Fermilab, which combines machine learning algorithms with Particle-in-Cell methods to advance beam dynamics in linacs. Linac_Gen employs techniques such as Random Forest, Genetic Algorithms, Support Vector Machines, and Neural Networks, achieving a tenfold increase in speed for phase-space matching in linacs over traditional methods through the use of genetic algorithms. Crucially, Linac_Gen's adept handling of 3D field maps elevates the precision and realism in simulating beam instabilities and resonances, marking a key advancement in the field. Benchmarked against established codes, Linac_Gen demonstrates not only improved efficiency and precision in beam dynamics studies but also in the design and optimization of linac systems, as evidenced in its application to Fermilab's PIP-II linac project. This work represents a notable advancement in accelerator physics, marrying ML with PIC methods to set new standards for efficiency and accuracy in accelerator design and research. Linac_Gen exemplifies a novel approach in accelerator technology, offering substantial improvements in both theoretical and practical aspects of beam dynamics.

43 PARTICLE ACCELERATORS↗

Multiresolution classification of turbulence features in image data through machine learning

During large-scale simulations, intermediate data products such as image databases have become popular due to their low relative storage cost and fast in-situ analysis. Serving as a form of data reduction, these image databases have become more acceptable to perform data analysis on. In this work, we present an image-space detection and classification system for extracting vortices at multiple scales through wavelet-based filtering. A custom image-space descriptor is used to encode a large variety of vortex-types and a machine learning system is trained for fast classification of vortex regions. By combining a radial-based histogram descriptor, a bag of visual words feature descriptor, and a support vector machine, our results show that we are able to detect and classify vortex features at various sizes at multiple scales. Once trained, our framework enables the fast extraction of vortices on new, unknown image datasets for flow analysis.

97 MATHEMATICS AND COMPUTING↗

Scalability Analysis of Quantum Models for Stress and Emotion Detection

Stress and emotion detection from high-dimensional physiological signals is a challenging task, particularly when aiming for accurate classification across diverse behavioral states. Quantum machine learning (QML) is promising for modeling such high-dimensional data, but scalability is limited by qubit resources and the exponential cost of classical statevector simulation. This work studies the scalability of quantum support vector machines (QSVMs) for binary stress detection and three-class emotion recognition (Negative/Neutral/Positive) under varying qubit counts and angle-encoding strategies. We also present a comparison study with one-feature-per-qubit (1:1) and two-features-per-qubit (2:1) mappings. Experiments are executed on HPC infrastructure using NVIDIA CUDA-Q to evaluate performance, variance, and class-dependent separability at higher-qubit setups. Results show that larger Hilbert spaces can improve peak accuracy but may increase instability. At the same time, dense 2:1 encoding yields more consistent stress detection performance. For emotion recognition, scaling improves discrimination for classes like Negative and Positive more than Neutral. We find that effective QML scaling is task-dependent and benefits more from encoding design than simply increasing qubit count.

Onim, Md. Saif Hassan [University of Tennessee, Kn↗

Machine learning assisted phase and size-controlled synthesis of iron oxide particles

Synthesis of iron oxides with specific phases and particle sizes is a crucial challenge in various fields, including materials science, energy storage, biomedical applications, environmental science, and earth science. However, despite significant advances in this area, much of the current palette of particle outcomes has been based on time-consuming trial-and-error exploration of synthesis conditions. The present study was designed to explore a very different approach to 1) predict the outcome of synthesis from specified reaction parameters based on using machine learning (ML) techniques, and 2) correlate sets of parameters to obtain products with desired outcomes by a newly designed recommendation algorithm. To achieve this, four ML algorithms were tested, namely random forest, logistic regression, support vector machine, and k-nearest neighbor. Among the models, random forest outperformed the others, attaining 96% and 81% accuracy when predicting the phase and size of iron oxide particles in the test dataset. Surprisingly, the permutation feature importance analysis revealed that volume, which may strongly relate to pressure, was one of the important features, along with precursor concentration, pH, temperature, and time, influencing the phase and size of iron oxide particles during synthesis. To verify the robustness of the random forest models, prediction and experimental results were compared based on 24 randomly generated methods in additive and non-additive systems not included in the datasets. The predictions of product phase and particle size from the models agreed well with the experimental results. Furthermore, a searching and ranking algorithm was developed to recommend potential synthesis parameters for obtaining iron oxide products with the desired phase and particle size from previous studies in the dataset. Furthermore, this study lays the foundation for a closed-loop approach in materials synthesis and preparation, beginning with suggesting potential reaction parameters from the dataset and predicting potential outcomes, followed by conducting experiments and analyses, and ultimately enriching the dataset.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimizing Classifiers for Radionuclide Identification

Identifying threat nuclear materials is a critical for the prevention of acts of nuclear terrorism on the homeland. For this purpose, many radionuclide identification devices are deployed in the field. However, these will not necessarily be in the hands of non-experts, therefore these devices need to provide ready-made answers for the personnel in the field. This is where advanced algorithms are employed to both interpret the data and provide the identification of the nuclear material being interrogated. We took a machine learning approach to identification, by using training and validation data sets to create and optimize classifiers which determine which radionuclide is consistent with the data. The classifiers investigated were the Random Forest, Decision Tree, Support Vector Machine, and XG Boost and their performance was judged using the F1 Score for both hyperparameter tuning and comparison. In the end, we found out that the Random Forest Classifier worked the best based off the F1 Score they got which was 0.98.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Data-driven analysis and prediction of wastewater treatment plant performance: Insights and forecasting for sustainable operations

Here this study presents a comprehensive performance and forecasting analysis of the As-Samra wastewater treatment plant (WWTP) in Jordan, with two main objectives. Firstly, a thorough evaluation of the plant's performance is conducted. The analysis involves independently assessing historical operational conditions, plant production, and their statistical correlations using various statistical techniques. The second objective focuses on developing a data-driven forecasting approach to predict the plant's production one month in advance, using multiple machine learning models. The results highlight the effectiveness of principal component analysis (PCA) in simplifying operational data, revealing distinct operational clusters, and identifying seasonal production patterns while showing correlations between operational conditions and overall power production. The support vector machine (SVM) forecasting model emerged as the top performer, showcasing the potential of a hybrid forecasting approach. The findings offer valuable perspectives for enhancing operational efficiency, refining production planning, and ultimately improving the environmental impact of the plant.

42 ENGINEERING↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Improving Data and Prediction Quality of High-Throughput Perovskite Synthesis with Model Fusion

Combinatorial fusion analysis (CFA) is an approach for combining multiple scoring systems using the rank-score characteristic function and cognitive diversity measure. One example is to combine diverse machine learning models to achieve better prediction quality. In this work, we apply CFA to the synthesis of metal halide perovskites containing organic ammonium cations via inverse temperature crystallization. Using a data set generated by high-throughput experimentation, four individual models (support vector machines, random forests, weighted logistic classifier, and gradient boosted trees) were developed. We characterize each of these scoring systems and explore 66 possible combinations of the models. When measured by the precision on predicting crystal formation, the majority of the combination models improves the individual model results. The best combination models outperform the best individual models by 3.9 percentage points in precision. In addition to improving prediction quality, we demonstrate how the fusion models can be used to identify mislabeled input data and address issues of data quality. In particular, we identify example cases where all single models and all fusion models do not give the correct prediction. Experimental replication of these syntheses reveals that these compositions are sensitive to modest temperature variations across the different locations of the heating element that can hinder or enhance the crystallization process. In summary, we demonstrate that model fusion using CFA can not only identify a previously unconsidered influence on reaction outcome but also be used as a form of quality control for high-throughput experimentation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Unraveling the Correlation between Raman and Photoluminescence in Monolayer MoS 2 through Machine‐Learning Models

Abstract 2D transition metal dichalcogenides (TMDCs) with intense and tunable photoluminescence (PL) have opened up new opportunities for optoelectronic and photonic applications such as light‐emitting diodes, photodetectors, and single‐photon emitters. Among the standard characterization tools for 2D materials, Raman spectroscopy stands out as a fast and non‐destructive technique capable of probing material's crystallinity and perturbations such as doping and strain. However, a comprehensive understanding of the correlation between photoluminescence and Raman spectra in monolayer MoS 2 remains elusive due to its highly nonlinear nature. Here, the connections between PL signatures and Raman modes are systematically explored, providing comprehensive insights into the physical mechanisms correlating PL and Raman features. This study's analysis further disentangles the strain and doping contributions from the Raman spectra through machine‐learning models. First, a dense convolutional network (DenseNet) to predict PL maps by spatial Raman maps is deployed. Moreover, a gradient boosted trees model (XGBoost) with Shapley additive explanation (SHAP) to bridge the impact of individual Raman features in PL features is applied. Last, a support vector machine (SVM) to project PL features on Raman frequencies is adopted. This work may serve as a methodology for applying machine learning to characterizations of 2D materials.

Lu, Ang‐Yu↗

Searching for Novel Chemistry in Exoplanetary Atmospheres Using Machine Learning for Anomaly Detection

Abstract The next generation of telescopes will yield a substantial increase in the availability of high-quality spectroscopic data for thousands of exoplanets. The sheer volume of data and number of planets to be analyzed greatly motivate the development of new, fast, and efficient methods for flagging interesting planets for reobservation and detailed analysis. We advocate the application of machine learning (ML) techniques for anomaly (novelty) detection to exoplanet transit spectra, with the goal of identifying planets with unusual chemical composition and even searching for unknown biosignatures. We successfully demonstrate the feasibility of two popular anomaly detection methods (local outlier factor and one-class support vector machine) on a large public database of synthetic spectra. We consider several test cases, each with different levels of instrumental noise. In each case, we use receiver operating characteristic curves to quantify and compare the performance of the two ML techniques.

Astronomy & Astrophysics↗

High–Resolution Maps of Near–Surface Permafrost for Three Watersheds on the Seward Peninsula, Alaska Derived From Machine Learning

Permafrost soils are a critical component of the global carbon cycle and are locally important because they regulate the hydrologic flux from uplands to rivers. Furthermore, degradation of permafrost soils causes land surface subsidence, damaging infrastructure that is crucial for local communities. Regional and hemispherical maps of permafrost are too coarse to resolve distributions at a scale relevant to assessments of infrastructure stability or to illuminate geomorphic impacts of permafrost thaw. Here we train machine learning models to generate meter–scale maps of near–surface permafrost for three watersheds in the discontinuous permafrost region. The models were trained using ground truth determinations of near–surface permafrost presence from measurements of soil temperature and electrical resistivity. We trained three classifiers: extremely randomized trees (ERTr), support vector machines (SVM), and an artificial neural network (ANN). Model uncertainty was determined using k–fold cross validation, and the modeled extents of near–surface permafrost were compared to the observed extents at each site. At–a–site near–surface permafrost distributions predicted by the ERTr produced the highest accuracy (70%–90%). However, the transferability of the ERTr to the sites outside of the training data set was poor, with accuracies ranging from 50% to 77%. The SVM and ANN models had lower accuracies for at–a–site prediction (70%–83%), yet they had greater accuracy when transferred to the non–training site (62%–78%). These models demonstrate the potential for integrating high–resolution spatial data and machine learning models to develop maps of near–surface permafrost extent at resolutions fine enough to assess infrastructure vulnerability and landscape morphology influenced by permafrost thaw.

54 ENVIRONMENTAL SCIENCES↗

High bias machine learning for antineutrino-based safeguards for small reactors

The statistical methods used for antineutrino detection will need to be improved to effectively monitor the inventory of next-generation nuclear reactors. In this sensitivity study, we evaluate machine learning models compared to previously used statistical approaches to identify diversion scenarios in a simulated Advanced Fast Reactor (AFR)-100. A chi-square goodness-of-fit technique, which individually compares the simulated antineutrino yields to the expected antineutrino yield, resulted in precise but low diversion detection probability. Various support vector machine (SVM) models were applied with diverse training datasets to evaluate the robustness of the method towards unexpected or “unseen” diversion scenarios. Furthermore, our results indicate that while the SVM models significantly improved the detection probability of near-field antineutrino-based safeguards, up to a probability of ~0.04, for the simulated small reactor, the detection system still needs improvements to reach the 0.2 detection limit established by the International Atomic Energy Agency.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Analysis of zebrafish periderm enhancers facilitates identification of a regulatory variant near human KRT8/18

Genome-wide association studies for non-syndromic orofacial clefting (OFC) have identified single nucleotide polymorphisms (SNPs) at loci where the presumed risk-relevant gene is expressed in oral periderm. The functional subsets of such SNPs are difficult to predict because the sequence underpinnings of periderm enhancers are unknown. We applied ATAC-seq to models of human palate periderm, including zebrafish periderm, mouse embryonic palate epithelia, and a human oral epithelium cell line, and to complementary mesenchymal cell types. We identified sets of enhancers specific to the epithelial cells and trained gapped-kmer support-vector-machine classifiers on these sets. We used the classifiers to predict the effects of 14 OFC-associated SNPs at 12q13 near KRT18. All the classifiers picked the same SNP as having the strongest effect, but the significance was highest with the classifier trained on zebrafish periderm. Reporter and deletion analyses support this SNP as lying within a periderm enhancer regulating KRT18/KRT8 expression.

59 BASIC BIOLOGICAL SCIENCES↗

Decentralized Microgrid Protection Through Relative Fault Direction Classification: Preprint

Protection in inverter-based resources (IBRs) dominated microgrids generally face significant challenges due to the low fault current and inconsistent fault behaviors from IBRs. Recently, machine learning-based approaches have attracted considerable attention to address these challenges. This paper introduces a novel decentralized protection strategy for microgrids. The proposed method decomposes the protection challenge into several distributed learning tasks, enabling individual relays to autonomously determine the direction of faults using a binary classification framework based on support vector machine (SVM) algorithms. Following the distributed fault direction estimation, classifier outcomes are shared among neighboring relays, facilitating a local decision-making process to ascertain the presence of faults within the neighborhood. Finally, a tripping signal is generated based on the classifier results of each relay to operate the circuit breaker. To test and validate this approach, a 100% renewable microgrid model is simulated in MATLAB/Simulink. In the numerical analysis, the application of SVM classifiers in our approach yields impressive results: an average relay classification accuracy of 98%, and a 96% accuracy in circuit breaker control. These findings highlight the potential of machine-learning-based approaches in enhancing the efficiency and reliability of microgrid protection systems.

decentralized algorithm↗

Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry

Advancing automated classification of atmospheric aerosols from Single-Particle Mass Spectrometry (SPMS) data remains challenging due to overlapping ion signatures, compositional diversity, and limited labeled data. This study evaluates supervised and semi-supervised learning frameworks to enhance aerosol identification by jointly leveraging labeled and unlabeled spectra. Four models were compared: a supervised Support Vector Machine (SVM), a self-training SVM, a stacked autoencoder classifier, and a stacked autoencoder trained using a temporal-ensembling Mean Teacher approach. All models achieved high and stable accuracies (90.0 %–91.1 %), surpassing previous results on the same dataset (87 %) and matching the performance of state-of-the-art deep learning methods. Despite small global metric differences (≤ 1 %), semi-supervised variants yielded up to 5 %–10 % improvements for compositionally rare particle types – such as soot (0.77 % of spectra, F1-score: 0.93–0.97) and hazelnut pollen (0.98 % of spectra, F1-score: 0.97–1.00) – equating to roughly ∼ 187 additional correctly classified spectra. These gains are scientifically significant, as such rare particles exert disproportionate influence on radiative absorption and ice nucleation processes; their improved detection reduces modeled uncertainties in aerosol absorption optical depth and mixed-phase cloud ice nucleation rates. The models' residual misclassifications (≈ 9 %) largely arise from true spectral overlap among chemically adjacent species (e.g., Na- vs. K-feldspar, coated vs. uncoated feldspars), reflecting physical compositional continuity rather than algorithmic error. Collectively, these findings demonstrate that leveraging unlabeled data to learn robust spectral representations and refine classification enhances both fidelity and interpretability, bridging data-driven analysis with aerosol–climate process understanding.

54 ENVIRONMENTAL SCIENCES↗