Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Support vector machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering↗

Nuclear Power Fault Diagnostics and Preventative Maintenance Optimization NPIC presentation

The nuclear industry is beginning to see reactors shut down—even after their operating licenses have been extended—because they are not economically competitive with other energy sources. These early closures happen primarily due to economic reasons, despite excellent safety records. Therefore, it is imperative to reduce costs in order to prevent these early closures. One of the contributors to these economic reasons is the large operations and maintenance costs. This paper showcases recent research on advanced fault diagnostics techniques and preventative maintenance optimization (PMO) for reducing NPP maintenance costs. Specifically, it focuses on the feedwater and condensate system (FWCS) for both pressurized- and boiling-water reactor (BWR) systems. The computerized maintenance management system (CMMS), which contains the plant’s digital record of all corrective maintenance (CM) and preventative maintenance (PM) work orders, provided the ground truth for locating potential faults and labeling the process data as either healthy or faulted. Various feature extraction techniques were used to further differentiate the faulted data from the healthy data. Through a cross-validation procedure, support vectors machines were used to label other test sets of process data as either healthy or faulted. With relatively few faults identified in the BWR system, the potential for PMO opens up, since an unnecessary amount of PM leads to inflated maintenance costs. The steps for PMO are summarized, from component health determinations to recommendations for action. An example of PMO assessment is presented for condensate pumps, condensate booster pumps, and the respective motors that drive them.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Nuclear Power Fault Diagnostics and Preventative Maintenance Optimization

The nuclear industry is beginning to see reactors shut down—even after their operating licenses have been extended—because they are not economically competitive with other energy sources. These early closures happen primarily due to economic reasons, despite excellent safety records. Therefore, it is imperative to reduce costs in order to prevent these early closures. One of the contributors to these economic reasons is the large operations and maintenance costs. This paper showcases recent research on advanced fault diagnostics techniques and preventative maintenance optimization (PMO) for reducing NPP maintenance costs. Specifically, it focuses on the feedwater and condensate system (FWCS) for both pressurized- and boiling-water reactor (BWR) systems. The computerized maintenance management system (CMMS), which contains the plant’s digital record of all corrective maintenance (CM) and preventative maintenance (PM) work orders, provided the ground truth for locating potential faults and labeling the process data as either healthy or faulted. Various feature extraction techniques were used to further differentiate the faulted data from the healthy data. Through a cross-validation procedure, support vectors machines were used to label other test sets of process data as either healthy or faulted. With relatively few faults identified in the BWR system, the potential for PMO opens up, since an unnecessary amount of PM leads to inflated maintenance costs. The steps for PMO are summarized, from component health determinations to recommendations for action. An example of PMO assessment is presented for condensate pumps, condensate booster pumps, and the respective motors that drive them.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Hybrid Data-Driven Based HVdc Ancillary Control for Multiple Frequency Data Attacks

The high voltage direct current (HVdc) intertie has been applied to provide ancillary-services for ac grids, utilizing the real-time feedback from phasor measurement units (PMUs). However, PMU data communication is vulnerable to false data injection attacks (FDIA) due to protocol defects, thus the HVdc ancillary control and system stability will be threatened. To address this issue, this article proposes a novel HVdc control strategy based on a hybrid data-driven (HDD) methodology. In this work, the HDD methodology is first proposed to detect the types and duration time of multiple frequency attacks. Specifically, the Hilbert Huang transform (HHT) is used to decompose the frequency data, using variational mode decomposition instead of the traditional empirical mode decomposition, to extract data features. Second, a multikernel support vector machine is proposed to classify the attacked data based on the designed distinctive features from HHT. Meanwhile, the attacking duration time is decided using an unsupervised technique. Third, an HDD-based HVdc ancillary control strategy is established to eliminate the effect of FDIAs on the HVdc frequency response. Comprehensive experiments of HDD-based HVdc ancillary controls under different FDIAs suggest that the proposed HDD could fast and accurately classify the FDIAs, and the HDD-based HVdc ancillary control strategy could significantly suppress the impact of the FDIAs.

97 MATHEMATICS AND COMPUTING↗

Leveraging design of experiments to build chemometric models for the quantification of uranium (VI) and HNO3 by Raman spectroscopy

Partial least squares regression (PLSR) and support vector regression (SVR) models were optimized for the quantification of U(VI) (10–320 g L −1 ) and HNO 3 (0.6–6 M) by Raman spectroscopy with optimized calibration sets chosen by optimal design of experiments. The designed approach effectively minimized the number of samples in the calibration set for PLSR and SVR by selecting sample concentrations with a quadratic process model, despite complex confounding and covarying spectral features in the spectra. The top PLS2 model resulted in percent root mean square errors of prediction for U(VI), HNO 3 , and NO 3 − of 3.7%, 3.6%, and 2.9%, respectively. PLS1 models performed similarly despite modeling an analyte with a majority linear response (i.e., uranyl symmetric stretch) and another with more covarying vibrational modes (i.e., HNO 3 ). Partial least squares (PLS) model loadings and regression coefficients were evaluated to better understand the relationship between weaker Raman bands and covarying spectral features. Support vector machine models outperformed PLS1 models, resulting in percent root mean square error of prediction values for U(VI) and HNO 3 of 1.5% and 3.1%, respectively. The optimal nonlinear SVR model was trained using a similar number of samples (11) compared with the PLSR model, even though PLS is a linear modeling approach. The generic D-optimal design presented in this work provides a robust statistical framework for selecting training set samples in disparate two-factor systems. This approach reinforces Raman spectroscopy for the quantification of species relevant to the nuclear fuel cycle and provides a robust chemometric modeling approach to bolster online monitoring in challenging process environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

iPNHOT: a knowledge-based approach for identifying protein-nucleic acid interaction hot spots

The interaction between proteins and nucleic acids plays pivotal roles in various biological processes such as transcription, translation, and gene regulation. Hot spots are a small set of residues that contribute most to the binding affinity of a protein-nucleic acid interaction. Compared to the extensive studies of the hot spots on protein-protein interfaces, the hot spot residues within protein-nucleic acids interfaces remain less well-studied, in part because mutagenesis data for protein-nucleic acids interaction are not as abundant as that for protein-protein interactions. In this study, we built a new computational model, iPNHOT, to effectively predict hot spot residues on protein-nucleic acids interfaces. One training data set and an independent test set were collected from dbAMEPNI and some recent literature, respectively. To build our model, we generated 97 different sequential and structural features and used a two-step strategy to select the relevant features. The final model was built based only on 7 features using a support vector machine (SVM). The features include two unique features such as ΔSASsa 1/2 and esp3, which are newly proposed in this study. Based on the cross validation results, our model gave F1 score and AUROC as 0.725 and 0.807 on the subset collected from ProNIT, respectively, compared to 0.407 and 0.670 of mCSM-NA, a state-of-the art model to predict the thermodynamic effects of protein-nucleic acid interaction. The iPNHOT model was further tested on the independent test set, which showed that our model outperformed other methods. Here, by collecting data from a recently published database dbAMEPNI, we proposed a new model, iPNHOT, to predict hotspots on both protein-DNA and protein-RNA interfaces. The results show that our model outperforms the existing state-of-art models. Our model is available for users through a webserver: http://zhulab.ahu.edu.cn/iPNHOT/.

59 BASIC BIOLOGICAL SCIENCES↗

Transmission risk of Oropouche fever across the Americas

Abstract Background Vector-borne diseases (VBDs) are important contributors to the global burden of infectious diseases due to their epidemic potential, which can result in significant population and economic impacts. Oropouche fever, caused by Oropouche virus (OROV), is an understudied zoonotic VBD febrile illness reported in Central and South America. The epidemic potential and areas of likely OROV spread remain unexplored, limiting capacities to improve epidemiological surveillance. Methods To better understand the capacity for spread of OROV, we developed spatial epidemiology models using human outbreaks as OROV transmission-locality data, coupled with high-resolution satellite-derived vegetation phenology. Data were integrated using hypervolume modeling to infer likely areas of OROV transmission and emergence across the Americas. Results Models based on one-support vector machine hypervolumes consistently predicted risk areas for OROV transmission across the tropics of Latin America despite the inclusion of different parameters such as different study areas and environmental predictors. Models estimate that up to 5 million people are at risk of exposure to OROV. Nevertheless, the limited epidemiological data available generates uncertainty in projections. For example, some outbreaks have occurred under climatic conditions outside those where most transmission events occur. The distribution models also revealed that landscape variation, expressed as vegetation loss, is linked to OROV outbreaks. Conclusions Hotspots of OROV transmission risk were detected along the tropics of South America. Vegetation loss might be a driver of Oropouche fever emergence. Modeling based on hypervolumes in spatial epidemiology might be considered an exploratory tool for analyzing data-limited emerging infectious diseases for which little understanding exists on their sylvatic cycles. OROV transmission risk maps can be used to improve surveillance, investigate OROV ecology and epidemiology, and inform early detection.

60 APPLIED LIFE SCIENCES↗

Understanding the role of segmentation on process-structure–property predictions made via machine learning

Here, the present study investigated the effect of porosity surface determination methods on performance of machine learning models used to predict the tensile properties of AlSi10Mg processed by laser powder bed fusion from micro-computed tomography data. Machine learning models applied in this work include support vector machines, neural networks, decision trees, and Bayesian classifiers. The effects of isosurface thresholding and local gradient approaches for porosity segmentation, as well as image filtering schemes, on model precision were evaluated for samples produced under differing levels of global energy density.

36 MATERIALS SCIENCE↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Deep Learning for In-Situ Layer Quality Monitoring during Laser-Based Directed Energy Deposition (LB-DED) Additive Manufacturing Process

Defects are a leading issue for the rejection of parts manufactured through the Directed Energy Deposition (DED) Additive Manufacturing (AM) process. In an attempt to illuminate and advance in situ quality monitoring and control of workpieces, we present an innovative data-driven method that synchronously collects sensing data and AM process parameters with a low sampling rate during the DED process. The proposed data-driven technique determines the important influences that individual printing parameters and sensing features have on prediction at the inter-layer qualification to perform feature selection. Three Machine Learning (ML) algorithms including Random Forest (RF), Support Vector Machine (SVM), and Convolutional Neural Network (CNN) are used. During post-production, a threshold is applied to detect low-density occurrences such as porosity sizes and quantities from CT scans that render individual layers acceptable or unacceptable. This information is fed to the ML models for training. Training/testing are completed offline on samples deemed “high-quality” and “low-quality”, utilizing only features recorded from the build process. CNN results show that the classification of acceptable/unacceptable layers can reach between 90% accuracy while training/testing on a “high-quality” sample and dip to 65% accuracy when trained/tested on “low-quality”/“high-quality” (respectively), indicating over-fitting but showing CNN as a promising inter-layer classifier.

36 MATERIALS SCIENCE↗

Machine learning models for estimating contamination across different curbside collection strategies

Contaminated recyclables, which are frequently discarded as waste, pose a significant challenge to the implementation of a circular economy. These contaminated recyclables impede the circulation of resources, resulting in higher processing costs at material recovery facilities (MRFs). Over the past few decades, machine learning (ML) models such as linear regression (LR), support vector machine (SVM), and random forest (RF) have evolved to provide new methods for predicting inbound contamination rates in addition to traditional statistical models. In this study, we applied ML models to predict inbound contamination rates using demographic features from 15 counties in the U.S. with different curbside collection strategies. In general, we found that ML models outperformed linear mixed models. Specifically, SVM models had the highest performance (R 2 = 0.75; mean absolute error (MAE) = 0.06), which may be due to their ability to model nonlinear relationships between features and inbound contamination rates. Further, the key predictor was population, with poverty rate being positively correlated and median age negatively correlated with inbound contamination rates. To improve the management of contamination and enhance the implementation of a circular economy, better models are needed to understand and estimate inbound contamination rates as well as identify critical factors in the present and future.

54 ENVIRONMENTAL SCIENCES↗

A Remote Sensing Technique to Upscale Methane Emission Flux in a Subtropical Peatland

Abstract Quantification of methane (CH 4 ) gas emission from peat is critical to understand CH 4 budget from natural wetlands under a climate warming scenario. Previous studies have focused on prediction and mapping of CH 4 emission flux using process‐based models, while application of statistical‐empirical models for upscaling spatially sparse in situ measurements is scarce. In this study, we developed an empirical remote sensing upscaling approach to estimate CH 4 emission flux in the Everglades using limited in situ point‐based CH 4 emission flux measurements and Landsat data during 2013–2018. We spatially and temporally linked in situ data with Landsat surface reflectance based on temporally composite data sets and developed an object‐based machine learning framework to model and map CH 4 emission flux. An ensemble analysis of two machine learning models, k ‐Nearest Neighbor ( k ‐NN) and Support Vector Machine (SVM), shows that the upscaling approach is promising for predicting CH 4 emission flux with a R 2 of 0.65 and 0.87 based on a fivefold cross‐validation for a dry season and wet season estimation, respectively. We generated emission flux map products that successfully revealed the spatial and temporal heterogeneity of CH 4 emission within the dominant freshwater marsh ecosystem in the Everglades. We conclude that Landsat is promising for upscaling and monitoring CH 4 emission flux and reducing the uncertainty in emission estimates from wetlands.

Zhang, Caiyun↗

Feasibility of Adding Twitter Data to Aid Drought Depiction: Case Study in Colorado

The use of social media, such as Twitter, has changed the information landscape for citizens’ participation in crisis response and recovery activities. Given that drought progression is slow and also spatially extensive, an interesting set of questions arise, such as how the usage of Twitter by a large population may change during the development of a major drought alongside how the changing usage facilitates drought detection. For this reason, contemporary analysis of how social media data, in conjunction with meteorological records, was conducted towards improvement in the detection of drought and its progression. The research utilized machine learning techniques applied over satellite-derived drought conditions in Colorado. Three different machine learning techniques were examined: the generalized linear model, support vector machines and deep learning, each applied to test the integration of Twitter data with meteorological records as a predictor of drought development. It is found that the integration of data resources is viable given that the Twitter-based model outperformed the control run which did not include social media input. Eight of the ten models tested showed quantifiable improvements in the performance over the control run model, suggesting that the Twitter-based model was superior in predicting drought severity. Future work lies in expanding this method to depict drought in the western U.S.

54 ENVIRONMENTAL SCIENCES↗

Measuring Young Stars in Space and Time. II. The Pre-main-sequence Stellar Content of N44

The Hubble Space Telescope survey Measuring Young Stars in Space and Time (MYSST) entails some of the deepest photometric observations of extragalactic star formation, capturing even the lowest-mass stars of the active star-forming complex N44 in the Large Magellanic Cloud. We employ the new MYSST stellar catalog to identify and characterize the content of young pre-main-sequence (PMS) stars across N44 and analyze the PMS clustering structure. To distinguish PMS stars from more evolved line of sight contaminants, a non-trivial task due to several effects that alter photometry, we utilize a machine-learning classification approach. This consists of training a support vector machine (SVM) and a random forest (RF) on a carefully selected subset of the MYSST data and categorize all observed stars as PMS or non-PMS. Combining SVM and RF predictions to retrieve the most robust set of PMS sources, we find ∼26,700 candidates with a PMS probability above 95% across N44. Employing a clustering approach based on a nearest neighbor surface density estimate, we identify 16 prominent PMS structures at 1σ significance above the mean density with sub-clusters persisting up to and beyond 3σ significance. The most active star-forming center, located at the western edge of N44's bubble, is a subcluster with an effective radius of ∼5.6 pc entailing more than 1100 PMS candidates. Furthermore, we confirm that almost all identified clusters coincide with known H ii regions and are close to or harbor massive young O stars or YSOs previously discovered by MUSE and Spitzer observations.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Deep Learning Estimation of Daily Ground–Level NO 2 Concentrations from Remote Sensing Data

The limited number of nitrogen dioxide (NO 2 ) surface measurements calls for the development of highly accurate approaches to estimating surface NO 2 concentrations. In this study, we leverage a new satellite instrument, the TROPOspheric Monitoring Instrument (TROPOMI), along with other predictor variables, to estimate daily surface NO 2 concentrations over Texas in 2019. We use the deep convolutional neural network (Deep-CNN), an advanced deep learning algorithm, to obtain estimates and achieve a correlation coefficient (R) of 0.91, an index of agreement (IOA) of 0.95, and a mean absolute bias (MAB) of 1.75 ppb in surface NO 2 estimation. Additionally, we leverage a novel approach, SHapley Additive exPlanations (SHAP), to describe how Deep-CNN understands each predictor variable. The SHAP results show that the Deep-CNN model has an advanced understanding of the dataset, revealing that TROPOMI closely captures levels of NO 2 . In addition, we show the superiority of our Deep-CNN model at estimating surface NO 2 over other well-known machine learning and regression models in the field, including the support vector machines (SVM), random forest (RF), and multiple linear regression (MLR). Although SVM and RF show strong capabilities at estimating surface NO 2 concentrations, their accuracy is inferior to that of the Deep-CNN model, ranking second and third in model accuracy in this study. The MLR, however, shows a poor ability at NO 2 estimation and ranks last among all models. Furthermore, testing the impact of sample size on model performance, we also show that, compared to other models, Deep-CNN needs more samples to trigger its strength at surface NO 2 estimation.

54 ENVIRONMENTAL SCIENCES↗

A low-complexity non-intrusive approach to predict the energy demand of buildings over short-term horizons

Reliable, non-intrusive, short-term (of up to 12 hours ahead) prediction of a building's energy demand is a critical component of intelligent energy management applications. A number of such approaches have been proposed over time, utilizing various statistical and, more recently, machine learning techniques, such as decision trees, neural networks and support vector machines. Importantly, all of these works barely outperform simple seasonal auto-regressive integrated moving average models, while their complexity is significantly higher. Here, we propose a novel low-complexity non-intrusive approach that improves the predictive accuracy of the state-of-the-art by up to ~10%. The backbone of our approach is a K-nearest neighbours search method, that exploits the demand pattern of the most similar historical days, and incorporates appropriate time-series pre-processing and easing. In the context of this work, we evaluate our approach against state-of-the-art methods and provide insights on their performance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Deep learning based event reconstruction for cyclotron radiation emission spectroscopy

The objective of the cyclotron radiation emission spectroscopy (CRES) technology is to build precise particle energy spectra. This is achieved by identifying the start frequencies of charged particle trajectories which, when exposed to an external magnetic field, leave semi-linear profiles (called tracks) in the time–frequency plane. Due to the need for excellent instrumental energy resolution in application, highly efficient and accurate track reconstruction methods are desired. Deep learning convolutional neural networks (CNNs) - particularly suited to deal with information-sparse data and which offer precise foreground localization—may be utilized to extract track properties from measured CRES signals (called events) with relative computational ease. In this work, we develop a novel machine learning based model which operates a CNN and a support vector machine in tandem to perform this reconstruction. A primary application of our method is shown on simulated CRES signals which mimic those of the Project 8 experiment—a novel effort to extract the unknown absolute neutrino mass value from a precise measurement of tritium β - -decay energy spectrum. When compared to a point-clustering based technique used as a baseline, we show a relative gain of 24.1% in event reconstruction efficiency and comparable performance in accuracy of track parameter reconstruction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Application of Quantum Machine Learning to High Energy Physics Analysis at LHC using IBM Quantum Computer Simulators and IBM Quantum Computer Hardware

One of the major objectives of the experimental programs at the LHC is the discovery of new physics. This requires the identification of rare signals in immense backgrounds. Using machine learning algorithms greatly enhances our ability to achieve this objective. With the progress of quantum technologies, quantum machine learning could become a powerful tool for data analysis in high energy physics. In this study, using IBM gate-model quantum computing systems, we employ the quantum variational classifier method and the quantum kernel estimator method in two recent LHC flagship physics analyses: $t\bar{t}H$ (Higgs boson production in association with a top quark pair) and $H\rightarrow\mu\mu$ (Higgs boson decays to two muons). We have obtained early results with 10 qubits on the IBM quantum simulator and the IBM quantum hardware. On the quantum simulator, the quantum machine learning methods perform similarly to classical algorithms such as SVM (support vector machine) and BDT (boosted decision tree), which are often employed in LHC physics analyses. On the quantum hardware, the quantum machine learning methods have shown promising discrimination power, comparable to that on the quantum simulator. This study demonstrates that quantum machine learning has the ability to differentiate between signal and background in realistic physics datasets.

Chan, Jay↗