Engineering PapersSearch

SEARCH · Engineering Papers

Results for “multiclass”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Multiclass Reduced-Set Support Vector Machines

There are well-established methods for reducing the number of support vectors in a trained binary support vector machine, often with minimal impact on accuracy. We show how reduced-set methods can be applied to multiclass SVMs made up of several binary SVMs, with significantly better results than reducing each binary SVM independently. Our approach is based on Burges' approach that constructs each reduced-set vector as the pre-image of a vector in kernel space, but we extend this by recomputing the SVM weights and bias optimally using the original SVM objective function. This leads to greater accuracy for a binary reduced-set SVM, and also allows vectors to be 'shared' between multiple binary SVMs for greater multiclass accuracy with fewer reduced-set vectors. We also propose computing pre-images using differential evolution, which we have found to be more robust than gradient descent alone. We show experimental results on a variety of problems and find that this new approach is consistently better than previous multiclass reduced-set methods, sometimes with a dramatic difference.

reduced set methods

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING

A parametric multiclass Bayes error estimator for the multispectral scanner spatial model performance evaluation

The author has identified the following significant results. The probability of correct classification of various populations in data was defined as the primary performance index. The multispectral data being of multiclass nature as well, required a Bayes error estimation procedure that was dependent on a set of class statistics alone. The classification error was expressed in terms of an N dimensional integral, where N was the dimensionality of the feature space. The multispectral scanner spatial model was represented by a linear shift, invariant multiple, port system where the N spectral bands comprised the input processes. The scanner characteristic function, the relationship governing the transformation of the input spatial, and hence, spectral correlation matrices through the systems, was developed.

Mobasseri, B. G.

Multiclass Flight Anomaly Detection Using Sensor Fusion Based on Dempster-Shafer Theory

As aviation systems in commercial operations continue to grow in complexity, the anomalies exhibited by these systems become more elaborate and difficult to detect. To address the challenge of detecting these complex anomalies, deep learning models have been used extensively in aviation anomaly detection studies, at the expense of end-user interpretability. Aiming to maintain the same level of interpretability as traditional threshold-exceedance methods, we continue our development of prediction models using ordinal patterns and their distributions throughout the flight. Specifically, this study extends our work into multiclass anomaly detection using sensor fusion based on Dempster-Shafer theory (DST), a second-order probability theory used to combine information from different sources of evidence. Our approach uses DST to reduce the uncertainty in the class predictions of an ensemble of classifiers. These classifiers rely on the similarity between flight data and class templates to make a prediction of the state of the aircraft. Our approach aims to take advantage of simple models trained on interpretable features (ordinal patterns) to correctly predict an anomaly and identify the flight dynamics linked to the anomaly. Our results show an improvement when using DST-based sensor fusion over simple majority voting. Additionally, our results provide insight into aircraft states linked to rare high-risk anomalies.

Risk detection

Multiclass Flight Anomaly Detection Using Sensor Fusion Based on Dempster-Shafer Theory

As aviation systems in commercial operations continue to grow in complexity, the anomalies exhibited by these systems become more elaborate and difficult to detect. To address the challenge of detecting these complex anomalies, deep learning models have been used extensively in aviation anomaly detection studies, at the expense of end-user interpretability. Aiming to maintain the same level of interpretability as traditional threshold-exceedance methods, we continue our development of prediction models using ordinal patterns and their distributions throughout the flight. Specifically, this study extends our work into multiclass anomaly detection using sensor fusion based on Dempster-Shafer theory (DST), a second-order probability theory used to combine information from different sources of evidence. Our approach uses DST toreduce the uncertainty in the class predictions of an ensemble of classifiers. These classifiers rely on the similarity between flight data and class templates to make a prediction of the state of the aircraft. Our approach aims to take advantage of simple models trained on interpretable features (ordinal patterns) to correctly predict an anomaly and identify the flight dynamics linked to the anomaly. Our results show an improvement when using DST-based sensor fusion over simple majority voting. Additionally, our results provide insight into aircraft states linked to rare high-risk anomalies.

Risk detection

Multiclass Bayes error estimation by a feature space sampling technique

A general Gaussian M-class N-feature classification problem is defined. An algorithm is developed that requires the class statistics as its only input and computes the minimum probability of error through use of a combined analytical and numerical integration over a sequence simplifying transformations of the feature space. The results are compared with those obtained by conventional techniques applied to a 2-class 4-feature discrimination problem with results previously reported and 4-class 4-feature multispectral scanner Landsat data classified by training and testing of the available data.

Mobasseri, B. G.

Multiclass Continuous Correspondence Learning

We extend the Structural Correspondence Learning (SCL) domain adaptation algorithm of Blitzer er al. to the realm of continuous signals. Given a set of labeled examples belonging to a 'source' domain, we select a set of unlabeled examples in a related 'target' domain that play similar roles in both domains. Using these 'pivot samples, we map both domains into a common feature space, allowing us to adapt a classifier trained on source examples to classify target examples. We show that when between-class distances are relatively preserved across domains, we can automatically select target pivots to bring the domains into correspondence.

correspondence learning

Investigation of spatial misregistration effects in multispectral scanner data

The author has identified the following significant results. A model for estimating the expected proportion of multiclass pixels in a scene was generalized and extended to include misregistration effects. Another substantial effort was the development of a simulation model to generate signatures to represent the distributions of signals from misregistered multiclass pixels, based on single class signatures. Spatial misregistration causes an increase in the proportion of multiclass pixels in a scene and a decorrelation between signals in misregistered data channels. The multiclass pixel proportion estimation model indicated that this proportion is strongly dependent on the pixel perimeter and on the ratio of the total perimeter of the fields in the scene to the area of the scene. Test results indicated that expected values computed with this model were similar to empirical measurements made of this proportion in four LACIE data segments.

Nalepka, R. F.

Arm and shoulder muscle segmentation in axial MRI with UNet deep learning model

Quantifying individual upper-limb muscle volumes from MRI provides key insight into muscle-specific strength, deficits, and adaptations. Manual delineation is the gold standard but time‑intensive, and the performance of current deep learning approaches, particularly for small or anatomically complex muscles, remains incompletely characterized. We evaluated a state‑of‑the‑art deep learning framework across the entire upper limb and analyzed factors governing segmentation performance, with attention to the forearm. Three previously published MRI datasets (1.5 T, 3D GRE T1‑weighted; total n = 39) spanning young, middle‑aged, and older adults were curated and quality‑checked, including expert manual segmentations for 31 muscles. Following multiclass mask reconstruction, we trained three 3D nnU‑Net multiclass models matched to the muscle subsets present across datasets, using five‑fold cross‑validation and a composite Dice Similarity Coefficient (DSC) + cross entropy loss. Segmentation accuracy was assessed with DSC. Performance varied across muscles (mean DSC = 0.806 ± 0.098), ranging from 0.920 (Deltoid) to 0.461 (Extensor pollicis brevis). In uncertainty‑weighted regressions, muscle volume was positively associated with DSC (R2 = 0.36, p < 0.001), whereas training segmentation count and muscle orientation showed negligible associations (R2 ≤ 0.06). A weighted mixed‑effects model identified volume as the strongest evaluated predictor, explaining 23.9% of variance in DSC; orientation and training count each contributed <1%, leaving 61.5% unexplained. These results indicate that deep learning–based segmentation can accurately quantify muscle volume for many upper‑limb muscles but remains constrained for small, low‑contrast forearm muscles.

Gillespie, Samuel

Performance Evaluation of Vertical Federated Machine Learning Against Adversarial Threats on Wide-Area Control System: Preprint

Federated machine learning (FL) is gaining significant popularity to develop cybersecurity solutions in power grids because of its advanced capability to support decentralized data handing at local devices, its privacy preservation, and its low-bandwidth requirement. However, the evolving adversarial machine learning (AML) threats raise significant concerns for the cybersecurity of FL architectures. The FL-based split neural network (SplitNN) achieves high performance through the decentralized training of local neural network models while preserving data privacy across multiple entities. In this paper, we propose a methodology for evaluating the performance of a vertical FLbased anomaly detector against different types of AML attacks, including denial-of-service attacks, adversarial data injection attacks, and replay attacks on the trained local models deployed in the grid network. For a case study, we consider the modified IEEE 13-bus system, and we develop SplitNN-based binary and multiclass classification models to detect, locate, and identify different types of data integrity attacks on the volt-watt control with two pooling layers: maximum pooling and AvgPool. Our experimental results, computed through performance metrics, reveal that the severity of these AML attacks varies with the integrated pooling mechanism, the type of classification model, and the nature of the cyberattack. Further, the AML attacks negatively impacted the prediction time per sample for the pretrained SplitNN during the online testing.

adversarial threats

Empirical Analysis and Automated Classification of Security Bug Reports

With the ever expanding amount of sensitive data being placed into computer systems, the need for effective cybersecurity is of utmost importance. However, there is a shortage of detailed empirical studies of security vulnerabilities from which cybersecurity metrics and best practices could be determined. This thesis has two main research goals: (1) to explore the distribution and characteristics of security vulnerabilities based on the information provided in bug tracking systems and (2) to develop data analytics approaches for automatic classification of bug reports as security or non-security related. This work is based on using three NASA datasets as case studies. The empirical analysis showed that the majority of software vulnerabilities belong only to a small number of types. Addressing these types of vulnerabilities will consequently lead to cost efficient improvement of software security. Since this analysis requires labeling of each bug report in the bug tracking system, we explored using machine learning to automate the classification of each bug report as a security or non-security related (two-class classification), as well as each security related bug report as specific security type (multiclass classification). In addition to using supervised machine learning algorithms, a novel unsupervised machine learning approach is proposed. An ac- curacy of 92%, recall of 96%, precision of 92%, probability of false alarm of 4%, F-Score of 81% and G-Score of 90% were the best results achieved during two-class classification. Furthermore, an accuracy of 80%, recall of 80%, precision of 94%, and F-score of 85% were the best results achieved during multiclass classification.

Cybersecurity

Measurements of inclusive and differential cross sections for top quark production in association with a Z boson in proton-proton collisions at $\sqrt{s} $ = 13 TeV

Measurements are presented of inclusive and differential cross sections for Z boson associated production of top quark pairs ($ \textrm{t}\overline{\textrm{t}}\textrm{Z} $) and single top quarks (tZq or tWZ). The data were recorded in proton-proton collisions at a center-of-mass energy of 13 TeV, corresponding to an integrated luminosity of 138 fb$^{−1}$. Events with three or more leptons, electrons or muons, are selected and a multiclass deep neural network is used to separate three event categories, the $ \textrm{t}\overline{\textrm{t}}\textrm{Z} $ and tWZ processes, the tZq process, and the backgrounds. A profile likelihood approach is used to unfold the differential cross sections, to account for systematic uncertainties, and to determine the correlations between the two signal categories in one global fit. The inclusive cross sections for a dilepton invariant mass between 70 and 110 GeV are measured to be 1.14 ± 0.07 pb for the sum of $ \textrm{t}\overline{\textrm{t}}\textrm{Z} $ and tWZ, and 0.81 ± 0.10 pb for tZq, in good agreement with theoretical predictions.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

HIV drug resistance during antiretroviral therapy scale-up in Uganda, 2012–19: a population-based, longitudinal study

Background With scale-up of antiretroviral therapy (ART) in sub-Saharan Africa, increasing pretreatment HIV drug resistance has been reported; however, the broader effect of ART expansion on population-level resistance patterns remains insufficiently quantified. We aimed to estimate the longitudinal prevalence of drug resistance and resistance-conferring mutations. Methods This study used data collected as part of the Rakai Community Cohort Study (RCCS), an open population-based census and cohort study conducted in southern Uganda. At each survey round, residents aged 15–49 years are invited to participate and receive a structured questionnaire that obtains sociodemographic, behavioural, and health information, including self-reported past and current ART use. Voluntary HIV testing is conducted using a rapid test algorithm and a venous blood sample. People with HIV provide samples for viral load quantification and deep sequencing. We analysed RCCS survey, HIV viral load, and deep sequencing (which was used to predict resistance) data from five survey rounds. The key outcomes were the population prevalence of viraemic people with HIV with non-nucleoside reverse transcriptase inhibitor (NNRTI), nucleoside reverse transcriptase inhibitor (NRTI), protease inhibitor, or multiclass resistance among all participants (regardless of HIV serostatus) in the 2015 and 2017 surveys. Prevalence of class-specific resistance and resistance-conferring substitutions were estimated using robust log-Poisson regression. Findings Between Aug 10, 2011, and Nov 4, 2020, there were 43 361 participants in the RCCS and 7923 (18·27%) people with HIV. Over five survey rounds, 93 622 participant visits occurred, among which 17 460 (18·65%) were from people with HIV. Over the analysis period, the median age of study participants remained similar (28 years [22–35] in 2012 and 29 years [21–38] in 2019). Sufficient data were available to reliably genotype 4072 (90·03%) of 4523 participant visits from 3407 people with HIV for at least one drug. Overall population prevalence of resistance contributed by viraemic pretreatment people with HIV decreased between 2012 and 2017 from 0·56% (95% CI 0·42–0·75) to 0·25% (0·18–0·33) for NNRTI and from 0·24% (0·15–0·37) to 0·05% (0·02–0·10) for NRTI (prevalence ratio 0·44 [0·29–0·68] for NNRTI and 0·21 [0·09–0·47] for NRTI). Between 2012 and 2017, NNRTI resistance among viraemic pretreatment people with HIV increased from 4·86% (3·69–6·42) to 9·61% (7·27–12·7; prevalence ratio 1·98 [1·34–2·91]). The prevalence of NNRTI and NRTI resistance was substantially higher among viraemic treatment-experienced people with HIV (51·49% [46·24–57·34] for NNRTI and 36·46% [30·06–44·22] for NRTI in 2017) than among pretreatment people with HIV. NNRTI and NRTI resistance was predominantly attributable to rtK103N and rtM184V. inT97A was observed at a similar prevalence among viraemic treatment-experienced (9·96% [6·41–15·48]) and viraemic pretreatment (10·56% [8·01–13·93]) people with HIV; no major dolutegravir resistance mutations were observed. Interpretation Despite rising NNRTI resistance among pretreatment people with HIV, overall population prevalence of pretreatment HIV drug-resistant viraemia decreased due to increasing ART uptake and viral suppression. This finding underscores the crucial role of achieving and maintaining high ART coverage in reducing transmission of drug-resistant HIV. The high prevalence of mutations conferring resistance to components of first-line ART regimens among viraemic people with HIV is potentially concerning. Funding National Institutes of Health, Johns Hopkins University Center for AIDS Research, Bill & Melinda Gates Foundation, and the US Centers for Disease Control and Prevention.

59 BASIC BIOLOGICAL SCIENCES

A unified large language model–based framework for heterogeneous PV image diagnosis

With advances in imaging technologies, modern photovoltaic (PV) systems generate large volumes of heterogeneous image data, including visible, electroluminescence (EL), and infrared (IR) images. Existing PV image analysis models, particularly deep learning approaches, are typically task-specific and lack cross-modality generalization. To address this limitation, this paper proposes an open-source large language model (LLM)–based unified framework for heterogeneous PV image diagnostics. Through task-aware diagnostic prompting, the framework enables analysis of visible, EL, and IR images within a single pipeline, supporting both zero-shot and few-shot inference and binary and multiclass classification. It is compatible with state-of-the-art multimodal LLMs, including ChatGPT, Gemini, Claude, Qwen, and CLIP. The framework is evaluated on PV module condition classification (clean, soiling, snow, hail, and bird droppings) using visible images, cell crack detection using EL images, and hotspot detection using IR images. GPT-5.1 in few-shot mode achieves the best performance, with classification accuracy exceeding 97.3%. Open-source models such as Qwen and CLIP also deliver competitive results on visible images (around 90% accuracy), though their performance is more limited on EL and IR modalities. On the full ELPV dataset, the framework achieves 83.5% zero-shot accuracy, within 2.8% of the supervised CNN baseline, confirming scalability to larger benchmarks. Practical aspects such as reproducibility, response latency, and confidence estimation are systematically analyzed. The framework operates across PV image modalities without modality- or task-specific training, making it well suited as a rapid pre-screening tool to support downstream detailed diagnostics. A benchmark dataset of diverse labeled PV images is also released.

Li, Baojie

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi

CaloChallenge 2022: a community challenge for fast calorimeter simulation

Here, we present the results of the ‘Fast Calorimeter Simulation Challenge 2022’—the CaloChallenge. We study state-of-the-art generative models on four calorimeter shower datasets of increasing dimensionality, ranging from a few hundred voxels to a few tens of thousand voxels. The 31 individual submissions span a wide range of current popular generative architectures, including variational autoencoders (VAEs), generative adversarial networks (GANs), normalizing flows, diffusion models, and models based on conditional flow matching. We compare all submissions in terms of quality of generated calorimeter showers, as well as shower generation time and model size. To assess the quality we use a broad range of different metrics including differences in one-dimensional histograms of observables, KPD/FPD scores, AUCs of binary classifiers, and the log-posterior of a multiclass classifier. The results of the CaloChallenge provide the most complete and comprehensive survey of cutting-edge approaches to calorimeter fast simulation to date. In addition, our work provides a uniquely detailed perspective on the important problem of how to evaluate generative models. As such, the results presented here should be applicable for other domains that use generative AI and require fast and faithful generation of samples in a large phase space.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND