Engineering PapersSearch

SEARCH · Engineering Papers

Results for “multiclass”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING

Arm and shoulder muscle segmentation in axial MRI with UNet deep learning model

Quantifying individual upper-limb muscle volumes from MRI provides key insight into muscle-specific strength, deficits, and adaptations. Manual delineation is the gold standard but time‑intensive, and the performance of current deep learning approaches, particularly for small or anatomically complex muscles, remains incompletely characterized. We evaluated a state‑of‑the‑art deep learning framework across the entire upper limb and analyzed factors governing segmentation performance, with attention to the forearm. Three previously published MRI datasets (1.5 T, 3D GRE T1‑weighted; total n = 39) spanning young, middle‑aged, and older adults were curated and quality‑checked, including expert manual segmentations for 31 muscles. Following multiclass mask reconstruction, we trained three 3D nnU‑Net multiclass models matched to the muscle subsets present across datasets, using five‑fold cross‑validation and a composite Dice Similarity Coefficient (DSC) + cross entropy loss. Segmentation accuracy was assessed with DSC. Performance varied across muscles (mean DSC = 0.806 ± 0.098), ranging from 0.920 (Deltoid) to 0.461 (Extensor pollicis brevis). In uncertainty‑weighted regressions, muscle volume was positively associated with DSC (R2 = 0.36, p < 0.001), whereas training segmentation count and muscle orientation showed negligible associations (R2 ≤ 0.06). A weighted mixed‑effects model identified volume as the strongest evaluated predictor, explaining 23.9% of variance in DSC; orientation and training count each contributed <1%, leaving 61.5% unexplained. These results indicate that deep learning–based segmentation can accurately quantify muscle volume for many upper‑limb muscles but remains constrained for small, low‑contrast forearm muscles.

Gillespie, Samuel

Performance Evaluation of Vertical Federated Machine Learning Against Adversarial Threats on Wide-Area Control System: Preprint

Federated machine learning (FL) is gaining significant popularity to develop cybersecurity solutions in power grids because of its advanced capability to support decentralized data handing at local devices, its privacy preservation, and its low-bandwidth requirement. However, the evolving adversarial machine learning (AML) threats raise significant concerns for the cybersecurity of FL architectures. The FL-based split neural network (SplitNN) achieves high performance through the decentralized training of local neural network models while preserving data privacy across multiple entities. In this paper, we propose a methodology for evaluating the performance of a vertical FLbased anomaly detector against different types of AML attacks, including denial-of-service attacks, adversarial data injection attacks, and replay attacks on the trained local models deployed in the grid network. For a case study, we consider the modified IEEE 13-bus system, and we develop SplitNN-based binary and multiclass classification models to detect, locate, and identify different types of data integrity attacks on the volt-watt control with two pooling layers: maximum pooling and AvgPool. Our experimental results, computed through performance metrics, reveal that the severity of these AML attacks varies with the integrated pooling mechanism, the type of classification model, and the nature of the cyberattack. Further, the AML attacks negatively impacted the prediction time per sample for the pretrained SplitNN during the online testing.

adversarial threats

Measurements of inclusive and differential cross sections for top quark production in association with a Z boson in proton-proton collisions at $\sqrt{s} $ = 13 TeV

Measurements are presented of inclusive and differential cross sections for Z boson associated production of top quark pairs ($ \textrm{t}\overline{\textrm{t}}\textrm{Z} $) and single top quarks (tZq or tWZ). The data were recorded in proton-proton collisions at a center-of-mass energy of 13 TeV, corresponding to an integrated luminosity of 138 fb$^{−1}$. Events with three or more leptons, electrons or muons, are selected and a multiclass deep neural network is used to separate three event categories, the $ \textrm{t}\overline{\textrm{t}}\textrm{Z} $ and tWZ processes, the tZq process, and the backgrounds. A profile likelihood approach is used to unfold the differential cross sections, to account for systematic uncertainties, and to determine the correlations between the two signal categories in one global fit. The inclusive cross sections for a dilepton invariant mass between 70 and 110 GeV are measured to be 1.14 ± 0.07 pb for the sum of $ \textrm{t}\overline{\textrm{t}}\textrm{Z} $ and tWZ, and 0.81 ± 0.10 pb for tZq, in good agreement with theoretical predictions.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

HIV drug resistance during antiretroviral therapy scale-up in Uganda, 2012–19: a population-based, longitudinal study

Background With scale-up of antiretroviral therapy (ART) in sub-Saharan Africa, increasing pretreatment HIV drug resistance has been reported; however, the broader effect of ART expansion on population-level resistance patterns remains insufficiently quantified. We aimed to estimate the longitudinal prevalence of drug resistance and resistance-conferring mutations. Methods This study used data collected as part of the Rakai Community Cohort Study (RCCS), an open population-based census and cohort study conducted in southern Uganda. At each survey round, residents aged 15–49 years are invited to participate and receive a structured questionnaire that obtains sociodemographic, behavioural, and health information, including self-reported past and current ART use. Voluntary HIV testing is conducted using a rapid test algorithm and a venous blood sample. People with HIV provide samples for viral load quantification and deep sequencing. We analysed RCCS survey, HIV viral load, and deep sequencing (which was used to predict resistance) data from five survey rounds. The key outcomes were the population prevalence of viraemic people with HIV with non-nucleoside reverse transcriptase inhibitor (NNRTI), nucleoside reverse transcriptase inhibitor (NRTI), protease inhibitor, or multiclass resistance among all participants (regardless of HIV serostatus) in the 2015 and 2017 surveys. Prevalence of class-specific resistance and resistance-conferring substitutions were estimated using robust log-Poisson regression. Findings Between Aug 10, 2011, and Nov 4, 2020, there were 43 361 participants in the RCCS and 7923 (18·27%) people with HIV. Over five survey rounds, 93 622 participant visits occurred, among which 17 460 (18·65%) were from people with HIV. Over the analysis period, the median age of study participants remained similar (28 years [22–35] in 2012 and 29 years [21–38] in 2019). Sufficient data were available to reliably genotype 4072 (90·03%) of 4523 participant visits from 3407 people with HIV for at least one drug. Overall population prevalence of resistance contributed by viraemic pretreatment people with HIV decreased between 2012 and 2017 from 0·56% (95% CI 0·42–0·75) to 0·25% (0·18–0·33) for NNRTI and from 0·24% (0·15–0·37) to 0·05% (0·02–0·10) for NRTI (prevalence ratio 0·44 [0·29–0·68] for NNRTI and 0·21 [0·09–0·47] for NRTI). Between 2012 and 2017, NNRTI resistance among viraemic pretreatment people with HIV increased from 4·86% (3·69–6·42) to 9·61% (7·27–12·7; prevalence ratio 1·98 [1·34–2·91]). The prevalence of NNRTI and NRTI resistance was substantially higher among viraemic treatment-experienced people with HIV (51·49% [46·24–57·34] for NNRTI and 36·46% [30·06–44·22] for NRTI in 2017) than among pretreatment people with HIV. NNRTI and NRTI resistance was predominantly attributable to rtK103N and rtM184V. inT97A was observed at a similar prevalence among viraemic treatment-experienced (9·96% [6·41–15·48]) and viraemic pretreatment (10·56% [8·01–13·93]) people with HIV; no major dolutegravir resistance mutations were observed. Interpretation Despite rising NNRTI resistance among pretreatment people with HIV, overall population prevalence of pretreatment HIV drug-resistant viraemia decreased due to increasing ART uptake and viral suppression. This finding underscores the crucial role of achieving and maintaining high ART coverage in reducing transmission of drug-resistant HIV. The high prevalence of mutations conferring resistance to components of first-line ART regimens among viraemic people with HIV is potentially concerning. Funding National Institutes of Health, Johns Hopkins University Center for AIDS Research, Bill & Melinda Gates Foundation, and the US Centers for Disease Control and Prevention.

59 BASIC BIOLOGICAL SCIENCES

A unified large language model–based framework for heterogeneous PV image diagnosis

With advances in imaging technologies, modern photovoltaic (PV) systems generate large volumes of heterogeneous image data, including visible, electroluminescence (EL), and infrared (IR) images. Existing PV image analysis models, particularly deep learning approaches, are typically task-specific and lack cross-modality generalization. To address this limitation, this paper proposes an open-source large language model (LLM)–based unified framework for heterogeneous PV image diagnostics. Through task-aware diagnostic prompting, the framework enables analysis of visible, EL, and IR images within a single pipeline, supporting both zero-shot and few-shot inference and binary and multiclass classification. It is compatible with state-of-the-art multimodal LLMs, including ChatGPT, Gemini, Claude, Qwen, and CLIP. The framework is evaluated on PV module condition classification (clean, soiling, snow, hail, and bird droppings) using visible images, cell crack detection using EL images, and hotspot detection using IR images. GPT-5.1 in few-shot mode achieves the best performance, with classification accuracy exceeding 97.3%. Open-source models such as Qwen and CLIP also deliver competitive results on visible images (around 90% accuracy), though their performance is more limited on EL and IR modalities. On the full ELPV dataset, the framework achieves 83.5% zero-shot accuracy, within 2.8% of the supervised CNN baseline, confirming scalability to larger benchmarks. Practical aspects such as reproducibility, response latency, and confidence estimation are systematically analyzed. The framework operates across PV image modalities without modality- or task-specific training, making it well suited as a rapid pre-screening tool to support downstream detailed diagnostics. A benchmark dataset of diverse labeled PV images is also released.

Li, Baojie

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi

CaloChallenge 2022: a community challenge for fast calorimeter simulation

Here, we present the results of the ‘Fast Calorimeter Simulation Challenge 2022’—the CaloChallenge. We study state-of-the-art generative models on four calorimeter shower datasets of increasing dimensionality, ranging from a few hundred voxels to a few tens of thousand voxels. The 31 individual submissions span a wide range of current popular generative architectures, including variational autoencoders (VAEs), generative adversarial networks (GANs), normalizing flows, diffusion models, and models based on conditional flow matching. We compare all submissions in terms of quality of generated calorimeter showers, as well as shower generation time and model size. To assess the quality we use a broad range of different metrics including differences in one-dimensional histograms of observables, KPD/FPD scores, AUCs of binary classifiers, and the log-posterior of a multiclass classifier. The results of the CaloChallenge provide the most complete and comprehensive survey of cutting-edge approaches to calorimeter fast simulation to date. In addition, our work provides a uniquely detailed perspective on the important problem of how to evaluate generative models. As such, the results presented here should be applicable for other domains that use generative AI and require fast and faithful generation of samples in a large phase space.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Calibrated particle identification for Belle II

We present several efforts aimed at improving charged particle identification at the Belle II experiment. We define an ablation test to quantify and evaluate the impact of each sub-detector on the global particle identification performance. We demonstrate that the performance of the identification scheme can be improved via a simple calibration of the sub-detector likelihoods through a set of per hypothesis, per sub-detector weights. Finally, we present preliminary results on an improved definition of the likelihoods contributed by the electromagnetic calorimeter. A set of multiclass boosted decision trees is trained to exploit the shape of energy depositions of different particle species. In simulated $B\bar{B}$ events, the pion-to-electron and muon fake rates are reduced by 55% and 31% respectively at low-medium momentum.

Hohmann, Marcel [Univ. of Melbourne, Parkville, VI

Divide and conquer: separating the two probabilities in seismic phase picking

There are two fundamental probabilities in the seismic phase picking process—the probability of the existence of a seismic phase (detection probability) and the probability associated with the phase arrival time estimation (timing probability). The nearly ubiquitous approach in developing deep learning phase picking models is to use a kernel, such as a truncated Gaussian, to mask the labelled phase arrival time and train a segmentation model. Once a model is trained, the times of the peaks in the output are taken as phase arrival times (picks), and the height of the peaks are taken as ‘probability’ of the picks. Here, we show that this ‘probability’ represents neither the detection nor the timing probability because this approach forces the output to follow the shape of the kernel. We introduce an approach using two models to estimate these two distinct probabilities. We use a binary classifier with a calibrated confidence to address the detection probability and a multiclass classifier to obtain a probability mass function to address the timing probability. This new approach can make the deep learning-based phase picking process more interpretable and provide options to logically control seismic monitoring workflows.

58 GEOSCIENCES

Method to simultaneously facilitate all jet physics tasks

Machine learning has become an essential tool in jet physics. Due to their complex, high-dimensional nature, jets can be explored holistically by neural networks in ways that are not possible manually. However, innovations in all areas of jet physics are proceeding in parallel. We show that specially constructed machine learning models trained for a specific jet classification task can improve the accuracy, precision, or speed of all other jet physics tasks. This is demonstrated by training on a particular multiclass generation and classification task and then using the learned representation for different generation and classification tasks, for datasets with a different (full) detector simulation, for jets from a different collision system ($pp$ versus $ep$), for generative models, for likelihood ratio estimation, and for anomaly detection. We consider our omnilearn approach thus as a jet-physics foundation model. It is made publicly available for use in any area where state-of-the-art precision is required for analyses involving jets and their substructure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Development of systematic uncertainty-aware neural network trainings for binned-likelihood analyses at the LHC

We propose a neural network training method capable of accounting for the effects of systematic variations of the data model in the training process and describe its extension towards neural network multiclass classification. The procedure is evaluated on the realistic case of the measurement of Higgs boson production via gluon fusion and vector boson fusion in the τ τ decay channel at the CMS experiment. The neural network output functions are used to infer the signal strengths for inclusive production of Higgs bosons as well as for their production via gluon fusion and vector boson fusion. We observe improvements of 12 and 16% in the uncertainty in the signal strengths for gluon and vector-boson fusion, respectively, compared with a conventional neural network training based on cross-entropy.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS