Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gaussian Process Classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Physics-Based Machine Learning Methods for U-235 Forensics Signatures

Signatures of low-intensity U-235 sources have been recently studied by utilizing a variety of machine learning (ML) classifiers using features derived from gamma spectral measurements collected under structured campaigns. Several ML classifiers, such as ensemble of tress and classification trees, revealed misleadingly-optimistic training error due to over-fitting, and furthermore, their performance is not directly relatable to the physical properties due to their data-driven, opaque designs. We present a regression-based ML method that first estimates the inverse distance to the source and then utilizes a threshold to infer its presence, by representing the background as a source located at an infinite distance. For the inverse distance estimation, we study the ensemble of trees and Gaussian process regression methods, and a hyper parameter auto-tuning and selection method that employs five regression estimators. These methods avoid the over-fitting observed in several ML classifiers, while providing the classification error nearly comparable to them based on independent test data. Their error is directly related to estimates of the inverse physical distance to source, and the precision of error determines the seperability property that determines the false alarm and missed detection rates. The property of monotonic decrease of the source strength with increasing detector distance combined with Poisson distribution of measurements is utilized to analytically validate these methods by deriving the generalization equations of underlying regression methods.

Rao, Nageswara↗

Automated Classification of Transient Contamination in Stationary Acoustic Data

An automated procedure for the classification of transient contamination of stationary acoustic data is proposed and analyzed. The procedure requires the assumption that the stationary acoustic data of interest can be modeled as a band-limited, Gaussian random process. It also requires that the transient contamination be of higher variance than the acoustic data of interest. When these assumptions are satisfied, it is a blind separation procedure, aside from the initial input specifying how to subdivide the time series of interest. No a priori threshold criterion is required. Simulation results show that for a sufficient number of blocks, the method performs well, as long as the occasional false positive or false negative is acceptable. The effectiveness of the procedure is demonstrated with an application to experimental wind tunnel acoustic test data which are contaminated by hydrodynamic gusts.

binary classification↗

Attention-Augmented Parametric Kernel Graph Neural Network (APKGNN) for Node Classification

We present a new graph neural network, the Attention-based Parametric-Kernel augmented Graph Neural Network (APKGNN), developed for node classification tasks. Despite extensive work on modeling multi-faceted relationships between connected nodes of a graph, the effect of attention on edge features mapped to relationships has not yet been analyzed through learning representation. This study derives such an attention vector by first calculating node features corresponding to endpoints of an edge and then aggregating these with extracted local intrinsic patches of a given graph to generate augmented local patch vectors. This process uses a parametric kernel based on Gaussian mixture models (GMMs) to embed local neighborhoods of the graph in local patches. The patch vectors then convolve with the above node features to produce an updated node representation. We show that this new learning representation (APKGNN) achieves higher node classification accuracy on tasks - both standard benchmarks (Cora, PubMed, Citeseer) and new experimental short text corpora where nodes correspond to text documents and words. This implementation of the GNN convolution layer outperforms state-of-the-art (SOTA) algorithms, achieving higher training, validation, and test accuracy by a significant margin on three standard benchmark data sets under both SOTA experimental settings and those for new testbeds.

Bose, Avishek↗

Evaluating Gaussian process metamodels and sequential designs for noisy level set estimation

Abstract We consider the problem of learning the level set for which a noisy black-box function exceeds a given threshold. To efficiently reconstruct the level set, we investigate Gaussian process (GP) metamodels. Our focus is on strongly stochastic simulators, in particular with heavy-tailed simulation noise and low signal-to-noise ratio. To guard against noise misspecification, we assess the performance of three variants: (i) GPs with Student- t observations; (ii) Student- t processes (TPs); and (iii) classification GPs modeling the sign of the response. In conjunction with these metamodels, we analyze several acquisition functions for guiding the sequential experimental designs, extending existing stepwise uncertainty reduction criteria to the stochastic contour-finding context. This also motivates our development of (approximate) updating formulas to efficiently compute such acquisition functions. Our schemes are benchmarked by using a variety of synthetic experiments in 1–6 dimensions. We also consider an application of level set estimation for determining the optimal exercise policy of Bermudan options in finance.

97 MATHEMATICS AND COMPUTING↗

Ensemble models for circuit topology estimation, fault detection and classification in distribution systems

This paper presents a methodology for simultaneous fault detection, classification, and topology estimation for adaptive protection of distribution systems. The methodology estimates the probability of the occurrence of each one of these events by using a hybrid structure that combines three sub-systems, a convolutional neural network for topology estimation, a fault detection based on predictive residual analysis, and a standard support vector machine with probabilistic output for fault classification. The input to all these sub-systems is the local voltage and current measurements. A convolutional neural network uses these local measurements in the form of sequential data to extract features and estimate the topology conditions. The fault detector is constructed with a Bayesian stage (a multitask Gaussian process) that computes a predictive distribution (assumed to be Gaussian) of the residuals using the input. Since the distribution is known, these residuals can be transformed into a Standard distribution, whose values are then introduced into a one-class support vector machine. The structure allows using a one-class support vector machine without parameter cross-validation, so the fault detector is fully unsupervised. Finally, a support vector machine uses the input to perform the classification of the fault types. All three sub-systems can work in a parallel setup for both performance and computation efficiency. In conclusion, we test all three sub-systems included in the structure on a modified IEEE123 bus system, and we compare and evaluate the results with standard approaches.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Automated Classification of Transient Contamination in Stationary Acoustic Data

An automated procedure for the classification of transient contamination of stationary acoustic data is proposed and analyzed. The procedure requires the assumption that the stationary acoustic data of interest can be modeled as a band-limited, Gaussian random process. It also requires that the transient contamination be of higher variance than the acoustic data of interest. When these assumptions are satisfied, it is a blind separation procedure, aside from the initial input specifying how to subdivide the time series of interest. No a priori threshold criterion is required. Simulation results show that for a sufficient number of blocks, the method performs well, as long as the occasional false positive or false negative is acceptable. The effectiveness of the procedure is demonstrated with an application to experimental wind tunnel acoustic test data which are contaminated by hydrodynamic gusts.

Bahr, Christopher J.↗

Synthetic aperture radar system design for random field classification

An optimum design study is carried out for synthetic aperture radar systems intended for classifying randomly reflecting areas (such as agricultural fields) characterized by a reflectivity density spectral density. The problem solution is obtained, neglecting interfield interference and assuming areas of known configuration and location, as well as a certain Gaussian signal field property. The optimum processor is nonlinear, but includes conventional matched filter processing. A set of summary design curves is plotted, and is applied to the design of a satellite synthetic aperture radar system.

Harger, R. O.↗

Learned adaptive properties for mitigation of weight perturbations in embedded spiking networks

Recent years have seen an increased importance of neural network inference in edge-based scenarios, which impose size and power constraints requiring novel computing devices. These same edge scenarios may require operating over long periods of time, or exposure to extreme environments, resulting in a drift of neural network weights that cause degraded performance. In searching for ways to develop neural network approaches that perform robustly under these conditions, we propose a biologically-inspired mechanism for the dynamic adaptation of within-neuron parameters that is guided by a global context signal carrying information about perturbations and variability in incoming stimuli. Specifically, we demonstrate that adaptive voltage thresholds or neuronal time constants, when informed by a global context signal, can enable network-level mechanisms to recover from perturbed synaptic weights. Consistent with prior literature, the context-modulated approach is effective for recurrent, but not feedforward networks, by modulating network level dynamics. We demonstrate this approach successfully recovers performance in image classification tasks and spatiotemporal tracking tasks under idealized and Gaussian noise as well as for realistic perturbations from a memristive device when exposed to ionizing radiation. Finally, we discuss how this approach enables the design of robust and energy-efficient neuromorphic systems that perform well, even in resource-constrained scenarios with extreme environments such as edge processing.

context modulation↗

Condition Monitoring for Helicopter Data

In this paper the classical "Westland" set of empirical accelerometer helicopter data is analyzed with the aim of condition monitoring for diagnostic purposes. The goal is to determine features for failure events from these data, via a proprietary signal processing toolbox, and to weigh these according to a variety of classification algorithms. As regards signal processing, it appears that the autoregressive (AR) coefficients from a simple linear model encapsulate a great deal of information in a relatively few measurements; it has also been found that augmentation of these by harmonic and other parameters can improve classification significantly. As regards classification, several techniques have been explored, among these restricted Coulomb energy (RCE) networks, learning vector quantization (LVQ), Gaussian mixture classifiers and decision trees. A problem with these approaches, and in common with many classification paradigms, is that augmentation of the feature dimension can degrade classification ability. Thus, we also introduce the Bayesian data reduction algorithm (BDRA), which imposes a Dirichlet prior on training data and is thus able to quantify probability of error in an exact manner, such that features may be discarded or coarsened appropriately.

Wen, Fang↗

On the use of stochastic process-based methods for the analysis of hyperspectral data

Further development in remote sensing technology requires refinement of information system design aspects, i.e., the ability to specify precisely the data to collect and the means to extract increasing amounts of information from the increasingly rich and complex data stream created. One of the principal directions of advance is that data from much larger numbers of spectral bands can be collected, but with significantly increased signal-to-noise ratio. The theory of stochastic or random processes may be applied to the modeling of second-order variations. A multispectral data set with a large number of spectral bands is analyzed using standard pattern recognition techniques. The data were classified using first a single spectral feature, then two, and continuing on with greater and greater numbers of features. Three different classification schemes are used: a standard maximum likelihood Gaussian scheme; the same approach with the mean values of all classes adjusted to be the same; and the use of a minimum distance to means scheme such that mean differences are used.

Landgrebe, David A.↗

Joint cosmic density reconstruction from photometric and spectroscopic samples

ABSTRACT We reconstruct the dark matter density field from spatially overlapping spectroscopic and photometric redshift catalogues through a field-level forward modelling approach. Instead of directly inferring the underlying density field, we find the best-fitting initial Gaussian fluctuations that will evolve into the observed cosmic volume. To account for the substantial uncertainty of photometric redshifts we employ a differentiable continuous Poisson process. As an initial test, we construct a mock based on the upcoming Prime Focus Spectrograph combined with photometric sample modelled on the Subaru Hyper Suprime-Cam. Depending on the statistic of interest, we find improvements in cosmic structure classification equivalent to 50–100 per cent more spectroscopic targets by combining relatively sparse spectroscopic with dense photometric samples.

Horowitz, B.↗

Classification Experiments on Real-World Texture

Many papers have been published concerning the analysis of visual texture and yet, very few application domains use texture for image classification. A possible reason for this low transfer of the technology is the lack of experience and testing in real-world imagery. In this paper, we assess the performance of texture-based classification methods on a number of real-world images relevant to autonomous navigation on cross-country terrain and to autonomous geology. Texture analysis will form part of the closed loop that allows a robotic system to navigate autonomously. We have implemented two different classifiers on features extracted by Gabor filter banks. The first classifier models feature distributions for each texture class using a mixture of Gaussians. Classification is performed using Maximum Likelihood. The second classifier represents local statistics using marginal histograms of the features over a region centered on the pixel to be classified. We measure system performance by comparison to ground truth image labels.

image segmentation↗

Deep Learning of Dark Energy Spectroscopic Instrument Mock Spectra to Find Damped Lyα Systems

We have updated and applied a convolutional neural network (CNN) machine-learning model to discover and characterize damped Ly α systems (DLAs) based on Dark Energy Spectroscopic Instrument (DESI) mock spectra. We have optimized the training process and constructed a CNN model that yields a DLA classification accuracy above 99% for spectra that have signal-to-noise ratios (S/N) above 5 per pixel. The classification accuracy is the rate of correct classifications. This accuracy remains above 97% for lower S/N ≈1 spectra. This CNN model provides estimations for redshift and H i column density with standard deviations of 0.002 and 0.17 dex for spectra with S/N above 3 pixel -1 . Also, this DLA finder is able to identify overlapping DLAs and sub-DLAs. Further, the impact of different DLA catalogs on the measurement of baryon acoustic oscillations (BAO) is investigated. The cosmological fitting parameter result for BAO has less than 0.61% difference compared to analysis of the mock results with perfect knowledge of DLAs. This difference is lower than the statistical error for the first year estimated from the mock spectra: above 1.7%. We also compared the performances of the CNN and Gaussian Process (GP) models. Our improved CNN model has moderately 14% higher purity and 7% higher completeness than an older version of the GP code, for S/N > 3. Both codes provide good DLA redshift estimates, but the GP produces a better column density estimate by 24% less standard deviation. A credible DLA catalog for the DESI main survey can be provided by combining these two algorithms.

79 ASTRONOMY AND ASTROPHYSICS↗

Digital processing of satellite imagery application to jungle areas of Peru

The author has identified the following significant results. The use of clustering methods permits the development of relatively fast classification algorithms that could be implemented in an inexpensive computer system with limited amount of memory. Analysis of CCTs using these techniques can provide a great deal of detail permitting the use of the maximum resolution of LANDSAT imagery. Potential cases were detected in which the use of other techniques for classification using a Gaussian approximation for the distribution functions can be used with advantage. For jungle areas, channels 5 and 7 can provide enough information to delineate drainage patterns, swamp and wet areas, and make a reasonable broad classification of forest types.

Pomalaza, J. C.↗

Automatic classification of soils and vegetation with ERTS-1 data

Preliminary results of a test of a computerized analysis method using ERTS 1 data are presented. The method consisted of a four-spectral-band supervised, maximum likelihood, Gaussian classifier with training statistics derived through a combination of clustering and manual methods. The multivariate analysis method leads to the assignment of each resolution element of the data to one of a preselected set of discrete classes. The data frame was an area over the Texas-Oklahoma border including Lake Texoma. The study suggests that multispectral scanner data coupled with machine processing shows promise for earth surface cover surveys. Futhermore, the processing time is short and consequently the costs are low; a full frame can be analyzed completely within 48 hours.

Landgrebe, D. A.↗

Toward Guided Mutagenesis: Gaussian Process Regression Predicts MHC Class II Antigen Mutant Binding

Antigen-specific immunotherapies (ASI) require successful loading and presentation of antigen peptides into the major histocompatibility complex (MHC) binding cleft. One route of ASI design is to mutate native antigens for either stronger or weaker binding interaction to MHC. Exploring all possible mutations is costly both experimentally and computationally. To reduce experimental and computational expense, here we investigate the minimal amount of prior data required to accurately predict the relative binding affinity of point mutations for peptide-MHC class II (pMHCII) binding. Using data from different residue subsets, we interpolate pMHCII mutant binding affinities by Gaussian process (GP) regression of residue volume and hydrophobicity. We apply GP regression to an experimental data set from the Immune Epitope Database, and theoretical data sets from NetMHCIIpan and Free Energy Perturbation calculations. We find that GP regression can predict binding affinities of nine neutral residues from a six-residue subset with an average R 2 coefficient of determination value of 0.62 ± 0.04 (±95% CI), average error of 0.09 ± 0.01 kcal/mol (±95% CI), and with an receiver operating characteristic (ROC) AUC value of 0.92 for binary classification of enhanced or diminished binding affinity. Similarly, metrics increase to an R2 value of 0.69 ± 0.04, average error of 0.07 ± 0.01 kcal/mol, and an ROC AUC value of 0.94 for predicting seven neutral residues from an eight-residue subset. Our work finds that prediction is most accurate for neutral residues at anchor residue sites without register shift. This work holds relevance to predicting pMHCII binding and accelerating ASI design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

Microstructure prediction for Ti-22Al-25Nb in laser powder bed fusion

This work presents a physics-informed framework for predicting solidification morphology and defect susceptibility in additively manufactured Ti–22Al–25Nb across a broad processing space. The framework integrates solidification microstructure selection (SMS) analysis with a single-track defect-based printability map to establish a unified methodology linking processing parameters to both interfacial morphology and manufacturability. Thermal gradients G and solidification rates R are first computed using the Thermo-Calc Additive Manufacturing (TC-AM) module, a finite-interface-dissipation (FID) phase-field (PF) model coupled with CALPHAD method is then employed to systematically distinguish planar and dendritic regimes as functions of $G$ and $R$. By superimposing the printability map onto the morphology projections, a comprehensive process–structure framework is obtained. Across most processing conditions, the predicted microstructure is predominantly dendritic, while planar growth emerges only under selected laser power $P$ and scan speed $v$ combinations. In addition to morphology classification, the framework quantifies the dendritic area fraction and introduces a width-based morphology descriptor to characterize the spatial extent of planar/dendritic regions within the melt pool. It provides mechanistic insight into the interplay between solidification physics and defect formation, offering practical guidance for parameter selection and microstructural control in Ti–22Al–25Nb additive manufacturing (AM).

36 MATERIALS SCIENCE↗