Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

AHIMSA - Ad hoc histogram information measure sensing algorithm for feature selection in the context of histogram inspired clustering techniques

An algorithm is proposed for dimensionality reduction in the context of clustering techniques based on histogram analysis. The approach is based on an evaluation of the hills and valleys in the unidimensional histograms along the different features and provides an economical means of assessing the significance of the features in a nonparametric unsupervised data environment. The method has relevance to remote sensing applications.

Dasarathy, B. V.

Applications of feature selection

The use of satellite-acquired (LANDSAT) multispectral scanner (MSS) data to conduct an inventory of some crop of economic interest such as wheat over a large geographical area is considered in relation to the development of accurate and efficient algorithms for data classification. The dimension of the measurement space and the computational load for a classification algorithm is increased by the use of multitemporal measurements. Feature selection/combination techniques used to reduce the dimensionality of the problem are described.

Guseman, L. F., Jr.

High resolution spectrophotometry of selected features in the 1.1 micron spectrum of Comet Kohoutek /1973f/

Fabry-Perot interferometry of Comet Kohoutek (1973f) at 1.1 microns with a resolution of 1.2 A showed emission features identified as OH and CN lines in addition to a strong Fraunhofer continuum. Central intensities have been derived for three cases (uniform, Gaussian, and Gaussian plus inverse-rho law) of brightness profiles in the comet coma. Limits for CH4, H2O, HeI, SiI and CrI are also derived.

Meisel, D. D.

Acreage estimation, feature selection, and signature extension dependent upon the maximum likelihood decision rule

A maximum likelihood estimation technique is used for the analysis of agricultural remote sensor data. The m-class probability of misclassification is estimated using unlabeled test samples and labeled training samples. A bound on the variance of a proposed unbiased estimator of the m-class probability of error is derived. The particular case in which each class density is assumed to be a mixture of multivariate normal densities is considered. The extension of spectral signatures in space and time is discussed.

Quirein, J. A.

The role of eigenvalues in linear feature selection theory

A particular measure of pattern class distinction called the average interclass divergence, or more simply, divergence, is considered. Here divergence will be the pairwise average of the expected interclass divergence derived from Hajek's two-class divergence.

Brown, D. R.

The role of eigenvalues in linear feature selection theory

The analysis concerns the role of eigenvalues in determining a particular measure of pattern class distinction called the divergence, which is the pairwise average of the expected interclass divergence derived from Hajek's two-class divergence. Decel and Quirein (1973) showed that there always exists a k x n real matrix B such that the transformation determined by B maximizes divergence in k-dimensional space, and that B can be written as a product involving an orthogonal n x n matrix U. In the present paper it is shown that divergence measure of pattern class distinction does not depend on the eigenvalues of U.

Brown, D. R.

Feature selection via entropy minimization: An example using LANDSAT satellite data

The author has identified the following significant results. The minimum entropy model may provide several useful advantages over traditional techniques for processing LANDSAT data. Total computer time to conduct a complete pattern recognition process is reduced. Subjective (transformed image), as well as statistically derived information is made available to the analyst/user much earlier in the analysis process. A rapid feedback loop in which numerous training set combinations can be tested for difference and representativeness is available. Additional tests of LANDSAT data processing using the minimum entropy model are clearly justified.

Zandonella, A.

Feature selection methodologies using simulated Thematic Mapper data

The present investigation is concerned with the determination of the intrinsic dimensionality of a simulated Thematic Mapper data set. In addition, the effectiveness and sensitivity of 'standard' statistics separability measures (i.e., transformed divergence) is evaluated in comparison to eigenvectors for identifying the optimum subset of the original Thematic Mapper Simulator (TMS) bands for classifying the various cover types. TMS data were collected on May 2, 1979 by NASA's NS001 aircraft multispectral scanner over a bottomland forested area in South Carolina near the city of Camden. It is found that the eigenvectors and eigenvalues of a covariance matrix from a multispectral scanner system (MSS) data set can be obtained without having to actually transform the data.

Dean, M. E.

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914