Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Myths and legends in learning classification rules

This paper is a discussion of machine learning theory on empirically learning classification rules. The paper proposes six myths in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, 'universal' learning algorithms, and interactive learnings. Some of the problems raised are also addressed from a Bayesian perspective. The paper concludes by suggesting questions that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

Optimization of Designs for Nanotube-based Scanning Probes

Optimization of designs for nanotube-based scanning probes, which may be used for high-resolution characterization of nanostructured materials, is examined. Continuum models to analyze the nanotube deformations are proposed to help guide selection of the optimum probe. The limitations on the use of these models that must be accounted for before applying to any design problem are presented. These limitations stem from the underlying assumptions and the expected range of nanotube loading, end conditions, and geometry. Once the limitations are accounted for, the key model parameters along with the appropriate classification of nanotube structures may serve as a basis for the design optimization of nanotube-based probe tips.

Harik, V. M.↗

Spoken Language Processing in the Clarissa Procedure Browser

Clarissa, an experimental voice enabled procedure browser that has recently been deployed on the International Space Station, is as far as we know the first spoken dialog system in space. We describe the objectives of the Clarissa project and the system's architecture. In particular, we focus on three key problems: grammar-based speech recognition using the Regulus toolkit; methods for open mic speech recognition; and robust side-effect free dialogue management for handling undos, corrections and confirmations. We first describe the grammar-based recogniser we have build using Regulus, and report experiments where we compare it against a class N-gram recogniser trained off the same 3297 utterance dataset. We obtained a 15% relative improvement in WER and a 37% improvement in semantic error rate. The grammar-based recogniser moreover outperforms the class N-gram version for utterances of all lengths from 1 to 9 words inclusive. The central problem in building an open-mic speech recognition system is being able to distinguish between commands directed at the system, and other material (cross-talk), which should be rejected. Most spoken dialogue systems make the accept/reject decision by applying a threshold to the recognition confidence score. NASA shows how a simple and general method, based on standard approaches to document classification using Support Vector Machines, can give substantially better performance, and report experiments showing a relative reduction in the task-level error rate by about 25% compared to the baseline confidence threshold method. Finally, we describe a general side-effect free dialogue management architecture that we have implemented in Clarissa, which extends the "update semantics'' framework by including task as well as dialogue information in the information state. We show that this enables elegant treatments of several dialogue management problems, including corrections, confirmations, querying of the environment, and regression testing.

Rayner, M.↗

Charged Particle Tracking via Edge-Classifying Interaction Networks

Recent work has demonstrated that geometric deep learning methods such as graph neural networks (GNNs) are well suited to address a variety of reconstruction problems in high-energy particle physics. In particular, particle tracking data are naturally represented as a graph by identifying silicon tracker hits as nodes and particle trajectories as edges, given a set of hypothesized edges, edge-classifying GNNs identify those corresponding to real particle trajectories. In this work, we adapt the physics-motivated interaction network (IN) GNN toward the problem of particle tracking in pileup conditions similar to those expected at the high-luminosity Large Hadron Collider. Assuming idealized hit filtering at various particle momenta thresholds, we demonstrate the IN’s excellent edge-classification accuracy and tracking efficiency through a suite of measurements at each stage of GNN-based tracking: graph construction, edge classification, and track building. The proposed IN architecture is substantially smaller than previously studied GNN tracking architectures; this is particularly promising as a reduction in size is critical for enabling GNN-based tracking in constrained computing environments. Furthermore, the IN may be represented as either a set of explicit matrix operations or a message passing GNN. Efforts are underway to accelerate each representation via heterogeneous computing resources towards both high-level and low-latency triggering applications.

accelerator physics↗

An empirical study of scanner system parameters

The selection of the current combination of parametric values (instantaneous field of view, number and location of spectral bands, signal-to-noise ratio, etc.) of a multispectral scanner is a complex problem due to the strong interrelationship these parameters have with one another. The study was done with the proposed scanner known as Thematic Mapper in mind. Since an adequate theoretical procedure for this problem has apparently not yet been devised, an empirical simulation approach was used with candidate parameter values selected by the heuristic means. The results obtained using a conventional maximum likelihood pixel classifier suggest that although the classification accuracy declines slightly as the IFOV is decreased this is more than made up by an improved mensuration accuracy. Further, the use of a classifier involving both spatial and spectral features shows a very substantial tendency to resist degradation as the signal-to-noise ratio is decreased. And finally, further evidence is provided of the importance of having at least one spectral band in each of the major available portions of the optical spectrum.

Landgrebe, D.↗

Nonparametric analysis of Minnesota spruce and aspen tree data and LANDSAT data

The application of nonparametric methods in data-intensive problems faced by NASA is described. The theoretical development of efficient multivariate density estimators and the novel use of color graphics workstations are reviewed. The use of nonparametric density estimates for data representation and for Bayesian classification are described and illustrated. Progress in building a data analysis system in a workstation environment is reviewed and preliminary runs presented.

Scott, D. W.↗

Asteroid proper elements and secular resonances

In a series of papers (e.g., Knezevic, 1991; Milani and Knezevic, 1990; 1991) we reported on the progress we were making in computing asteroid proper elements, both as regards their accuracy and long-term stability. Additionally, we reported on the efficiency and 'intelligence' of our software. At the same time, we studied the associated problems of resonance effects, and we introduced the new class of 'nonlinear' secular resonances; we determined the locations of these secular resonances in proper-element phase space and analyzed their impact on the asteroid family classification. Here we would like to summarize the current status of our work and possible further developments.

Knezevic, Zoran↗

Cascaded VLSI neural network architecture for on-line learning

High-speed, analog, fully-parallel and asynchronous building blocks are cascaded for larger sizes and enhanced resolution. A hardware-compatible algorithm permits hardware-in-the-loop learning despite limited weight resolution. A comparison-intensive feature classification application has been demonstrated with this flexible hardware and new algorithm at high speed. This result indicates that these building block chips can be embedded as application-specific-coprocessors for solving real-world problems at extremely high data rates.

Duong, Tuan A.↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

Classification of Orbits in Poincare Maps Using Machine Learning

Poincare plots, also called Poincare maps, are used by plasma physicists to understand the behavior of magnetically confined plasma in numerical simulations of a tokamak. These plots are created by the intersection of field lines with a two-dimensional poloidal plane that is perpendicular to the axis of the torus representing the tokamak. A plot is composed of multiple orbits, each created by a different field line as it goes around the torus. Each orbit can have one of four distinct shapes, or classes, that indicate changes in the topology of the magnetic fields confining the plasma. Given the (x, y) coordinates of the points that form an orbit, the analysis task is to assign a class to the orbit, a task that appears ideally suited for a machine learning approach. In this paper, we describe how we overcame two major challenges in solving this problem - creating a high-quality training set, with few mislabeled orbits, and converting the coordinates of the points into features that are discriminating, despite the variation within the orbits of a class and the apparent similarities between orbits of different classes. Our automated approach is not only more objective and accurate than visual classification, but is also less tedious, making it easier for plasma physicists to analyze the topology of magnetic fields from numerical simulations of the tokamak.

97 MATHEMATICS AND COMPUTING↗

Homotopy characterization of non-Hermitian Hamiltonians

We revisit the problem of classifying topological band structures in non-Hermitian systems. Recently, a solution has been proposed, which is based on redefining the notion of energy band gap in two different ways, leading to the so-called “point-gap” and “line-gap” schemes. However, simple Hamiltonians without band degeneracies can be constructed which correspond to neither of the two schemes. Here, we resolve this shortcoming of the existing classifications by developing the most general topological characterization of non-Hermitian bands for systems without a symmetry. Our approach, which is based on homotopy theory, makes no particular assumptions on the band gap, and predicts significant extensions to the previous classification frameworks. In particular, we show that the one-dimensional invariant generalizes from $\mathbb{Z}$ winding number to the non-Abelian braid group, and that depending on the braid group invariants, the two-dimensional invariants can be cyclic groups $\mathbb{Z}_n$ (rather than $\mathbb{Z}$ Chern number). Finally, we interpret these results in terms of a correspondence with gapless systems, and we illustrate them in terms of analogies with other problems in band topology, namely, the fragile topological invariants in Hermitian systems and the topological defects and textures of nematic liquids.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A Siamese CNN + KNN-Based Classification Framework for Non-intrusive Load Monitoring

Through the development of smart grids, programs such as demand side response, have been presented as auxiliary services to the real-time operation of distributed networks. In order to provide consumers information on their energy consumption, so that a modulation in consumption is possible, non-intrusive load monitoring has been introduced as an solution to this pattern recognition problem. Non-intrusive load monitoring enables the modeling of electrical loads connected to the low-voltage system, considering only a single measurement point. Presented state-of-the-art solutions though, consider availability of data as well as representation of all possible classes of the environment. This is of course a most conservative hypothesis, since in real-life applications availability of such data is much difficult, as well as the dynamic behavior of models is implicitly evolving in time. Here, a framework that uses neural Siamese networks with k-nearest neighbor clustering is presented toward non-intrusive load monitoring. Online learning feature is implemented, which relaxes the hypothesis of data requirements as well addresses the evolving nature of load profile. k-nearest clustering allows nonlinear characteristic space modelling. Test results using synthetics and real-life data show that the solution, besides obtaining a good generalizability in the classification, also obtained results with an accuracy of 95.77%.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Dynamic Reduction Network for Point Clouds

Classifying whole images is a classic problem in machine learning, and graph neural networks are a powerful methodology to learn highly irregular geometries. It is often the case that certain parts of a point cloud are more important than others when determining overall classification. On graph structures this started by pooling information at the end of convolutional filters, and has evolved to a variety of staged pooling techniques on static graphs. In this paper, a dynamic graph formulation of pooling is introduced that removes the need for predetermined graph structure. It achieves this by dynamically learning the most important relationships between data via an intermediate clustering. The network architecture yields interesting results considering representation size and efficiency. It also adapts easily to a large number of tasks from image classification to energy regression in high energy particle physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Planning applications in East Central Florida

The author has identified the following significant results. This is a study of applications of ERTS data to planning problems, especially as applicable to East Central Florida. The primary method has been computer analysis of digital data, with visual analysis of images serving to supplement the digital analysis. The principal method of analysis was supervised maximum likelihood classification, supplemented by density slicing and mapping of ratios of band intensities. Land-use maps have been prepared for several urban and non-urban sectors. Thematic maps have been found to be a useful form of the land-use maps. Change-monitoring has been found to be an appropriate and useful application. Mapping of marsh regions has been found effective and useful in this region. Local planners have participated in selecting training samples and in the checking and interpretation of results.

Hannah, J. W.↗

Applications of remote sensing to hydrologic planning

The transfer of LANDSAT remote sensing technology from the research sector to user operational applications requires demonstration of the utility and accuracy of LANDSAT data in solving real problems. This report describes such a demonstration project in the area of water resources, specifically the estimation of non-point source pollutant loads. Non-point source pollutants were estimated from land cover data from LANDSAT images. Classification accuracies for three small watersheds were above 95%. Land cover was converted to pollutant loads for a fourth watershed through the use of coefficients relating significant pollutants to land use and storm runoff volume. These data were input into a simulator model which simulated runoff from average rainfall. The result was the estimation of monthly expected pollutant loads for the 17 subbasins comprising the Magothy watershed.

Loats, H., Jr.↗

Lake water quality mapping from Landsat

In the project described remote sensing was used to check the quality of lake waters. The lakes of three Landsat scenes were mapped with the Bendix MDAS multispectral analysis system. From the MDAS color coded maps, the lake with the worst algae problem was easily located. The lake was closely checked, and the presence of 100 cows in the springs which fed the lake could be identified as the pollution source. The laboratory and field work involved in the lake classification project is described.

Scherz, J. P.↗

Landsat digital data application to forest vegetation and land use classification in Minnesota

Landsat digital data were used to map eleven categories of land cover in north central Minnesota. The classification accuracy of these maps was found to be very low and they were not adequate for use by field level resource managers. A discussion of the advantages and disadvantages of various processing systems, different algorithms, and the problems in selecting training sets, is included.

Mead, R. A.↗

Large Area Crop Inventory Experiment (LACIE). LACIE phase 1 and phase 2 accuracy assessment

The author has identified the following significant results. The initial CAS estimates, which were made for each month from April through August, were considerably higher than the USDA/SRS estimates. This was attributed to: (1) the practice of considering bare ground as potential wheat and counting it as wheat; (2) overestimation of the wheat proportions in segments having only a small amount of wheat; and (3) the classification of confusion crops as wheat. At the end of the season most of the segments were reworked using improved methods based on experience gained during the season. In particular, new procedures were developed to solve the three problems listed above. These and other improvements used in the rework experiment resulted in at-harvest estimates that were much closer to the USDA/SRS estimates than those obtained during the regular season.

Source record↗