Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

The problem of the structure of the surface of the lunar maria

The results are described of geological and morphological interpretation of large-scale photographs of the surface of lunar maria. A morphological classification is proposed for small craters. It is shown that the classification of a crater in a given morphological category, reflecting the degree of its morphological maturity, is determined by the age of the given crater. A check on the nature of distribution of small craters, as done on the basis of the nearest neighbor method, show that the distribution is governed by a near-random law. It is found that variations of the probability density function for the craters conform to a normal distribution law. An empirical relation is found for the coefficient of variation as a function of crater diameter and the dimensions of the area on which the counts are made. The rockfalls near craters of various morphological classes are described.

Florenskiy, K. P.↗

Hydra: Computer Vision for Online Data Quality Monitoring

Hydra is a system utilizing computer vision for near real-time data quality monitoring. Currently operational across all of Jefferson Lab’s experimental halls, it reduces the workload of shift takers by autonomously monitoring diagnostic plots during experiments. Hydra uses "off-the-shelf" supervised learning technologies and is supported by a comprehensive MySQL database. To simplify access, web apps have been developed to facilitate both labeling and monitoring of Hydra’s inferences. Hydra can connect with the alarm system and incorporates complete historical tracking, enabling it to identify issues that shift takers could miss. When issues are detected, a natural first question is: "Why does Hydra think there is a problem?" To answer, Hydra employs Gradient-weighted Class Activation Maps (GradCAM) to identify regions of the image that are important for the specific classification. This interpretive layer enhances transparency and trustworthiness, which is essential for integration with experiment workflows and operation. The Hydra system, results, and sociological considerations for deployment will be discussed.

Jeske, Torri↗

An Expression for the Transformed Covariance Matrix of Multivariate Normal Populations

The problem is considered of classifying into one of m distinct n-variate classes an arbitrary n-channel multispectral measurement vector x. The classification procedure used is the maximum likelihood procedure. Information loss in compressing the n-channel data to k channels is taken to be the difference in the average interclass divergences (or probability of misclassification) in n-space and in k-space. Data compression is accomplished by kxn linear transformation i.e., multiplication of the spectral n-vector by a kxn matrix of rank k.

Decell, H. P., Jr.↗

Maximum likelihood estimation of label imperfections and its use in the identification of mislabeled patterns

The problem of estimating label imperfections and the use of the estimation in identifying mislabeled patterns is presented. Expressions for the maximum likelihood estimates of classification errors and a priori probabilities are derived from the classification of a set of labeled patterns. Expressions also are given for the asymptotic variances of probability of correct classification and proportions. Simple models are developed for imperfections in the labels and for classification errors and are used in the formulation of a maximum likelihood estimation scheme. Schemes are presented for the identification of mislabeled patterns in terms of threshold on the discriminant functions for both two-class and multiclass cases. Expressions are derived for the probability that the imperfect label identification scheme will result in a wrong decision and are used in computing thresholds. The results of practical applications of these techniques in the processing of remotely sensed multispectral data are presented.

Chittineni, C. B.↗

Satellite recovery - Attitude dynamics of the targets

The problems of categorizing and modeling the attitude dynamics of uncontrolled artificial earth satellites which may be targets in recovery attempts are addressed. Methods of classification presented are based on satellite rotational kinetic energy, rotational angular momentum and orbit and on the type of control present prior to the benign failure of the control system. The use of approximate analytical solutions and 'exact' numerical solutions to the equations governing satellite attitude motions to predict uncontrolled attitude motion is considered. Analytical and numerical results are presented for the evolution of satellite attitude motions after active control termination.

Cochran, J. E., Jr.↗

On the development of an expert system for wheelchair selection

The presentation of wheelchairs for the Multiple Sclerosis (MS) patients involves the examination of a number of complicated factors including ambulation status, length of diagnosis, and funding sources, to name a few. Consequently, only a few experts exist in this area. To aid medical therapists with the wheelchair selection decision, a prototype medical expert system (ES) was developed. This paper describes and discusses the steps of designing and developing the system, the experiences of the authors, and the lessons learned from working on this project. Wheelchair Advisor, programmed in CLIPS, serves as diagnosis, classification, prescription, and training tool in the MS field. Interviews, insurance letters, forms, and prototyping were used to gain knowledge regarding the wheelchair selection problem. Among the lessons learned are that evolutionary prototyping is superior to the conventional system development life-cycle (SDLC), the wheelchair selection is a good candidate for ES applications, and that ES can be applied to other similar medical subdomains.

Madey, Gregory R.↗

Learning time series for intelligent monitoring

We address the problem of classifying time series according to their morphological features in the time domain. In a supervised machine-learning framework, we induce a classification procedure from a set of preclassified examples. For each class, we infer a model that captures its morphological features using Bayesian model induction and the minimum message length approach to assign priors. In the performance task, we classify a time series in one of the learned classes when there is enough evidence to support that decision. Time series with sufficiently novel features, belonging to classes not present in the training set, are recognized as such. We report results from experiments in a monitoring domain of interest to NASA.

Manganaris, Stefanos↗

ML Classifier Fusion for Three Data Streams with Quality Inversely Proportional to Time Resolution

We consider a monitoring scenario of phenomenon using three different streams of measurements whose quality is proportional to their constant inter-arrival times. Each measurement of a stream needs to be binary-classified to reflect the state of interest of the phenomenon. A set of classifiers is separately trained and fused for each stream at its time resolution using measurements collected under known states. We present a machine learning method to fuse the outputs of these fusers to provide a final classification at the finest time resolution. We show that this fused-fusers method provides decisions with likely superior classification probability compared to the best individual classifiers and fused-classifiers. We derive generalization equations that guarantee a superior classification probability of fused-fusers with a confidence probability specified by the classifiers’ generalization equations. We apply these results to study a practical problem of classifying Pu/Np target dissolution events at a radiochemical processing facility using gamma spectral measurements of effluent flows.

Rao, Nageswara↗

Star cluster classification using deep transfer learning with PHANGS- HST

Currently available star cluster catalogues from the Hubble Space Telescope (HST) imaging of nearby galaxies heavily rely on visual inspection and classification of candidate clusters. Here, the time-consuming nature of this process has limited the production of reliable catalogues and thus also post-observation analysis. To address this problem, deep transfer learning has recently been used to create neural network models that accurately classify star cluster morphologies at production scale for nearby spiral galaxies (D ≲ 20 Mpc). Here, we use HST ultraviolet (UV)–optical imaging of over 20 000 sources in 23 galaxies from the Physics at High Angular resolution in Nearby GalaxieS (PHANGS) survey to train and evaluate two new sets of models: (i) distance-dependent models, based on cluster candidates binned by galaxy distance (9–12, 14–18, and 18–24 Mpc), and (ii) distance-independent models, based on the combined sample of candidates from all galaxies. We find that the overall accuracy of both sets of models is comparable to previous automated star cluster classification studies (~60–80 per cent) and shows improvement by a factor of 2 in classifying asymmetric and multipeaked clusters from PHANGS-HST. Somewhat surprisingly, while we observe a weak negative correlation between model accuracy and galactic distance, we find that training separate models for the three distance bins does not significantly improve classification accuracy. We also evaluate model accuracy as a function of cluster properties such as brightness, colour, and spectral energy distribution (SED)-fit age. Based on the success of these experiments, our models will provide classifications for the full set of PHANGS-HST candidate clusters (N ~ 200 000) for public release.

79 ASTRONOMY AND ASTROPHYSICS↗

Selective Sampling for Sensor Type Classification in Buildings

A key barrier to applying any smart technology to a building is the requirement of locating and connecting to the necessary resources among the thousands of sensing and control points, i.e., the metadata mapping problem. Existing solutions depend on exhaustive manual annotation of sensor metadata --- a laborious, costly, and hardly scalable process. To reduce the amount of manual effort required, this paper presents a multi-oracle selective sampling framework to leverage noisy labels from information sources with unknown reliability such as existing buildings, which we refer to as weak oracles, for metadata mapping. This framework involves an interactive process, where a small set of sensor instances are progressively selected and labeled for it to learn how to aggregate the noisy labels as well as to predict sensor types. Two key challenges arise in designing the framework, namely, weak oracle reliability estimation and instance selection for querying. To address the first challenge, we develop a clustering-based approach for weak oracle reliability estimation to capitalize on the observation that weak oracles perform differently in different groups of instances. For the second challenge, we propose a disagreement-based query selection strategy to combine the potential effect of a labeled instance on both reducing classifier uncertainty and improving the quality of label aggregation. We evaluate our solution on a large collection of real-world building sensor data from 5 buildings with more than 11,000 sensors of 18 different types. The experiment results validate the effectiveness of our solution, which outperforms a set of state-of-the-art baselines.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Automated image processing of Landsat II digital data for watershed runoff prediction

Digital image processing of Landsat data from a 230 sq km area was examined as a possible means of generating soil cover information for use in the watershed runoff prediction of Kern County, California. The soil cover information included data on brush, grass, pasture lands and forests. A classification accuracy of 94% for the Landsat-based soil cover survey suggested that the technique could be applied to the watershed runoff estimate. However, problems involving the survey of complex mountainous environments may require further attention

Sasso, R. R.↗

On the error in crop acreage estimation using satellite (LANDSAT) data

The problem of crop acreage estimation using satellite data is discussed. Bias and variance of a crop proportion estimate in an area segment obtained from the classification of its multispectral sensor data are derived as functions of the means, variances, and covariance of error rates. The linear discriminant analysis and the class proportion estimation for the two class case are extended to include a third class of measurement units, where these units are mixed on ground. Special attention is given to the investigation of mislabeling in training samples and its effect on crop proportion estimation. It is shown that the bias and variance of the estimate of a specific crop acreage proportion increase as the disparity in mislabeling rates between two classes increases. Some interaction is shown to take place, causing the bias and the variance to decrease at first and then to increase, as the mixed unit class varies in size from 0 to 50 percent of the total area segment.

Chhikara, R.↗

Pargo Chasma and its relationship to global tectonics

Pargo Chasma was first identified on Pioneer Venus data as a 10,000 km long lineation extending from Atla Regio in the north terminating in the plains south of Phoebe Regio. More recent Magellan data have revealed this feature to be one of the longest chains of coronae so far identified on the planet. Stofan et al have identified 60 coronae and 2 related features associated with this chain; other estimates differ according to the classification scheme adopted, for example Head et al. identify only 29 coronae but 43 arachnoids in the same region. This highlights one of the major problems associated with the preliminary mapping of the Magellan data: there has been an emphasis on identifying particular features on Venus without a universally accepted scheme to classify those features. Nevertheless, Pargo Chasma is clearly identified as a major tectonic belt of global significance. Together with the Artemis-Atla-Beta tectonic zone and the Beta-Phoebe rift belt, Pargo Chasma defines a region on Venus with an unusually high concentration of tectonic and volcanic features. Thus, an understanding of the processes involved in the formation of Pargo Chasma may lend significant insight into the evolution of the region and the planet as a whole. I have produced a detailed 1 to 10 million scale map of Pargo Chasma and the surrounding area from preliminary USGS controlled mosaiced image maps of Venus constructed from Magellan data. In view of the problems highlighted above in relation the efforts already made at identifying a particular set of features I have mapped the region purely on the basis of the geomorphology visible in the magellan data without any attempt at identifying a particular set or class of features. Thus, the map produced distinguishes between areas of different brightness and texture. This has the advantage of highlighting the tectonic fabric of Pargo Chasma and clearly illustrates the close inter-relationship between individual coronae and the surrounding tectonic belts.

Ghail, R. C.↗

A Provably Accurate Randomized Sampling Algorithm for Logistic Regression

In statistics and machine learning, logistic regression is a widely-used supervised learning technique primarily employed for binary classification tasks. When the number of observations greatly exceeds the number of predictor variables, we present a simple, randomized sampling-based algorithm for logistic regression problem that guarantees high-quality approximations to both the estimated probabilities and the overall discrepancy of the model. Our analysis builds upon two simple structural conditions that boil down to randomized matrix multiplication, a fundamental and well-understood primitive of randomized numerical linear algebra. We analyze the properties of estimated probabilities of logistic regression when leverage scores are used to sample observations, and prove that accurate approximations can be achieved with a sample whose size is much smaller than the total number of observations. To further validate our theoretical findings, we conduct comprehensive empirical evaluations. Overall, our work sheds light on the potential of using randomized sampling approaches to efficiently approximate the estimated probabilities in logistic regression, offering a practical and computationally efficient solution for large-scale datasets.

Chowdhury, Agniva↗

Investigations in adaptive processing of multispectral data

Adaptive data processing procedures are applied to the problem of classifying objects in a scene scanned by multispectral sensor. These procedures show a performance improvement over standard nonadaptive techniques. Some sources of error in classification are identified and those correctable by adaptive processing are discussed. Experiments in adaptation of signature means by decision-directed methods are described. Some of these methods assume correlation between the trajectories of different signature means; for others this assumption is not made.

Kriegler, F. J.↗

The large area crop inventory experiment - A major demonstration of space remote sensing

The NASA-U.S. Department of Agriculture Large Area Crop Inventory Experiment (LACIE), aimed at using multispectral remote sensing data from Landsat 1 and 2 to generate accurate annual global crop production forecasts, is discussed. The forecasts take into account meteorological conditions as well as yield and acreage, and may be used to increase the discrimination of U.S. harvest estimates down to regional levels and to provide more accurate early-season predictions. Sample problems involving the determination of wheat harvests and the monitoring of drought conditions are described. Difficulties related to misidentification of abnormally-developing plantations, the automatic classification of homogeneous spectral groups, the computerized generation of colored maps, and the estimation of yields during years when exceptional meteorological conditions prevail are also considered. Samples of Landsat-generated classification maps for Western U.S. and for the Saratov, U.S.S.R. crop regions are given.

Macdonald, R. B.↗

Fuzzy sets, rough sets, and modeling evidence: Theory and Application. A Dempster-Shafer based approach to compromise decision making with multiattributes applied to product selection

The Dempster-Shafer theory of evidence is applied to a multiattribute decision making problem whereby the decision maker (DM) must compromise with available alternatives, none of which exactly satisfies his ideal. The decision mechanism is constrained by the uncertainty inherent in the determination of the relative importance of each attribute element and the classification of existing alternatives. The classification of alternatives is addressed through expert evaluation of the degree to which each element is contained in each available alternative. The relative importance of each attribute element is determined through pairwise comparisons of the elements by the decision maker and implementation of a ratio scale quantification method. Then the 'belief' and 'plausibility' that an alternative will satisfy the decision maker's ideal are calculated and combined to rank order the available alternatives. Application to the problem of selecting computer software is given.

Dekorvin, Andre↗

Automated system function allocation and display format: Task information processing requirements

An important consideration when designing the interface to an intelligent system concerns function allocation between the system and the user. The display of information could be held constant, or 'fixed', leaving the user with the task of searching through all of the available information, integrating it, and classifying the data into a known system state. On the other hand, the system, based on its own intelligent diagnosis, could display only relevant information in order to reduce the user's search set. The user would still be left the task of perceiving and integrating the data and classifying it into the appropriate system state. Finally, the system could display the patterns of data. In this scenario, the task of integrating the data is carried out by the system, and the user's information processing load is reduced, leaving only the tasks of perception and classification of the patterns of data. Humans are especially adept at this form of display processing. Although others have examined the relative effectiveness of alphanumeric and graphical display formats, it is interesting to reexamine this issue together with the function allocation problem. Currently, Johnson Space Center is the test site for an intelligent Thermal Control System (TCS), TEXSYS, being tested for use with Space Station Freedom. Expert TCS engineers, as well as novices, were asked to classify several displays of TEXSYS data into various system states (including nominal and anomalous states). Three different display formats were used: fixed, subset, and graphical. The hypothesis tested was that the graphical displays would provide for fewer errors and faster classification times by both experts and novices, regardless of the kind of system state represented within the display. The subset displays were hypothesized to be the second most effective display format/function allocation condition, based on the fact that the search set is reduced in these displays. Both the subset and the graphic display conditions were hypothesized to be processed more efficiently than the fixed display conditions.

Czerwinski, Mary P.↗