Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Health Monitoring and Prognostics for Electric Aircrafts

As more and more electric vehicles emerge in our daily operation progressively, a very critical challenge lies in the prediction of remaining driving flying time/distance for the flying vehicles. This information is important, particularly in the case of auto vehicles, because such vehicles can become self-aware, autonomously compute its own capabilities, and identify how to best plan and successfully complete vehicular missions safely. In case of electric aircrafts, computing the remaining flying time is also safety-critical, since an aircraft that runs out of power (battery charge) while in the air will eventually lose control leading to catastrophe. To facilitate and solve the prediction problem, awareness of the current health state of the system is key, since it is necessary to perform condition-based predictions. To accurately predict the future state of any system, it is required to possess knowledge of its current health state and future operational conditions. Latest achievements of data-driven algorithms in regression of complex nonlinear functions and classification tasks have generated a growing interest in artificial intelligence for industrial applications. Complex multi-physics models as well as digital twins, once purely built on physics and corresponding simplified lumped parameter iterations, can now benefit from machine learning algorithms to mitigate the lack of understanding of some complex behavior. Given models of the current and future system behavior, a general approach of model-based prognostics can solve the prediction problem and further decision-making. A systematic prediction framework is implemented to identify all possible sources of uncertainty, quantify each of them individually, and mathematically estimate their combined effect on the system-level quantity of interest, in this case, the remaining flying time/distance of the unmanned aircraft. Note - This presentation contains all previously approved and published information.

Systems Health Managent↗

Peri-Net-Pro: the neural processes with quantified uncertainty for crack patterns

Abstract This paper develops a deep learning tool based on neural processes (NPs) called the Peri-Net-Pro, to predict the crack patterns in a moving disk and classifies them according to the classification modes with quantified uncertainties. In particular, image classification and regression studies are conducted by means of convolutional neural networks (CNNs) and NPs. First, the amount and quality of the data are enhanced by using peridynamics to theoretically compensate for the problems of the finite element method (FEM) in generating crack pattern images. Second, case studies are conducted with the prototype microelastic brittle (PMB), linear peridynamic solid (LPS), and viscoelastic solid (VES) models obtained by using the peridynamic theory. The case studies are performed to classify the images by using CNNs and determine the suitability of the PMB, LBS, and VES models. Finally, a regression analysis is performed on the crack pattern images with NPs to predict the crack patterns. The regression analysis results confirm that the variance decreases when the number of epochs increases by using the NPs. The training results gradually improve, and the variance ranges decrease to less than 0.035. The main finding of this study is that the NPs enable accurate predictions, even with missing or insufficient training data. The results demonstrate that if the context points are set to the 10th, 100th, 300th, and 784th, the training information is deliberately omitted for the context points of the 10th, 100th, and 300th, and the predictions are different when the context points are significantly lower. However, the comparison of the results of the 100th and 784th context points shows that the predicted results are similar because of the Gaussian processes in the NPs. Therefore, if the NPs are employed for training, the missing information of the training data can be supplemented to predict the results.

Mathematics↗

Practical applications of remote sensing technology

Land managers increasingly are becoming dependent upon remote sensing and automated analysis techniques for information gathering and synthesis. Remote sensing and geographic information system (GIS) techniques provide quick and economical information gathering for large areas. The outputs of remote sensing classification and analysis are most effective when combined with a total natural resources data base within the capabilities of a computerized GIS. Some examples are presented of the successes, as well as the problems, in integrating remote sensing and geographic information systems. The need to exploit remotely sensed data and the potential that geographic information systems offer for managing and analyzing such data continues to grow. New microcomputers with vastly enlarged memory, multi-fold increases in operating speed and storage capacity that was previously available only on mainframe computers are a reality. Improved raster GIS software systems have been developed for these high performance microcomputers. Vector GIS systems previously reserved for mini and mainframe systems are available to operate on these enhanced microcomputers. One of the more exciting areas that is beginning to emerge is the integration of both raster and vector formats on a single computer screen. This technology will allow satellite imagery or digital aerial photography to be presented as a background to a vector display.

Whitmore, Roy A., Jr.↗

A High Performance Computing Approach to Tree Cover Delineation in 1-m NAIP Imagery Using a Probabilistic Learning Framework

Tree cover delineation is a useful instrument in deriving Above Ground Biomass (AGB) density estimates from Very High Resolution (VHR) airborne imagery data. Numerous algorithms have been designed to address this problem, but most of them do not scale to these datasets, which are of the order of terabytes. In this paper, we present a semi-automated probabilistic framework for the segmentation and classification of 1-m National Agriculture Imagery Program (NAIP) for tree-cover delineation for the whole of Continental United States, using a High Performance Computing Architecture. Classification is performed using a multi-layer Feedforward Backpropagation Neural Network and segmentation is performed using a Statistical Region Merging algorithm. The results from the classification and segmentation algorithms are then consolidated into a structured prediction framework using a discriminative undirected probabilistic graphical model based on Conditional Random Field, which helps in capturing the higher order contextual dependencies between neighboring pixels. Once the final probability maps are generated, the framework is updated and re-trained by relabeling misclassified image patches. This leads to a significant improvement in the true positive rates and reduction in false positive rates. The tree cover maps were generated for the whole state of California, spanning a total of 11,095 NAIP tiles covering a total geographical area of 163,696 sq. miles. The framework produced true positive rates of around 88% for fragmented forests and 74% for urban tree cover areas, with false positive rates lower than 2% for both landscapes. Comparative studies with the National Land Cover Data (NLCD) algorithm and the LiDAR canopy height model (CHM) showed the effectiveness of our framework for generating accurate high-resolution tree-cover maps.

Segments↗

Follow-up Water Quality Analysis at the Little Patuxent River in Anne Arundel County, Maryland andImpact to the Health of the Chesapeake Bay

In accordance with the federal Clean Water Act, it is of utmost importance to identify impaired water bodies and subsequently implement Total Maximum Daily Loads (TMDL) as needed to achieve desirable water quality standards. One of these water bodies include the Little Patuxent River (LPR) in the Anne Arundel and Howard Counties in Maryland, which is the subject of this report. In the 1996, 1998, and then 2006 Maryland Integrated Report, the LPR was found to be impaired by nutrients, bacteria, suspended sediments, and cadmium (Cd). As a result, TMDLs were created to limit factors that were contributing to these problems (Maryland Department of the Environment 2011, December). But later in 2008 and 2009, water quality analyses (WQAs) found that the LPR’s condition had improved. From a Category 5 Cd and phosphorus classification in 1996 it was reduced to Category 2 in 2008 for Cd and in 2009 for phosphorus (Maryland Department of the Environment 2011, December). However, it has been more than a decade since the last WQA and the environmental conditions may have changed. This lack of testing could result in untracked, accumulated damages to the overall health of the LPR. We present testing results for the 13 site locations on the LPR to evaluate turbidity, pH, phosphates, nitrates, water temperature, dissolved oxygen (D.O.), coliform bacteria, and cadmium concentrations in sediments. We will also present next steps including remediation strategies.

Aaban Syed↗

The probabilistic neural network architecture for high speed classification of remotely sensed imagery

In this paper we discuss a neural network architecture (the Probabilistic Neural Net or the PNN) that, to the best of our knowledge, has not previously been applied to remotely sensed data. The PNN is a supervised non-parametric classification algorithm as opposed to the Gaussian maximum likelihood classifier (GMLC). The PNN works by fitting a Gaussian kernel to each training point. The width of the Gaussian is controlled by a tuning parameter called the window width. If very small widths are used, the method is equivalent to the nearest neighbor method. For large windows, the PNN behaves like the GMLC. The basic implementation of the PNN requires no training time at all. In this respect it is far better than the commonly used backpropagation neural network which can be shown to take O(N6) time for training where N is the dimensionality of the input vector. In addition the PNN can be implemented in a feed forward mode in hardware. The disadvantage of the PNN is that it requires all the training data to be stored. Some solutions to this problem are discussed in the paper. Finally, we discuss the accuracy of the PNN with respect to the GMLC and the backpropagation neural network (BPNN). The PNN is shown to be better than GMLC and not as good as the BPNN with regards to classification accuracy.

Chettri, Samir R.↗

Ligand-Based Compound Activity Prediction via Few-Shot Learning

Predicting the activities of new compounds against biophysical or phenotypic assays based on the known activities of one or a few existing compounds is a common goal in early stage drug discovery. This problem can be cast as a “few-shot learning” challenge, and prior studies have developed few-shot learning methods to classify compounds as active versus inactive. However, the ability to go beyond classification and rank compounds by expected affinity is more valuable. We describe Few-Shot Compound Activity Prediction (FS-CAP), a novel neural architecture trained on a large bioactivity data set to predict compound activities against an assay outside the training set, based on only the activities of a few known compounds against the same assay. Our model aggregates encodings generated from the known compounds and their activities to capture assay information and uses a separate encoder for the new compound whose activity is to be predicted. The new method provides encouraging results relative to traditional chemical-similarity-based techniques as well as other state-of-the-art few-shot learning methods in tests on a variety of ligand-based drug discovery settings and data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ERIM progress report on use of ERTS-1 data: Summary report of work on ten tasks

The author has identified the following significant results. Several of the tasks have produced significant results which are summarized: (1) Absolute water depth can be calculated from a ratio of signals from bands MSS 4 and MSS 5. (2) A 13 category terrain feature classification map of Yellowstone National Park has been produced using supervised pattern recognition techniques. (3) ERTS-1 data has been shown to provide a detection and monitoring capability for a number of water quality problems associated with off-shore ocean dumping sites and inland lakes. (4) A corrected ratio of bands MSS-5 and MSS-7 signals has been formed. (5) A concise format has been devised for storing the ratio signatures of geologic rock and mineral materials determined from laboratory reflectance spectra. (6) Results of work in information extraction demonstrate: signal variability exists among ERTS-1 detectors in any one spectral band that will impact users doing quantitative analysis on successive ERTS-1 images; a newly developed computer-aided procedure for correlating ERTS-1 pixels to ground features; the strong influence of atmospheric effects in ERTS-1 data; and area estimation accuracies are better using the ERIM proportion estimation algorithm than for conventional recognition techniques.

Thomson, F. J.↗

Preliminary results of mapping urban land cover with Seasat SAR imagery

The detectability of urban land cover types is explored using digitally processed Seasat SAR imagery of the Denver, Colorado area. Test sites within the metropolitan area were selected to include a cross section of Anderson, et. al. Level II land cover classes and cover types representative of the urban area growth stages. Using the Image 100 interactive processing system each test site was level sliced in an attempt to define specific reflectance boundaries for each cover type and to determine the spectral and spatial characteristics of homogeneous response regions. The rural-urban fringe boundary was readily definable, but a precise Level I and Level II land cover classification was not possible. High density housing could be separated from low density housing and from parks, but reflectance values were often look angle dependent. Confusion between some water and vegetation responses also posed problems.

Henderson, F. M.↗

Myths and legends in learning classification rules

A discussion is presented of machine learning theory on empirically learning classification rules. Six myths are proposed in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, universal learning algorithms, and interactive learning. Some of the problems raised are also addressed from a Bayesian perspective. Questions are suggested that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

Cascaded VLSI neural network architecture for on-line learning

High-speed, analog, fully-parallel, and asynchronous building blocks are cascaded for larger sizes and enhanced resolution. A hardware compatible algorithm permits hardware-in-the-loop learning despite limited weight resolution. A computation intensive feature classification application was demonstrated with this flexible hardware and new algorithm at high speed. This result indicates that these building block chips can be embedded as an application specific coprocessor for solving real world problems at extremely high data rates.

Thakoor, Anilkumar P.↗

Cognitive representations of flight-deck information attributes

A large number of aviation issues are generically being called fligh-deck information management issues, underscoring the need for an organization or classification structure. One objective of this study was to empirically determine how pilots organize flight-deck information attributes and -- based upon that data -- develop a useful taxonomy (in terms of better understanding the problems and directing solutions) for classifying flight-deck information management issues. This study also empirically determined how pilots model the importance of flight-deck information attributes for managing information. The results of this analysis suggest areas in which flight-deck researchers and designers may wish to consider focusing their efforts.

Ricks, Wendell R.↗

Application of remote sensing technology to the solution of problems in the management of resources in Indiana

The author has identified the following significant results. The Lydick, South Bend West, South Bend East, and Osceola quadrangles were successfully classified into twenty-six cover types with a high degree of accuracy. The ability of this computer-assisted classification system to delineate various stages of urban development, from heavy industry to new suburban development, was of particular interest to the planning commission. The classification is clearly more beneficial than the existing agricultural soils and topographic maps, because it shows the current ground cover conditions all on one map. It shows how an area is developing along with the specific type and location of new development. The classification also shows at a glance whether development is taking place in an area suitable for development or if growth is taking place in prime agricultural land, areas of poor foundation material, or other places where development is not desirable.

Weismiller, R. A.↗

Optimal decision trees for categorical data via integer programming

Decision trees have been a very popular class of predictive models for decades due to their interpretability and good performance on categorical features. However, they are not always robust and tend to overfit the data. Additionally, if allowed to grow large, they lose interpretability. In this paper, we present a mixed integer programming formulation to construct optimal decision trees of a prespecified size. We take the special structure of categorical features into account and allow combinatorial decisions (based on subsets of values of features) at each node. Our approach can also handle numerical features via thresholding. Here we show that very good accuracy can be achieved with small trees using moderately-sized training sets. The optimization problems we solve are tractable with modern solvers.

97 MATHEMATICS AND COMPUTING↗

Adversarial classification via distributional robustness with Wasserstein ambiguity

Abstract We study a model for adversarial classification based on distributionally robust chance constraints. We show that under Wasserstein ambiguity, the model aims to minimize the conditional value-at-risk of the distance to misclassification, and we explore links to adversarial classification models proposed earlier and to maximum-margin classifiers. We also provide a reformulation of the distributionally robust model for linear classification, and show it is equivalent to minimizing a regularized ramp loss objective. Numerical experiments show that, despite the nonconvexity of this formulation, standard descent methods appear to converge to the global minimizer for this problem. Inspired by this observation, we show that, for a certain class of distributions, the only stationary point of the regularized ramp loss minimization problem is the global minimizer.

Ho-Nguyen, Nam (ORCID:0000000344647730)↗

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An interdisciplinary analysis of ERTS data for Colorado mountain environments using ADP techniques. An early analysis of ERTS-1 data

There are no author-identified significant results in this report. The principal problem encountered has been the lack of good quality, small scale baseline photography for the test areas. Analysis of the ERTS-1 data for the San Juan Site will emphasize development of a preliminary spectral classification defining grass cover categories, and then selection of subframes for intensive investigation of the forestry, geologic, and hydrologic properties of the area. Primary work has been devoted to the selection and digitization of areas for topographic modeling, and compilation of ground based data maps necessary for computer analysis. Study effort has emphasized: geomorphic features; macro-vegetation; micro-vegetation; snow-hydrology; insect/disease damage; and blow-down. Analysis of a frame of the Lake Texoma area indicates a great deal of potential in the analysis and interpretation of ERTS imagery. Preliminary results of investigations of geologic, forest, range, cropland, and water resources of the area are summarized.

Hoffer, R. M.↗

Near ground level sensing for spatial analysis of vegetation

Measured changes in vegetation indicate the dynamics of ecological processes and can identify the impacts from disturbances. Traditional methods of vegetation analysis tend to be slow because they are labor intensive; as a result, these methods are often confined to small local area measurements. Scientists need new algorithms and instruments that will allow them to efficiently study environmental dynamics across a range of different spatial scales. A new methodology that addresses this problem is presented. This methodology includes the acquisition, processing, and presentation of near ground level image data and its corresponding spatial characteristics. The systematic approach taken encompasses a feature extraction process, a supervised and unsupervised classification process, and a region labeling process yielding spatial information.

Sauer, Tom↗