Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Land cover/use classification of Cairns, Queensland, Australia: A remote sensing study involving the conjunctive use of the airborne imaging spectrometer, the large format camera and the thematic mapper simulator

In an attempt to improve the land cover/use classification accuracy obtainable from remotely sensed multispectral imagery, Airborne Imaging Spectrometer-1 (AIS-1) images were analyzed in conjunction with Thematic Mapper Simulator (NS001) Large Format Camera color infrared photography and black and white aerial photography. Specific portions of the combined data set were registered and used for classification. Following this procedure, the resulting derived data was tested using an overall accuracy assessment method. Precise photogrammetric 2D-3D-2D geometric modeling techniques is not the basis for this study. Instead, the discussion exposes resultant spectral findings from the image-to-image registrations. Problems associated with the AIS-1 TMS integration are considered, and useful applications of the imagery combination are presented. More advanced methodologies for imagery integration are needed if multisystem data sets are to be utilized fully. Nevertheless, research, described herein, provides a formulation for future Earth Observation Station related multisensor studies.

Heric, Matthew↗

Supervised Machine Learning Approach for Classifying Earth Science Publications

The data collections archived and distributed by the GES DISC NASA data center are widely utilized for various Earth Science studies. As these collections are created, many research works are published regarding these collections' algorithms, their validation, and their applications. As NASA data centers collect these publications for public use, it is helpful to categorize them based on how they relate to their associated datasets. Specifically, whether the publication linked to the GES DISC dataset is using it for applicational research, describing the algorithm used for the dataset creation, validating the dataset, or providing a general overview of the data collection. Currently, this process requires simple manual labeling, and as such, it may be possible to solve via automation. To approach this problem, machine learning classifiers were developed to predict a publication's category. Manually labeled publications were used as the training data for the supervised machine learning algorithms, specifically Random Forest and Multinomial Naïve Bayes. After balancing the dataset and implementing the Multinomial Naïve Bayes algorithm, the classification accuracy achieved was substantially higher than the baseline accuracy, thus significantly improving the efficiency of publication labeling.

Rohan Dayal↗

Adversarial Attacks on Deep Neural Network-based Power System Event Classification Models

Online event classification is essential to strengthening the reliability of the power transmission system. Recently, deep learning based methods have achieved great success in numerous domains such as computer vision and natural language processing. Researchers began to adopt deep learning based methods to solve the power system event identification problem and achieved effective results. However, these previous works do not consider that deep learning models are vulnerable to adversarial attacks, potentially influencing real-world applications' reliability. In this paper, we adopt several adversarial attack mechanisms by adding tailored noise signal to the input Phasor Measurement Units (PMU) time series and make the deep learning model misclassify the power system event. This numerical study discloses that current state-of-the-art deep learning based power system event classifiers are extremely vulnerable to adversarial attacks, which may jeopardize the reliability of the power transmission system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

CAN-D: A Modular Four-Step Pipeline for Comprehensively Decoding Controller Area Network Data

Controller area networks (CANs) are a broadcast protocol for real-time communication of critical vehicle subsystems. Original equipment manufacturers of passenger vehicles hold secret their mappings of CAN data to vehicle signals, and these definitions vary according to make, model, and year. Without these mappings, the wealth of real-time vehicle information hidden in the CAN packets is uninterpretable, severely impeding vehicle-related research, including CAN cybersecurity and privacy studies, aftermarket tuning, efficiency and performance monitoring, and fault diagnosis to name a few. Guided by the four-part CAN signal definition, we present CAN-D (CAN-Decoder), a modular, four-step pipeline for identifying each signal's boundaries (start bit and length), endianness (byte ordering), signedness (bit-to-integer encoding), and by leveraging diagnostic standards, augmenting a subset of the extracted signals with meaningful, physical interpretation. En route to CAN-D, we provide a comprehensive review of the CAN signal reverse engineering research. All previous methods ignore endianness and signedness, rendering them incapable of decoding many standard CAN signal definitions. Incorporating endianness grows the search space from 128 to 4.72E21 signal tokenizations and introduces a web of changing dependencies. In response, we formulate, formally analyze, and provide an efficient solution to an optimization problem, allowing identification of the optimal set of signal boundaries and byte orderings. In addition, we provide two novel, state-of-the-art signal boundary classifiers—both of which are superior to previous approaches in precision and recall in three different test scenarios—and the first signedness classification algorithm, which exhibits a $>$ 97% F-score. Altogether, CAN-D is the only solution with the potential to extract any CAN signal that is also the state of the art. In evaluation on 10 vehicles of different makes, CAN-D's average $\ell ^1$ error is five times better (81% less) than all previous methods and exhibits lower average error, even when considering only signals that meet prior methods’ assumptions. Finally, CAN-D is implemented in lightweight hardware, allowing for an on-board diagnostic (OBD-II) plugin for real-time in-vehicle CAN decoding.

42 ENGINEERING↗

Estimating the exceedance probability of rain rate by logistic regression

Recent studies have shown that the fraction of an area with rain intensity above a fixed threshold is highly correlated with the area-averaged rain rate. To estimate the fractional rainy area, a logistic regression model, which estimates the conditional probability that rain rate over an area exceeds a fixed threshold given the values of related covariates, is developed. The problem of dependency in the data in the estimation procedure is bypassed by the method of partial likelihood. Analyses of simulated scanning multichannel microwave radiometer and observed electrically scanning microwave radiometer data during the Global Atlantic Tropical Experiment period show that the use of logistic regression in pixel classification is superior to multiple regression in predicting whether rain rate at each pixel exceeds a given threshold, even in the presence of noisy data. The potential of the logistic regression technique in satellite rain rate estimation is discussed.

Chiu, Long S.↗

A neural net approach to space vehicle guidance

The space vehicle guidance problem is formulated using a neural network approach, and the appropriate neural net architecture for modeling optimum guidance trajectories is investigated. In particular, an investigation is made of the incorporation of prior knowledge about the characteristics of the optimal guidance solution into the neural network architecture. The online classification performance of the developed network is demonstrated using a synthesized network trained with a database of optimum guidance trajectories. Such a neural-network-based guidance approach can readily adapt to environment uncertainties such as those encountered by an AOTV during atmospheric maneuvers.

Caglayan, Alper K.↗

MARLOWE: An Untargeted Proteomics, Statistical Approach to Taxonomic Classification for Forensics

General proteomics research for fundamental science typically addresses laboratory- or patient-derived samples of known origin and composition. However, in a few research areas, such as environmental proteomics, clinical identification of infectious organisms, archeology, art/cultural history, and forensics, attributing the origin of a protein-containing sample to the organisms that produced it is a central focus. A small number of groups have approached this problem and developed software tools for taxonomic characterization and/or identification using bottom-up proteomics. Most such tools identify peptides via database search, and many rely on organism-specific peptides as markers. Our group recently introduced MARLOWE, a software tool for taxonomic characterization of unknown samples based on de novo peptide identification and signal-erosion-resistant strong peptides, which are shared peptides distributed in a taxonomy-dependent manner. In the current work, we further characterize the utility of MARLOWE using publicly available proteomics data from forensically-relevant samples. MARLOWE characterizes samples based on their protein profile, and returns ranked organism lists of potential contributors and taxonomic scores based on shared strong peptides between organisms. Overall, the correct characterization rate ranges between 44 and 100%, depending on the sample type and data acquisition parameters (with lower numbers associated with lower-quality data sets). MARLOWE demonstrates successful characterization of true contributors and close relatives, and provides sufficient specificity to distinguish certain microbial species. MARLOWE demonstrates its ability to provide insight into potential taxonomic sources for a wide range of sample types without prior assumptions about sample contents. As a result, this approach can find utility in forensic science and also broadly in bioanalytical applications that utilize proteomics approaches for taxonomic characterization.

Bacteria↗

Sources of error in thematic classification of remotely sensed imagery

From a statistician's point of view, the input datasets are rarely examined to determine their underlying frequency distribution; it is just assumed that the data are normal enough, and that the deviations from normality are unimportant. It is not clear how deviations from a hypothetical multivariate normal might affect the power of the classification process, and there is ample evidence in the literature that, at a minimum, the spectral channels are correlated. From a practioner's point of view, in a supervised classification the number of training fields for developing a statistical description of a given class is usually arbitrary. It is unclear how small changes in the details of the training field selection process affect the quality of the derived thematic information. The start of an examination of this latter problem is discussed.

Star, Jeffrey L.↗

Parameter trade-offs for imaging spectroscopy systems

With the advent of the EOS era and of configurable sensors, users of these instruments are faced with the twin problems of specifying data acquisition parameters and extracting desired information from the voluminous data. An application of a system model is made to explore system parameter trade-offs for a model sensor based on the High Resolution Imaging Spectrometer. Radiometric performance was studied, along with the effect on classification accuracy of several system parameters. Using a model scene based on typical agricultural reflectance and atmospheric conditions, the atmosphere and sensor are seen to have significant effects on the mean received signal and noise performance. The effect of random uncorrelated errors in the radiometric calibration of the detector array is seen to degrade system performance, especially in the spectral bands below 1 micron. Accurate pixel-to-pixel relative radiometric calibration and the use of the Image Motion Compensation option are seen to improve classification accuracy, especially at high solar zenith angles. Feature sets chosen from characteristics of the scene performed best overall, but ones chosen based on signal-to-noise ratios were seen to be more robust.

Kerekes, John P.↗

Evaluation of global teleconnections in CMIP6 climate projections using complex networks

In climatological research, the evaluation of climate models is one of the central research subjects. As an expression of large-scale dynamical processes, global teleconnections play a major role in interannual to decadal climate variability. Their realistic representation is an indispensable requirement for the simulation of climate change, both natural and anthropogenic. Therefore, the evaluation of global teleconnections is of utmost importance when assessing the physical plausibility of climate projections. We present an application of the graph-theoretical analysis tool δ-MAPS, which constructs complex networks on the basis of spatio-temporal gridded data sets, here sea surface temperature and geopotential height at 500 hPa. Complex networks complement more traditional methods in the analysis of climate variability, like the classification of circulation regimes or empirical orthogonal functions, assuming a new non-linear perspective. While doing so, a number of technical tools and metrics, borrowed from different fields of data science, are implemented into the δ-MAPS framework in order to overcome specific challenges posed by our target problem. Those are trend empirical orthogonal functions (EOFs), distance correlation and distance multicorrelation, and the structural similarity index. δ-MAPS is a two-stage algorithm. In the first place, it assembles grid cells with highly coherent temporal evolution into so-called domains. In a second step, the teleconnections between the domains are inferred by means of the non-linear distance correlation. We construct 2 unipartite and 1 bipartite network for 22 historical CMIP6 climate projections and 2 century-long coupled reanalyses (CERA-20C and 20CRv3). Potential non-stationarity is taken into account by the use of moving time windows. The networks derived from projection data are compared to those from reanalyses. Our results indicate that no single climate projection outperforms all others in every aspect of the evaluation. But there are indeed models which tend to perform better/worse in many aspects. Differences in model performance are generally low within the geopotential height unipartite networks but higher in sea surface temperature and most pronounced in the bipartite network representing the interaction between ocean and atmosphere.

58 GEOSCIENCES↗

Selection of the Australian indicator region

Each Australian state was examined for the availability of LANDSAT data, area, yield, and production characteristics, statistics, crop calendars, and other ancillary data. Agrophysical conditions that could influence labeling and classification accuracies were identified in connection with the highest producing states as determined from available Australian crop statistics. Based primarily on these production statistics, Western Australia and New South Wales were selected as the wheat indicator region for Australia. The general characteristics of wheat in the indicator region, with potential problems anticipated for proportion estimation are considered. The varieties of wheat, the diseases and pests common to New South Wales, and the wheat growing regions of both states are examined.

Reed, C. R.↗

Vortical Flows Research Program of the Fluid Dynamics Research Branch

The research interests of the staff of the Fluid Dynamics Research Branch in the general area of vortex flows are summarized. A major factor in the development of enchanced maneuverability and reduced drag by aerodynamic means is the use of effective vortex control devices. The key to control is the use of emerging computational tools for predicting viscous fluid flow in close coordination with fundamental experiments. In fact, the extremely complex flow fields resulting from numerical solutions to boundary value problems based on the Navier-Stokes equations requires an intimate relationship between computation and experiment. The field of vortex flows is important in so many practical areas that a concerted effort in this area is justified. A brief background of the research activity undertaken is presented, including a proposed classification of the research areas. The classification makes a distinction between issues related to vortex formation and structure, and work on vortex interactions and evolution. Examples of current research results are provided, along with references where available. Based upon the current status of research and planning, speculation on future research directions of the group is also given.

Source record↗

Rule groupings in expert systems using nearest neighbour decision rules, and convex hulls

Expert System shells are lacking in many areas of software engineering. Large rule based systems are not semantically comprehensible, difficult to debug, and impossible to modify or validate. Partitioning a set of rules found in CLIPS (C Language Integrated Production System) into groups of rules which reflect the underlying semantic subdomains of the problem, will address adequately the concerns stated above. Techniques are introduced to structure a CLIPS rule base into groups of rules that inherently have common semantic information. The concepts involved are imported from the field of A.I., Pattern Recognition, and Statistical Inference. Techniques focus on the areas of feature selection, classification, and a criteria of how 'good' the classification technique is, based on Bayesian Decision Theory. A variety of distance metrics are discussed for measuring the 'closeness' of CLIPS rules and various Nearest Neighbor classification algorithms are described based on the above metric.

Anastasiadis, Stergios↗

Spatial and Temporal Dust Source Variability in Northern China Identified Using Advanced Remote Sensing Analysis

The aim of this research is to provide a detailed characterization of spatial patterns and temporal trends in the regional and local dust source areas within the desert of the Alashan Prefecture (Inner Mongolia, China). This problem was approached through multi-scale remote sensing analysis of vegetation changes. The primary requirements for this regional analysis are high spatial and spectral resolution data, accurate spectral calibration and good temporal resolution with a suitable temporal baseline. Landsat analysis and field validation along with the low spatial resolution classifications from MODIS and AVHRR are combined to provide a reliable characterization of the different potential dust-producing sources. The representation of intra-annual and inter-annual Normalized Difference Vegetation Index (NDVI) trend to assess land cover discrimination for mapping potential dust source using MODIS and AVHRR at larger scale is enhanced by Landsat Spectral Mixing Analysis (SMA). The combined methodology is to determine the extent to which Landsat can distinguish important soils types in order to better understand how soil reflectance behaves at seasonal and inter-annual timescales. As a final result mapping soil surface properties using SMA is representative of responses of different land and soil cover previously identified by NDVI trend. The results could be used in dust emission models even if they are not reflecting aggregate formation, soil stability or particle coatings showing to be critical for accurately represent dust source over different regional and local emitting areas.

spectral mixing analysis↗

Sea ice feature and type identification in merged ERS-1 SAR and LANDSAT Thematic Mapper imagery

LANDSAT Thematic Mapper (TM) and ERS-1 SAR (Synthetic Aperture Radar) C band images were acquired for the same area in the Beaufort Sea, 18 Apr. 1992. The two images were co-located to the same grid (25 m resolution) and supervised classification was performed on the TM channel 3 scene in order to classify open water, nilas, grey ice, first year ice, and multiyear ice. Comparison of the LANDSAT classification and the corresponding dB values from the SAR scene showed that, under the given circumstances (high surface winds), open water/nilas/grey ice as defined by a single ice category by the SAR classifier could not be distinguished from first year ice. Surface roughening due to wind appears to be a major problem for the SAR classifier, as the range of the dB values is large enough to overlap all the other ice categories. Additional information, such as surface wind speed is necessary to overcome part of this problem.

Steffen, K.↗

Black Box Testing: Experiments with Runway Incursion Advisory Alerting System

This report summarizes our research findings on the Black box testing of Runway Incursion Advisory Alerting System (RIAAS) and Runway Safety Monitor (RSM) system. Developing automated testing software for such systems has been a problem because of the extensive information that has to be processed. Customized software solutions have been proposed. However, they are time consuming to develop. Here, we present a less expensive, and a more general test platform that is capable of performing complete black box testing. The technique is based on the classification of the anomalies that arise during Monte Carlo simulations. In addition, we also discuss a generalized testing tool (prototype) that we have developed.

Mukkamala, Ravi↗

Anderson Acceleration for Distributed Training of Deep Learning Models

Anderson acceleration (AA) is an extrapolation technique that has recently gained interest in the deep learning (DL) community to speed-up the sequential training of DL models. However, when performed at large scale, the DL training is exposed to a higher risk of getting trapped into steep local minima of the training loss function, and standard AA does not provide sufficient acceleration to escape from these steep local minima. This results in poor generalizability and makes AA ineffective. To restore AA’s advantage to speed-up the training of DL models on large scale computing platforms, we combine AA with an adaptive moving average procedure that boosts the training to escape from steep local minima. By monitoring the relative standard deviation between consecutive iterations, we also introduce a criterion to automatically assess whether the moving average is needed. We applied the method to the following DL instantiations for image classification: (i) ResNet50 trained on the open-source CIFAR100 dataset and (ii) ResNet50 trained on the open-source ImageNet1k dataset. Numerical results obtained using up to 1,536 NVIDIA V100 GPUs on the OLCF supercomputer Summit showed the stabilizing effect of the moving average on AA for all the problems above.

Lupo Pasini, Massimiliano↗

Deep convolutional neural networks for multi-scale time-series classification and application to disruption prediction in fusion devices

The multi-scale, mutli-physics nature of fusion plasmas makes predicting plasma events challenging. Recent advances in deep convolutional neural network architectures (CNN) utilizing dilated convolutions enable accurate predictions on sequences which have long-range, multi-scale characteristics, such as the time-series generated by diagnostic instruments observing fusion plasmas. Here we apply this neural network architecture to the popular problem of disruption prediction in fusion tokamaks, utilizing raw data from a single diagnostic, the Electron Cyclotron Emission imaging (ECEi) diagnostic from the DIII-D tokamak. ECEi measures a fundamental plasma quantity (electron temperature) with high temporal resolution over the entire plasma discharge, making it sensitive to a number of potential pre-disruptions markers with different temporal and spatial scales. Promising, initial disruption prediction results are obtained training a deep CNN with large receptive field ({$\sim$}30k), achieving an $F_1$-score of {$\sim$}91\% on individual time-slices using only the ECEi data.

convolutional neural networks↗