Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A geometrical interpretation of the 2n-th central difference

Many algorithms used for data smoothing, data classification and error detection require the calculation of the distance from a point to the polynomial interpolating its 2n neighbors (n on each side). This computation, if performed naively, would require the solution of a system of equations and could create numerical problems. This note shows that if the data is equally spaced, then this calculation can be performed using a simple recursion formula.

Tapia, R. A.↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Fast decay of classification error in variational quantum circuits

Variational quantum circuits (VQCs) have shown great potential in near-term applications. However, the discriminative power of a VQC, in connection to its circuit architecture and depth, is not understood. To unleash the genuine discriminative power of a VQC, we propose a VQC system with the optimal classical post-processing—maximum-likelihood estimation on measuring all VQC output qubits. Via extensive numerical simulations, we find that the error of VQC quantum data classification typically decays exponentially with the circuit depth, when the VQC architecture is extensive—the number of gates does not shrink with the circuit depth. This fast error suppression ends at the saturation towards the ultimate Helstrom limit of quantum state discrimination. On the other hand, non-extensive VQCs such as quantum convolutional neural networks are sub-optimal and fail to achieve the Helstrom limit, demonstrating a trade-off between ansatz complexity and classification performance in general. To achieve the best performance for a given VQC, the optimal classical post-processing is crucial even for a binary classification problem. To simplify VQCs for near-term implementations, we find that utilizing the symmetry of the input properly can improve the performance, while oversimplification can lead to degradation.

Zhang, Bingzhi↗

Processing LiDAR Data to Predict Natural Hazards

ELF-Base and ELF-Hazards (wherein 'ELF' signifies 'Extract LiDAR Features' and 'LiDAR' signifies 'light detection and ranging') are developmental software modules for processing remote-sensing LiDAR data to identify past natural hazards (principally, landslides) and predict future ones. ELF-Base processes raw LiDAR data, including LiDAR intensity data that are often ignored in other software, to create digital terrain models (DTMs) and digital feature models (DFMs) with sub-meter accuracy. ELF-Hazards fuses raw LiDAR data, data from multispectral and hyperspectral optical images, and DTMs and DFMs generated by ELF-Base to generate hazard risk maps. Advanced algorithms in these software modules include line-enhancement and edge-detection algorithms, surface-characterization algorithms, and algorithms that implement innovative data-fusion techniques. The line-extraction and edge-detection algorithms enable users to locate such features as faults and landslide headwall scarps. Also implemented in this software are improved methodologies for identification and mapping of past landslide events by use of (1) accurate, ELF-derived surface characterizations and (2) three LiDAR/optical-data-fusion techniques: post-classification data fusion, maximum-likelihood estimation modeling, and hierarchical within-class discrimination. This software is expected to enable faster, more accurate forecasting of natural hazards than has previously been possible.

Fairweather, Ian↗

Satellite inventory of Minnesota forest resources

The methods and results of using Landsat Thematic Mapper (TM) data to classify and estimate the acreage of forest covertypes in northeastern Minnesota are described. Portions of six TM scenes covering five counties with a total area of 14,679 square miles were classified into six forest and five nonforest classes. The approach involved the integration of cluster sampling, image processing, and estimation. Using cluster sampling, 343 plots, each 88 acres in size, were photo interpreted and field mapped as a source of reference data for classifier training and calibration of the TM data classifications. Classification accuracies of up to 75 percent were achieved; most misclassification was between similar or related classes. An inverse method of calibration, based on the error rates obtained from the classifications of the cluster plots, was used to adjust the classification class proportions for classification errors. The resulting area estimates for total forest land in the five-county area were within 3 percent of the estimate made independently by the USDA Forest Service. Area estimates for conifer and hardwood forest types were within 0.8 and 6.0 percent respectively, of the Forest Service estimates. A trial of a second method of estimating the same classes as the Forest Service resulted in standard errors of 0.002 to 0.015. A study of the use of multidate TM data for change detection showed that forest canopy depletion, canopy increment, and no change could be identified with greater than 90 percent accuracy. The project results have been the basis for the Minnesota Department of Natural Resources and the Forest Service to define and begin to implement an annual system of forest inventory which utilizes Landsat TM data to detect changes in forest cover.

Bauer, Marvin E.↗

Hybrid NN/SVM Computational System for Optimizing Designs

A computational method and system based on a hybrid of an artificial neural network (NN) and a support vector machine (SVM) (see figure) has been conceived as a means of maximizing or minimizing an objective function, optionally subject to one or more constraints. Such maximization or minimization could be performed, for example, to optimize solve a data-regression or data-classification problem or to optimize a design associated with a response function. A response function can be considered as a subset of a response surface, which is a surface in a vector space of design and performance parameters. A typical example of a design problem that the method and system can be used to solve is that of an airfoil, for which a response function could be the spatial distribution of pressure over the airfoil. In this example, the response surface would describe the pressure distribution as a function of the operating conditions and the geometric parameters of the airfoil. The use of NNs to analyze physical objects in order to optimize their responses under specified physical conditions is well known. NN analysis is suitable for multidimensional interpolation of data that lack structure and enables the representation and optimization of a succession of numerical solutions of increasing complexity or increasing fidelity to the real world. NN analysis is especially useful in helping to satisfy multiple design objectives. Feedforward NNs can be used to make estimates based on nonlinear mathematical models. One difficulty associated with use of a feedforward NN arises from the need for nonlinear optimization to determine connection weights among input, intermediate, and output variables. It can be very expensive to train an NN in cases in which it is necessary to model large amounts of information. Less widely known (in comparison with NNs) are support vector machines (SVMs), which were originally applied in statistical learning theory. In terms that are necessarily oversimplified to fit the scope of this article, an SVM can be characterized as an algorithm that (1) effects a nonlinear mapping of input vectors into a higher-dimensional feature space and (2) involves a dual formulation of governing equations and constraints. One advantageous feature of the SVM approach is that an objective function (which one seeks to minimize to obtain coefficients that define an SVM mathematical model) is convex, so that unlike in the cases of many NN models, any local minimum of an SVM model is also a global minimum.

Rai, Man Mohan↗

A method for classification of multisource data using interval-valued probabilities and its application to HIRIS data

A method of classifying multisource data in remote sensing is presented. The proposed method considers each data source as an information source providing a body of evidence, represents statistical evidence by interval-valued probabilities, and uses Dempster's rule to integrate information based on multiple data source. The method is applied to the problems of ground-cover classification of multispectral data combined with digital terrain data such as elevation, slope, and aspect. Then this method is applied to simulated 201-band High Resolution Imaging Spectrometer (HIRIS) data by dividing the dimensionally huge data source into smaller and more manageable pieces based on the global statistical correlation information. It produces higher classification accuracy than the Maximum Likelihood (ML) classification method when the Hughes phenomenon is apparent.

Kim, H.↗

A method for classification of multisource data using interval-valued probabilities and its application to HIRIS data

A method of classifying multisource data in remote sensing is presented. The proposed method considers each data source as an information source providing a body of evidence, represents statistical evidence by interval-valued probabilities, and uses Dempster's rule to integrate information based on multiple data sources. The method is applied to the problems of ground-cover classification of multispectral data combined with digital terrain data such as elevation, slope, and aspect. Then this method is applied to simulated 201-band High Resolution Imaging Spectrometer (HIRIS) data by dividing the dimensionally huge data source into smaller and more manageable pieces based on the global statistical correlation information. It produces higher classification accuracy than the Maximum Likelihood (ML) classification method when the Hughes phenomenon is apparent.

Kim, H.↗

Petrology, Geochemistry, and Pairing of Lunar Meteorites from the Dominion Range

Introduction: During the 2018-2019 Antarctic Search for Meteorites (ANSMET) field season in the Dominion Range (DOM), 7 lunar meteorite stones were collected: DOM 18242 (15.1 g), DOM 18244 (25.1 g), DOM 18262 (6.8 g), DOM 18509 (16.5 g), DOM 18543 (13.6 g), DOM 18666 (45.9 g), and DOM 18678 (11.6 g). Here we present the initial results of electron microprobe and X-ray computed tomography (XCT) studies of these stones and look at the details of their petrography and mineral chemistry, as well as investigate possible pairing relationships, both with each other and with previously described lunar meteorites. Most of the work presented here is on the DOM 18509, 18543, and 18678 stones; subsamples of the other stones are in hand and similar measurements will be made on them by the time of the meeting. Methods: Textures in DOM 18509, DOM 18543, and DOM 18678 were characterized in 2D by optical microscopy, backscattered electron (BSE) imaging and elemental X-ray images on thin sections, as well as in 3D by X-ray computed tomography (XCT) on sample chips. Mineral compositions were assessed through a combination of wavelength dispersive spectroscopy EPMA (electron probe microanalysis) and x-ray mapping on the JEOL 8530 at NASA JSC. The bulk composition of all three meteorites was determined based on analyses of the fusion crust glass. XCT analyses were done on the Nikon XTH 320 at NASA JSC. ICP-MS data on bulk rock subsamples for each meteorite will be carried out in the near future. Results: The stones are all similar in macroscopic appearance with a dark aphanitic matrix hosting a variety of small- to medium-sized angular mineral and lithic fragments (often light colored in nature) [1,2]. Based on EPMA and XCT results, the three meteorites are polymict regolith breccias comprised of mineral, glass, and lithic clasts ranging up to several mm in length. Melt veins run through all three meteorite samples. Mineral clasts in all 3 stones are dominated by pyroxene and plagioclase (An82-96), with minor amounts of SiO2, olivine (Fo1-52), and FeTi-oxides. Pyroxene grains are mostly Fe-rich pigeonite and augite, and larger clasts are normally zoned and have fine exsolution lamellae. The lithic clasts in all stones consists of: (1) basalt clasts that contain zoned pyroxene, plagioclase laths, and ilmenite, with minor silica and Fe-rich olivine; (2) granulitic clasts; (3) anorthosite clasts; (4) Si-rich clasts that also contain ilmenite, troilite, high-Ca pyroxene, fayalite, and K-feldspar likely mesostasis from late stage basalts). All three meteorites contain glassy fusion crust that is highly vesicular, high in FeO and Al2O3 (15-17 wt% each), ferroan (Mg# of 23-24), and moderately rich in TiO2 (1.4-1.8 wt%). The composition of the fusion crust can serve as a proxy for the bulk meteorite composition and is identical within error for all three meteorites. Implications: The lithic and mineral clasts in all three stones are similar in clast population and assemblages as well as mineral chemistry. In addition, the fusion crust composition, a proxy for bulk composition, is within error of each other for all three stones. Thus DOM 18509, DOM 18543, and DOM 18678 are almost certainly paired. Based on similarities in macroscopic description as well as preliminary classification data [1,2], all 7 lunar stones from DOM are likely paired, though additional quantitative analyses are needed to confirm this. The presence of spherules and vesicular fusion crusts indicates that the DOM pairing group is a regolith breccia. The presence of basaltic and gabbroic clasts as well as more feldspathic materials, suggest that the regolith from which these meteorites formed contained a mixture of feldspathic highland material and mare material, suggesting a possible provenance near a marehighlands boundary. No evidence of KREEPy lithologies have been observed so far in these meteorites, however, future ICP-MS data on bulk rock chips for the stones will reveal any KREEP component if present. The DOM pairing group has many similarities to previously described lunar breccia meteorite MET 01210, however more detailed compositional data will be needed to make a definitive comparison.

R.A. Zeigler↗

Using spatial logic in classification of Landsat TM data

A strategy for spatial/spectral classification of Landsat TM data is presented. The strategy is founded upon 'spatial logic', a logic that seeks to emulate important aspects of visual image interpretation. The carefully structured classification process begins with spectral stratification of the data into water, vegetated and non-vegetated pixels. A region growing algorithm is then used to define 'fields' of similar land cover composition. Fields are characterized by cover composition, size and neighborhood characteristics. A supervised iterative contextual classification algorithm is developed to assign final land use/land cover labels. Maps are generalized using a spatial post-processing technique. Positive, though preliminary, results are presented.

Merchant, J. W.↗

Mechanisms Behind the Long‐Distance Diurnal Offshore Precipitation Propagation in Northwestern South America

Abstract Northwestern South America (NWSA) is the rainiest region on Earth, with diurnal precipitation exhibiting extensive westward offshore propagation of up to about 1,200 km in boreal spring (March‐May). The diurnal offshore precipitation propagation begins slowly (3–10 m s −1 ) near the coast of NWSA (<200 km) but accelerates significantly (∼20 m s −1 ) and shows an afternoon enhancement far from the coast (>400 km). However, the driving mechanisms behind this long‐distance precipitation propagation remain unclear. Using a new cloud tracking and classification data set, we found that mesoscale convective systems (MCSs) are the dominant precipitation contributors in the offshore region of NWSA. Cloud tracking shows that the long‐distance propagation and the afternoon enhancement of diurnal precipitation primarily originate from MCSs initiated in the early morning, either over open oceans or from the coast of Central America. Composite tendency analysis shows that MCSs initiated near the coast of Central America have significant upward cooling and moistening signals starting from the surface before initiation. Further analysis of surface diurnal perturbation fields indicates that the land breeze is the primary driving mechanism for MCS initiation. Conversely, for MCSs initiated over open oceans, a significant downward cooling signal from 400 hPa is observed ∼7 hr before initiation, corresponding to the passage of diurnal gravity waves emitted from the Andes. Additionally, our findings highlight the critical role of lower and mid‐level moisture conditions in MCS initiation, alongside the influence of gravity waves.

Hu, Jingyi [Department of Meteorology and Atmosphe↗

Quantum materials for energy-efficient neuromorphic computing: Opportunities and challenges

Neuromorphic computing approaches become increasingly important as we address future needs for efficiently processing massive amounts of data. The unique attributes of quantum materials can help address these needs by enabling new energy-efficient device concepts that implement neuromorphic ideas at the hardware level. In particular, strong correlations give rise to highly non-linear responses, such as conductive phase transitions that can be harnessed for short- and long-term plasticity. Similarly, magnetization dynamics are strongly non-linear and can be utilized for data classification. This Perspective discusses select examples of these approaches and provides an outlook on the current opportunities and challenges for assembling quantum-material-based devices for neuromorphic functionalities into larger emergent complex network systems.

36 MATERIALS SCIENCE↗

The use of the temporal dimension in classifying and mapping ERTS-1 MSS data

Multispectral data from two ERTS-1 scenes of the same central Pennsylvania area were brought into registration by translation and then merged. The two scenes were viewed on different dates, but from adjacent ground tracks, as frequent cloud cover in Pennsylvania made it impossible to choose two scenes from the same track. Targets selected to be mapped included river water, railroad yards, creeks, urban areas, industrial areas, and vegetation. Equivalent training areas were chosen from each of the original scenes and from the merged data. Classification maps were produced for each, and a comparison was made. Scene brightness was found to have the most important effect on classification differences.

Borden, F. Y.↗

Applications of feature selection

The use of satellite-acquired (LANDSAT) multispectral scanner (MSS) data to conduct an inventory of some crop of economic interest such as wheat over a large geographical area is considered in relation to the development of accurate and efficient algorithms for data classification. The dimension of the measurement space and the computational load for a classification algorithm is increased by the use of multitemporal measurements. Feature selection/combination techniques used to reduce the dimensionality of the problem are described.

Guseman, L. F., Jr.↗