Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

LANL Regular Employee Population and Demographics, FY15-21

This document summarizes an effort to compile and analyze the LANL regular employee population’s demographic attributes and trends between FY15-FY21 using MicroStrategy, a web-based data analytics and visualization tool. LANL’s HR Division collected and prepared data for the Weapons Program’s annual reporting activities for the 2017 through 2023 NNSA Stockpile Stewardship and Management Plan (SSMP). Los Alamos workforce data represent the fiscal year-end snapshot of the permanent employee population who are categorized by the Common Occupational Classification System (COCS). This work enables further analyses, including trends for all workforce attributes presented in the SSMPs as well as a definition of a crosswalk between COCS categorization and the internal laboratory job classification hierarchy. Data for other sites and the federal workforce, as shown in the SSMPs, are included and corresponding figures can be visualized within the MicroStrategy tool. A portable, Excel-based version of the tool was developed and shared with the multi-site workforce working group. This summary highlights selected aspects of the LANL regular workforce. The data analyses and visualizations communicate insights within three major themes regarding permanent LANL employees: 1) demographic transformation, 2) attrition, and 3) skills. Regarding demographic transformation, and consistent with the general trend across the enterprise, early career regular employees are the fastest growing group, with growth ranging between 12-34% year-to-year over the seven-year period analyzed. The data also reflect effects of internal and external factors on attrition; the large spike in separations in FY18 aligns with the contract transition from LANS to Triad and the dip in attrition between FY20-FY21 is likely a result of the global pandemic. The population of regular employees in general management, engineer, and operator COCS categories have the fastest growth rates amongst all the occupational categories. The MicroStrategy and Excel tools provide additional details and trends.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A noninteractive procedure for land-use determination

A complete operational procedure designed for use by the U.S. Army Corps of Engineers for the determination of land-use information for hydrologic planning purposes is reported. The procedure combines photo interpretation techniques and the batch-mode computer analysis of Landsat digital data. Since the operational constraints preclude the use of dedicated, interactive image processing facilities, several novel approaches to geometric correction, classification, and data input/output were developed. The procedure is summarized, and examples of the application of the procedure to urban watersheds are described. In spite of the constraints, the procedure provides results comparable in accuracy and lower in cost than those provided by commercial services using interactive techniques.

Ford, G. E.↗

Computer-implemented land use classification with pattern recognition software and ERTS digital data

Significant progress has been made in the classification of surface conditions (land uses) with computer-implemented techniques based on the use of ERTS digital data and pattern recognition software. The supervised technique presently used at the NASA Earth Resources Laboratory is based on maximum likelihood ratioing with a digital table look-up approach to classification. After classification, colors are assigned to the various surface conditions (land uses) classified, and the color-coded classification is film recorded on either positive or negative 9 1/2 in. film at the scale desired. Prints of the film strips are then mosaicked and photographed to produce a land use map in the format desired. Computer extraction of statistical information is performed to show the extent of each surface condition (land use) within any given land unit that can be identified in the image. Evaluations of the product indicate that classification accuracy is well within the limits for use by land resource managers and administrators. Classifications performed with digital data acquired during different seasons indicate that the combination of two or more classifications offer even better accuracy.

Joyce, A. T.↗

Mapping of taiga forest units using AIRSAR data and/or optical data, and retrieval of forest parameters

A maximum a posteriori Bayesian classifier for multifrequency polarimetric SAR data is used to perform a supervised classification of forest types in the floodplains of Alaska. The image classes include white spruce, balsam poplar, black spruce, alder, non-forests, and open water. The authors investigate the effect on classification accuracy of changing environmental conditions, and of frequency and polarization of the signal. The highest classification accuracy (86 percent correctly classified forest pixels, and 91 percent overall) is obtained combining L- and C-band frequencies fully polarimetric on a date where the forest is just recovering from flooding. The forest map compares favorably with a vegetation map assembled from digitized aerial photos which took five years for completion, and address the state of the forest in 1978, ignoring subsequent fires, changes in the course of the river, clear-cutting of trees, and tree growth. HV-polarization is the most useful polarization at L- and C-band for classification. C-band VV (ERS-1 mode) and L-band HH (J-ERS-1 mode) alone or combined yield unsatisfactory classification accuracies. Additional data acquired in the winter season during thawed and frozen days yield classification accuracies respectively 20 percent and 30 percent lower due to a greater confusion between conifers and deciduous trees. Data acquired at the peak of flooding in May 1991 also yield classification accuracies 10 percent lower because of dominant trunk-ground interactions which mask out finer differences in radar backscatter between tree species. Combination of several of these dates does not improve classification accuracy. For comparison, panchromatic optical data acquired by SPOT in the summer season of 1991 are used to classify the same area. The classification accuracy (78 percent for the forest types and 90 percent if open water is included) is lower than that obtained with AIRSAR although conifers and deciduous trees are better separated due to the presence of leaves on the deciduous trees. Optical data do not separate black spruce and white spruce as well as SAR data, cannot separate alder from balsam poplar, and are of course limited by the frequent cloud cover in the polar regions. Yet, combining SPOT and AIRSAR offers better chances to identify vegetation types independent of ground truth information using a combination of NDVI indexes from SPOT, biomass numbers from AIRSAR, and a segmentation map from either one.

Rignot, Eric↗

Effectiveness of Deep Learning Trained on SynthCity Data for Urban Point-Cloud Classification

3D object recognition is one of the most popular areas of study in computer vision. Many of the more recent algorithms focus on indoor point clouds, classifying 3D geometric objects, and segmenting outdoor 3D scenes. One of the challenges of the classification pipeline is finding adequate and accurate training data. Hence, this article seeks to evaluate the accuracy of a synthetically generated data set called SynthCity, tested on two mobile laser-scan data sets. Varying levels of noise were applied to the training data to reflect varying levels of noise in different scanners. The chosen deep-learning algorithm was Kernel Point Convolution, a convolutional neural network that uses kernel points in Euclidean space for convolution weights.

Geology↗

MicroFisher: Fungal taxonomic classification for metatranscriptomic and metagenomic data using multiple short hypervariable markers

AbstractProfiling the taxonomic and functional composition of microbes using metagenomic (MG) and metatranscriptomic (MT) sequencing is advancing our understanding of microbial functions. However, the sensitivity and accuracy of microbial classification using genome– or core protein-based approaches, especially the classification of eukaryotic organisms, is limited by the availability of genomes and the resolution of sequence databases. To address this, we propose the MicroFisher, a novel approach that applies multiple hypervariable marker genes to profile fungal communities from MGs and MTs. This approach utilizes the hypervariable regions of ITS and large subunit (LSU) rRNA genes for fungal identification with high sensitivity and resolution. Simultaneously, we propose a computational pipeline (MicroFisher) to optimize and integrate the results from classifications using multiple hypervariable markers. To test the performance of our method, we applied MicroFisher to the synthetic community profiling and found high performance in fungal prediction and abundance estimation. In addition, we also used MGs from forest soil and MTs of root eukaryotic microbes to test our method and the results showed that MicroFisher provided more accurate profiling of environmental microbiomes compared to other classification tools. Overall, MicroFisher serves as a novel pipeline for classification of fungal communities from MGs and MTs.

Wang, Haihua↗

A study of the utilization of ERTS-1 data from the Wabash River Basin

The author has identified the following significant results. The most significant results were obtained in the water resources research, urban land use mapping, and soil association mapping projects. ERTS-1 data was used to classify water bodies to determine acreages and high agreement was obtained with USGS figures. Quantitative evaluation was achieved of urban land use classifications from ERTS-1 data and an overall test accuracy of 90.3% was observed. ERTS-1 data classifications of soil test sites were compared with soil association maps scaled to match the computer produced map and good agreement was observed. In some cases the ERTS-1 results proved to be more accurate than the soil association map.

Landgrebe, D. A.↗

A Determination of the Optimum Time of Year for Remotely Classifying Marsh Vegetation from LANDSAT Multispectral Scanner Data

The author has identified the following significant results. A technique was used to determine the optimum time for classifying marsh vegetation from computer-processed LANDSAT MSS data. The technique depended on the analysis of data derived from supervised pattern recognition by maximum likelihood theory. A dispersion index, created by the ratio of separability among the class spectral means to variability within the classes, defined the optimum classification time. Data compared from seven LANDSAT passes acquired over the same area of Louisiana marsh indicated that June and September were optimum marsh mapping times to collectively classify Baccharis halimifolia, Spartina patens, Spartina alterniflora, Juncus roemericanus, and Distichlis spicata. The same technique was used to determine the optimum classification time for individual species. April appeared to be the best month to map Juncus roemericanus; May, Spartina alterniflora; June, Baccharis halimifolia; and September, Spartina patens and Distichlis spicata. This information is important, for instance, when a single species is recognized to indicate a particular environmental condition.

Butera, M. K.↗

Land cover change detection using a GIS-guided, feature-based classification of Landsat thematic mapper data

Landsat TM data were combined with land cover and planimetric data layers contained in the State of Michigan's geographic information system (GIS) to identify changes in forestlands, specifically new oil/gas wells. A GIS-guided feature-based classification method was developed. The regions extracted by the best image band/operator combination were studied using a set of rules based on the characteristics of the GIS oil/gas pads.

Enslin, William R.↗

Remote Sensing Information Classification

This viewgraph presentation reviews the classification of Remote Sensing data in relation to epidemiology. Classification is a way to reduce the dimensionality and precision to something a human can understand. Classification changes SCALAR data into NOMINAL data.

Rickman, Douglas L.↗

Event analysis in an electric power system

According to some embodiments, system and methods are provided including receiving, via a communication interface of an event detection and classification module comprising a processor, data from one or more sensors in a system; determining an event occurred based on the received data; applying a coherency similarity process to the received data via a classification module; determining whether the event is an actual event or a mal-doer event based on an output of the classification module; transmitting the determination of the event as the actual or the mal-doer event; and modifying operation of the system based on the transmitted output. Numerous other aspects are provided.

Hart, Philip Joseph↗

Improving an Acoustic Vehicle Detector Using an Iterative Self-Supervision Procedure

In many non-canonical data science scenarios, obtaining, detecting, attributing, and annotating enough high-quality training data is the primary barrier to developing highly effective models. Moreover, in many problems that are not sufficiently defined or constrained, manually developing a training dataset can often overlook interesting phenomena that should be included. To this end, we have developed and demonstrated an iterative self-supervised learning procedure, whereby models are successfully trained and applied to new data to extract new training examples that are added to the corpus of training data. Successive generations of classifiers are then trained on this augmented corpus. Using low-frequency acoustic data collected by a network of infrasound sensors deployed around the High Flux Isotope Reactor and Radiochemical Engineering Development Center at Oak Ridge National Laboratory, we test the viability of our proposed approach to develop a powerful classifier with the goal of identifying vehicles from continuously streamed data and differentiating these from other sources of noise such as tools, people, airplanes, and wind. Using a small collection of exhaustively manually labeled data, we test several implementation details of the procedure and demonstrate its success regardless of the fidelity of the initial model used to seed the iterative procedure. Finally, we demonstrate the method’s ability to update a model to accommodate changes in the data-generating distribution encountered during long-term persistent data collection.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Proceedings of the Eleventh International Symposium on Remote Sensing of Environment, volume 2

Application and processing of remotely sensed data are discussed. Areas of application include: pollution monitoring, water quality, land use, marine resources, ocean surface properties, and agriculture. Image processing and scene analysis are described along with automated photointerpretation and classification techniques. Data from infrared and multispectral band scanners onboard LANDSAT satellites are emphasized.

Source record↗

Grass Evolutionary Lineages Can Be Identified Using Hyperspectral Leaf Reflectance

Hyperspectral remote sensing has the potential to map numerous attributes of the Earth’s surface, including spatial patterns of biological diversity. Grasslands are one of the largest biomes on Earth. Accurate mapping of grassland biodiversity relies on spectral discrimination of endmembers of species or plant functional types. We focused on spectral separation of grass lineages that dominate global grassy biomes: Andropogoneae (C 4 ), Chloridoideae (C 4 ), and Pooideae (C 3 ). We examined leaf reflectance spectra (350–2,500 nm) from 43 grass species representing these grass lineages from four representative grassland sites in the Great Plains region of North America. Here, we assessed the utility of leaf reflectance data for classification of grass species into three major lineages and by collection site. Classifications had very high accuracy (94%) that were robust to site differences in species and environment. We also show an information loss using multispectral sensors, that is, classification accuracy of grass lineages using spectral bands provided by current multispectral satellites is much lower (accuracy of 85.2% and 61.3% using Sentinel 2 and Landsat 8 bands, respectively). Our results suggest that hyperspectral data have an exciting potential for mapping grass functional types as informed by phylogeny. Leaf-level hyperspectral separability of grass lineages is consistent with the potential increase in biodiversity and functional information content from the next generation of satellite-based spectrometers.

59 BASIC BIOLOGICAL SCIENCES↗

Data reduction through optimized scalar quantization for more compact neural networks

Raw data generation for several existing and planned large physics experiments now exceeds TB/s rates, generating untenable data sets in very little time. Those data often demonstrate high dimensionality while containing limited information. Meanwhile, Machine Learning algorithms are now becoming an essential part of data processing and data analysis. Those algorithms can be used offline for post processing and post data analysis, or they can be used online for real time processing providing ultra low latency experiment monitoring. Both use cases would benefit from data throughput reduction while preserving relevant information: one by reducing the offline storage requirements by several orders of magnitude and the other by allowing ultra fast online inferencing with low complexity Machine Learning models. Moreover, reducing the data source throughput also reduces material cost, power and data management requirements. In this work we demonstrate optimized nonuniform scalar quantization for data source reduction. This data reduction allows lower dimensional representations while preserving the relevant information of the data, thus enabling high accuracy Tiny Machine Learning classifier models for online fast inferences. We demonstrate this approach with an initial proof of concept targeting the CookieBox, an array of electron spectrometers used for angular streaking, that was developed for LCLS-II as an online beam diagnostic tool. We used the Lloyd-Max algorithm with the CookieBox dataset to design an optimized nonuniform scalar quantizer. Optimized quantization lets us reduce input data volume by 69% with no significant impact on inference accuracy. When we tolerate a 2% loss on inference accuracy, we achieved 81% of input data reduction. Finally, the change from a 7-bit to a 3-bit input data quantization reduces our neural network size by 38%.

97 MATHEMATICS AND COMPUTING↗

Data fusion with artificial neural networks (ANN) for classification of earth surface from microwave satellite measurements

A data fusion system with artificial neural networks (ANN) is used for fast and accurate classification of five earth surface conditions and surface changes, based on seven SSMI multichannel microwave satellite measurements. The measurements include brightness temperatures at 19, 22, 37, and 85 GHz at both H and V polarizations (only V at 22 GHz). The seven channel measurements are processed through a convolution computation such that all measurements are located at same grid. Five surface classes including non-scattering surface, precipitation over land, over ocean, snow, and desert are identified from ground-truth observations. The system processes sensory data in three consecutive phases: (1) pre-processing to extract feature vectors and enhance separability among detected classes; (2) preliminary classification of Earth surface patterns using two separate and parallely acting classifiers: back-propagation neural network and binary decision tree classifiers; and (3) data fusion of results from preliminary classifiers to obtain the optimal performance in overall classification. Both the binary decision tree classifier and the fusion processing centers are implemented by neural network architectures. The fusion system configuration is a hierarchical neural network architecture, in which each functional neural net will handle different processing phases in a pipelined fashion. There is a total of around 13,500 samples for this analysis, of which 4 percent are used as the training set and 96 percent as the testing set. After training, this classification system is able to bring up the detection accuracy to 94 percent compared with 88 percent for back-propagation artificial neural networks and 80 percent for binary decision tree classifiers. The neural network data fusion classification is currently under progress to be integrated in an image processing system at NOAA and to be implemented in a prototype of a massively parallel and dynamically reconfigurable Modular Neural Ring (MNR).

Lure, Y. M. Fleming↗

Towards an Automated Classification of Transient Events in Synoptic Sky Surveys

We describe the development of a system for an automated, iterative, real-time classification of transient events discovered in synoptic sky surveys. The system under development incorporates a number of Machine Learning techniques, mostly using Bayesian approaches, due to the sparse nature, heterogeneity, and variable incompleteness of the available data. The classifications are improved iteratively as the new measurements are obtained. One novel featrue is the development of an automated follow-up recommendation engine, that suggest those measruements that would be the most advantageous in terms of resolving classification ambiguities and/or characterization of the astrophysically most interesting objects, given a set of available follow-up assets and their cost funcations. This illustrates the symbiotic relationship of astronomy and applied computer science through the emerging disciplne of AstroInformatics.

classification↗

Analysis of the Tanana River Basin using LANDSAT data

Digital image classification techniques were used to classify land cover/resource information in the Tanana River Basin of Alaska. Portions of four scenes of LANDSAT digital data were analyzed using computer systems at Ames Research Center in an unsupervised approach to derive cluster statistics. The spectral classes were identified using the IDIMS display and color infrared photography. Classification errors were corrected using stratification procedures. The classification scheme resulted in the following eleven categories; sedimented/shallow water, clear/deep water, coniferous forest, mixed forest, deciduous forest, shrub and grass, bog, alpine tundra, barrens, snow and ice, and cultural features. Color coded maps and acreage summaries of the major land cover categories were generated for selected USGS quadrangles (1:250,000) which lie within the drainage basin. The project was completed within six months.

Morrissey, L. A.↗