Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

traffic flow management↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

weather↗

Application of Bayesian Classification to Content-Based Data Management

The high volume of Earth Observing System data has proven to be challenging to manage for data centers and users alike. At the Goddard Earth Sciences Distributed Active Archive Center (GES DAAC), about 1 TB of new data are archived each day. Distribution to users is also about 1 TB/day. A substantial portion of this distribution is MODIS calibrated radiance data, which has a wide variety of uses. However, much of the data is not useful for a particular user's needs: for example, ocean color users typically need oceanic pixels that are free of cloud and sun-glint. The GES DAAC is using a simple Bayesian classification scheme to rapidly classify each pixel in the scene in order to support several experimental content-based data services for near-real-time MODIS calibrated radiance products (from Direct Readout stations). Content-based subsetting would allow distribution of, say, only clear pixels to the user if desired. Content-based subscriptions would distribute data to users only when they fit the user's usability criteria in their area of interest within the scene. Content-based cache management would retain more useful data on disk for easy online access. The classification may even be exploited in an automated quality assessment of the geolocation product. Though initially to be demonstrated at the GES DAAC, these techniques have applicability in other resource-limited environments, such as spaceborne data systems.

Lynnes, Christopher↗

Classification of orthostatic intolerance through data analytics

Imbalance in the autonomic nervous system can lead to orthostatic intolerance manifested by dizziness, lightheadedness, and a sudden loss of consciousness (syncope); these are common conditions, but they are challenging to diagnose correctly. Uncertainties about the triggering mechanisms and the underlying pathophysiology have led to variations in their classification. This study uses machine learning to categorize patients with orthostatic intolerance. Here we use random forest classification trees to identify a small number of markers in blood pressure, and heart rate time-series data measured during head-up tilt to (a) distinguish patients with a single pathology and (b) examine data from patients with a mixed pathophysiology. Next, we use Kmeans to cluster the markers representing the time-series data. We apply the proposed method analyzing clinical data from 186 subjects identified as control or suffering from one of four conditions: postural orthostatic tachycardia (POTS), cardioinhibition, vasodepression, and mixed cardioinhibition and vasodepression. Classification results confirm the use of supervised machine learning. We were able to categorize more than 95% of patients with a single condition and were able to subgroup all patients with mixed cardioinhibitory and vasodepressor syncope. Clustering results confirm the disease groups and identify two distinct subgroups within the control and mixed groups. The proposed study demonstrates how to use machine learning to discover structure in blood pressure and heart rate time-series data. The methodology is used in classification of patients with orthostatic intolerance. Diagnosing orthostatic intolerance is challenging, and full characterization of the pathophysiological mechanisms remains a topic of ongoing research. This study provides a step toward leveraging machine learning to assist clinicians and researchers in addressing these challenges.

60 APPLIED LIFE SCIENCES↗

Computer-aided analysis of Skylab multispectral scanner data in mountainous terrain for land use, forestry, water resource, and geologic applications

The author has identified the following significant results. One of the most significant results of this Skylab research involved the geometric correction and overlay of the Skylab multispectral scanner data with the LANDSAT multispectral scanner data, and also with a set of topographic data, including elevation, slope, and aspect. The Skylab S192 multispectral scanner data had distinct differences in noise level of the data in the various wavelength bands. Results of the temporal evaluation of the SL-2 and SL-3 photography were found to be particularly important for proper interpretation of the computer-aided analysis of the SL-2 and SL-3 multispectral scanner data. There was a quality problem involving the ringing effect introduced by digital filtering. The modified clustering technique was found valuable when working with multispectral scanner data involving many wavelength bands and covering large geographic areas. Analysis of the SL-2 scanner data involved classification of major cover types and also forest cover types. Comparison of the results obtained wth Skylab MSS data and LANDSAT MSS data indicated that the improved spectral resolution of the Skylab scanner system enabled a higher classification accuracy to be obtained for forest cover types, although the classification performance for major cover types was not significantly different.

Hoffer, R. M.↗

Raw Lidar and Camera Data Synchronized with Precipitation and Present Weather Data

As part of the sensor characterization task of the SMART 2.0 project, this dataset includes raw data from three spinning lidars ([Ouster OS2-128](https://ouster.com/products/scanning-lidar/os2-sensor/), [Velodyne Puck (VLP-16)](https://velodynelidar.com/products/puck/), and [Velodyne Ultra Puck (VLP-32)](https://velodynelidar.com/products/ultra-puck/)), one camera ([Mako G-319](https://www.alliedvision.com/en/camera-selector/detail/mako/g-319/)), and one present weather sensor ([Vaisala FD-70](https://www.vaisala.com/en/products/weather-environmental-sensors/forward-scatter-fd70)). All data were synchronized, with the log start time indicated in the file name (HHMMSS). The data can be filtered by date, log time (HHMMSS), sensor, frame ID, and weather classification. These data were gathered statically at the Argonne Testbed for Multiscale Observational Science (ATMOS). Two target stop signs were placed in view of the sensors to contribute a target for comparing sensor data under different conditions. The weather data for each day are stored in netCDF “.nc” files. The lidar data contain the X, Y, Z, intensity, reflectivity, and ring from Ouster OS2-128 rev6, Velodyne VLP-16, and Velodyne VLP-32 lidars. ![raw lidar image](LiDAR_pointcloud_ATMOS.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Classification of ERTS-1 MSS data by canonical analysis

The objective of canonical analysis is to obtain the maximum separability among a number of catergories. The application of canonical analysis was investigated using the merged MSS ERTS-1 data for one area viewed on two dates. The effect of threshold values on classification regions and confusion regions was investigated.

Lachowski, H. M.↗

A case study in the practical use of LANDSAT data

The use of computer aided classification of LANDSAT data in developing water quality plans for New Jersey watersheds is used to exemplify how a state natural resource management program benefits from satellite imagery. The transition of a research and development system into an operational remote sensing system to help decision makers is demonstrated. Nontechnial issues that can assist (or hinder) an agency in adopting a new technology are examined. The progress of LANDSAT use by state government from the earliest stage of curiosity through to incorporation in actual state planning methods is charted. Potential applications of LANDSAT data to real information needs and solutions to management problems are examined. The problems and mistakes that occurred in using LANDSAT data in the past are discussed as well as the ways by which these problems were overcome.

Cox, S.↗

Classification of simulated and actual NOAA-6 AVHRR data for hydrologic land-surface feature definition

An examination of the possibilities of using Landsat data to simulate NOAA-6 Advanced Very High Resolution Radiometer (AVHRR) data on two channels, as well as using actual NOAA-6 imagery, for large-scale hydrological studies is presented. A running average was obtained of 18 consecutive pixels of 1 km resolution taken by the Landsat scanners were scaled up to 8-bit data and investigated for different gray levels. AVHRR data comprising five channels of 10-bit, band-interleaved information covering 10 deg latitude were analyzed and a suitable pixel grid was chosen for comparison with the Landsat data in a supervised classification format, an unsupervised mode, and with ground truth. Landcover delineation was explored by removing snow, water, and cloud features from the cluster analysis, and resulted in less than 10% difference. Low resolution large-scale data was determined useful for characterizing some landcover features if weekly and/or monthly updates are maintained.

Ormsby, J. P.↗

Remote-sensing applications in the State of Mississippi

A computer derived land use classification scheme for infrared LANDSAT imagery was developed and applied to update existing Mississippi coastline data. Inventory classifications were accomplished by photographic enlargement and photointerpretations showing color coded resources on the ground.

Bankston, P. T.↗

Thematic mapping, land use, geological structure and water resources in central Spain

The author has identified the following significant results. The images can be positioned in an absolute reference system (geographical coordinates or polar stereographic coordinates) by means of their marginal indicators. By digital analysis of LANDSAT data and geometric positioning of pixels in UTM projection, accuracy was achieved for corrected MSS information which could be used for updating maps at scale 1:200,000 or smaller. Results show that adjustment of the UTM grid was better obtained by a first order, or even second order, algorithm of geometric correction. Digital analysis of LANDSAT data from the Madrid area showed that this line of study was promising for automatic classification of data applied to thematic cartography and soils identification.

Delascuevas, N.↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

The existing and forthcoming data bases from NASA missions contain an abundance of information whose complexity cannot be efficiently tapped with simple statistical techniques. Powerful multivariate statistical methods already exist which can be used to harness much of the richness of these data. Automatic classification techniques have been developed to solve the problem of identifying known types of objects in multiparameter data sets, in addition to leading to the discovery of new physical phenomena and classes of objects. We propose an exploratory study and integration of promising techniques in the development of a general and modular classification/analysis system for very large data bases, which would enhance and optimize data management and the use of human research resource.

Djorgovski, George↗

Mapping Species Composition of Forests and Tree Plantations in Northeastern Costa Rica with an Integration of Hyperspectral and Multitemporal Landsat Imagery

An efficient means to map tree plantations is needed to detect tropical land use change and evaluate reforestation projects. To analyze recent tree plantation expansion in northeastern Costa Rica, we examined the potential of combining moderate-resolution hyperspectral imagery (2005 HyMap mosaic) with multitemporal, multispectral data (Landsat) to accurately classify (1) general forest types and (2) tree plantations by species composition. Following a linear discriminant analysis to reduce data dimensionality, we compared four Random Forest classification models: hyperspectral data (HD) alone; HD plus interannual spectral metrics; HD plus a multitemporal forest regrowth classification; and all three models combined. The fourth, combined model achieved overall accuracy of 88.5%. Adding multitemporal data significantly improved classification accuracy (p less than 0.0001) of all forest types, although the effect on tree plantation accuracy was modest. The hyperspectral data alone classified six species of tree plantations with 75% to 93% producer's accuracy; adding multitemporal spectral data increased accuracy only for two species with dense canopies. Non-native tree species had higher classification accuracy overall and made up the majority of tree plantations in this landscape. Our results indicate that combining occasionally acquired hyperspectral data with widely available multitemporal satellite imagery enhances mapping and monitoring of reforestation in tropical landscapes.

hyperspectral fusion↗

Improvements in lake water budget computations using Landsat data

A supervised multispectral classification was performed on Landsat data for Lake Okeechobee's extensive littoral zone to provide two types of information. First, the acreage of a given plant species as measured by satellite was combined with a more accurate transpiration rate to give a better estimate of evapotranspiration from the littoral zone. Second, the surface area coupled by plant communities was used to develop a better estimate of the water surface as a function of lake stage. Based on this information, more detailed representations of evapotranspiration and total water surface (and hence total lake volume) were provided to the water balance budget model for lake volume predictions. The model results based on information derived from satellite demonstrated a 94 percent reduction in cumulative lake stage error and a 70 percent reduction in the maximum deviation of the lake stage.

Gervin, J. C.↗

Analysis of data acquired by synthetic aperture radar over Dade County, Florida, and Acadia Parish, Louisiana

Results of digital processing of airborne X-band synthetic aperture radar (SAR) data acquired over Dade County, Florida, and Acadia Parish, Louisiana are presented. The goal was to investigate the utility of SAR data for land cover mapping and area estimation under the AgRISTARS Domestic Crops and Land Cover Project. In the case of the Acadia Paris study area, LANDSAT multispectral scanner (MSS) data were also used to form a combined SAR and MSS data set. The results of accuracy evaluation for the SAR, MSS, and SAR/MSS data using supervised classification show that the combined SAR/MSS data set results in an improved classification accuracy of the five land cover classes as compared with SAR-only and MSS-only data sets. In the case of the Dade County study area, the results indicate that both HH and VV polarization data are highly responsive to the row orientation of the row crop but not to the specific vegetation which forms the row structure. On the other hand, the HV polarization data are relatively insensitive to the orientation of row crop. Therefore, the HV polarization data may be used to discriminate the specific vegetation that forms the row structure.

Wu, S. T.↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

The existing and forthcoming data bases from NASA missions contain an abundance of information whose complexity cannot be efficiently tapped with simple statistical techniques. Powerful multivariate statistical methods already exist which can be used to harness much of the richness of these data. Automatic classification techniques have been developed to solve the problem of identifying known types of objects in multi parameter data sets, in addition to leading to the discovery of new physical phenomena and classes of objects. We propose an exploratory study and integration of promising techniques in the development of a general and modular classification/analysis system for very large data bases, which would enhance and optimize data management and the use of human research resources.

Djorgovski, Stanislav↗

Smart Optical RAM for Fast Information Management and Analysis

Statement of Problem Instruments for high speed and high capacity in-situ data identification, classification and storage capabilities are needed by NASA for the information management and analysis of extremely large volume of data sets in future space exploration, space habitation and utilization, in addition to the various missions to planet-earth programs. Parameters such as communication delays, limited resources, and inaccessibility of human manipulation require more intelligent, compact, low power, and light weight information management and data storage techniques. New and innovative algorithms and architecture using photonics will enable us to meet these challenges. The technology has applications for other government and public agencies.

Liu, Hua-Kuang↗

Crop classification using airborne radar and Landsat data

NASA 13.3 GHz airborne radar data from a soil moisture measurement analysis is used to investigate the statistical nature of the radar backscattering coefficient for bare ground and three different crop types, and to evaluate the crop classification rates using Landsat data alone or combined with the airborne survey. The scatterometer was a fan-beam Doppler system, VV polarized, and is considered only for 50 deg angles of incidence. A total of 36 fields were covered a week apart by the aircraft and Landsat, and Rayleigh statistics were used in the frequency averaging to eliminate fluctuations due to random fluctuations. Within-field variances were calculated for the Landsat and the radar data and used to design optimum crop classification procedures. The Landsat Band 4 readings were 67% accurate, and an increase in accuracy of 10% was achieved by the addition of the radar data.

Ulaby, F. T.↗