Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

New Horizons Successful Completes the Historic First Flyby of Pluto and Its Moons

On July 14, 2015, after a 9.5 year trek across the solar system, NASA's New Horizons spacecraft flew by the dwarf planet Pluto and its system of moons, taking imagery, spectra and in-situ particle data. Data from New Horizons will address numerous outstanding questions on the geology and composition of Pluto and Charon, plus measurements of Pluto's atmosphere, and provide revised understanding of the formation and evolution of Pluto and Charon and its smaller moons. This data set is an invaluable glimpse into the outer Third Zone of the solar system. Data from the intense July 14th fly-by sequence will be downlinked to Earth over a period of 16 months, the duration set by the large data set (over 60 GBits) and the limited transmitted bandwidth rates (approx. 1-2 kbps) and sharing the three 70 m DSN assets with our missions. The small fraction (approx. 1%) of data downlinked during the early phase of the flyby has already revealed Pluto and Charon to be very different worlds, with increasing and dynamic complexity.

Pluto↗

Self-Adjusting Hash Tables for Embedded Flight Applications

A common practice in computer science to associate a value with a key is to use a class of algorithms called a hash-table. These algorithms enable rapid storage and retrieval of values based upon a key. This approach assumes that many keys will need to be stored immediately. A new set of hash-table algorithms optimally uses system resources to ideally represent keys and values in memory such that the information can be stored and retrieved with a minimal amount of time and space. These hash-tables support the efficient addition of new entries. Also, for large data sets, the look-up time for large data-set searches is independent of the number of items stored, i.e., O(1), provided that the chance of collision is low.

James, Mark↗

Visualization tools for the processing of airglow data from RAIDS

In anticipation of large data sets associated with a number of atmospheric imaging instruments being prepared for long term global coverage, NRL is developing graphical interfaces for all aspects of the program. For the first of these projects, RAIDS (the Remote Atmospheric and Ionospheric Detection System), a graphical approach to data handling, visualization, and analysis is envisioned and will set the stage for the satellites that follow. An overall system of hardware and a set of software 'tools,' that will allow for both the routine handling of all data and the analysis of large data sets assembled by scientists and instrument engineers, are currently being developed. The software for standard processing and visualization of instrument data is independent of computer platform and will allow for easy adaptation from one experiment to another. The processing will produce data sets that have similar characteristics, allowing for easy comparison of data obtained under similar circumstances. The visualization of both the engineering and scientific data is an important part of the system. By creating graphical environments for engineering evaluations and for scientific analysis data sets can be viewed and analyzed rapidly. This rapid analysis of data will contribute towards a greater portion of the RAIDS data being utilized.

Miller, Gordon J.↗

Efficient Implementation of an Optimal Interpolator for Large Spatial Data Sets

Interpolating scattered data points is a problem of wide ranging interest. A number of approaches for interpolation have been proposed both from theoretical domains such as computational geometry and in applications' fields such as geostatistics. Our motivation arises from geological and mining applications. In many instances data can be costly to compute and are available only at nonuniformly scattered positions. Because of the high cost of collecting measurements, high accuracy is required in the interpolants. One of the most popular interpolation methods in this field is called ordinary kriging. It is popular because it is a best linear unbiased estimator. The price for its statistical optimality is that the estimator is computationally very expensive. This is because the value of each interpolant is given by the solution of a large dense linear system. In practice, kriging problems have been solved approximately by restricting the domain to a small local neighborhood of points that lie near the query point. Determining the proper size for this neighborhood is a solved by ad hoc methods, and it has been shown that this approach leads to undesirable discontinuities in the interpolant. Recently a more principled approach to approximating kriging has been proposed based on a technique called covariance tapering. This process achieves its efficiency by replacing the large dense kriging system with a much sparser linear system. This technique has been applied to a restriction of our problem, called simple kriging, which is not unbiased for general data sets. In this paper we generalize these results by showing how to apply covariance tapering to the more general problem of ordinary kriging. Through experimentation we demonstrate the space and time efficiency and accuracy of approximating ordinary kriging through the use of covariance tapering combined with iterative methods for solving large sparse systems. We demonstrate our approach on large data sizes arising both from synthetic sources and from real applications.

Memarsadeghi, Nargess↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.

Memarsadeghi, Nargess↗

Accessing and Visualizing scientific spatiotemporal data

This paper discusses work done by JPL 's Parallel Applications Technologies Group in helping scientists access and visualize very large data sets through the use of multiple computing resources, such as parallel supercomputers, clusters, and grids These tools do one or more of the following tasks visualize local data sets for local users, visualize local data sets for remote users, and access and visualize remote data sets The tools are used for various types of data, including remotely sensed image data, digital elevation models, astronomical surveys, etc The paper attempts to pull some common elements out of these tools that may be useful for others who have to work with similarly large data sets.

data sets↗

Benchmarking Memory Performance with the Data Cube Operator

Data movement across a computer memory hierarchy and across computational grids is known to be a limiting factor for applications processing large data sets. We use the Data Cube Operator on an Arithmetic Data Set, called ADC, to benchmark capabilities of computers and of computational grids to handle large distributed data sets. We present a prototype implementation of a parallel algorithm for computation of the operatol: The algorithm follows a known approach for computing views from the smallest parent. The ADC stresses all levels of grid memory and storage by producing some of 2d views of an Arithmetic Data Set of d-tuples described by a small number of integers. We control data intensity of the ADC by selecting the tuple parameters, the sizes of the views, and the number of realized views. Benchmarking results of memory performance of a number of computer architectures and of a small computational grid are presented.

Frumkin, Michael A.↗

Distributed Queries of Large Numerical Data Sets

We have extended a previously developed high-level data model, which combines numerical quantities and meta-data into a unified hybrid model, to distributed data. An elegant query language based on SQL is extended further to allow queries against such a distributed hybrid data base. The extension is realized by allowing statements in a non-SQL programming language to be embedded in SQL view definitions.

Nemes, Richard M.↗

Integrating Large Scale Data Sets to Develop Predictive Hypotheses of Low-Dose Radiation-Induced Health Effects

Over one hundred years of radiation biology research has revealed much about the DNA damages induced by the deposition of energy from exposure to ionizing radiation and the subsequent cellular responses. However, there are still significant gaps in our understanding of how these might lead to detrimental health effects, particularly at low doses (100 mGy (milligray)). Recent advances in high throughput omics technologies enable interrogation of induced radiation effects at the genomic, proteomic and metabolomic levels. These include changes in gene expression, protein modifications, e.g., phosphorylation, acetylation, and methylation, and metabolic changes. We will discuss the integration of data obtained from multiple omics platforms to understand radiation dose, and dose rate effects in a complex human tissue model as a function of time. We will use as an example our results on the low dose responses in a 3D human skin model.

ionizing radiation↗

Encounter-Based Simulation Architecture for Detect-And-Avoid Modeling

This paper presents an encounter-based simulation architecture developed at NASA to facilitate flexible and efficient Detect and Avoid modeling in parametric or tradespace studies on large data sets. The basic premise of this tool is that large-scale input data can be reduced to a set of `canonical encounters' and that using the reduced data in simulations does not lead to loss of fidelity. A canonical encounter is specified as ownship and intruder flight portions potentially resulting in a loss of well clear along with a set of properties that characterize the encounter. The advantages of using canonical encounters include faster simulations, reduced memory footprint, ability to select encounters based on user-specified criteria, shared encounters across multiple teams, peer-reviewed encounters, and a better understanding of the input data set, to name a few.

MOPS↗

Encounter-Based Simulation Architecture for Detect and Avoid Modeling

This paper presents an encounter-based simulation architecture developed at NASA to facilitate flexible and efficient Detect and Avoid modeling in parametric or tradespace studies on large data sets. The basic premise of this tool is that large-scale input data can be reduced to a set of `canonical encounters' and that using the reduced data in simulations does not lead to loss of fidelity. A canonical encounter is specified as ownship and intruder flight portions potentially resulting in a loss of well clear along with a set of properties that characterize the encounter. The advantages of using canonical encounters include faster simulations, reduced memory footprint, ability to select encounters based on user-specified criteria, shared encounters across multiple teams, peer-reviewed encounters, and a better understanding of the input data set, to name a few.

MOPS↗

The Application of Principal Component Analysis Using Fixed Eigenvectors to the Infrared Thermographic Inspection of the Space Shuttle Thermal Protection System

The Nondestructive Evaluation Sciences Branch at NASA s Langley Research Center has been actively involved in the development of thermographic inspection techniques for more than 15 years. Since the Space Shuttle Columbia accident, NASA has focused on the improvement of advanced NDE techniques for the Reinforced Carbon-Carbon (RCC) panels that comprise the orbiter s wing leading edge. Various nondestructive inspection techniques have been used in the examination of the RCC, but thermography has emerged as an effective inspection alternative to more traditional methods. Thermography is a non-contact inspection method as compared to ultrasonic techniques which typically require the use of a coupling medium between the transducer and material. Like radiographic techniques, thermography can be used to inspect large areas, but has the advantage of minimal safety concerns and the ability for single-sided measurements. Principal Component Analysis (PCA) has been shown effective for reducing thermographic NDE data. A typical implementation of PCA is when the eigenvectors are generated from the data set being analyzed. Although it is a powerful tool for enhancing the visibility of defects in thermal data, PCA can be computationally intense and time consuming when applied to the large data sets typical in thermography. Additionally, PCA can experience problems when very large defects are present (defects that dominate the field-of-view), since the calculation of the eigenvectors is now governed by the presence of the defect, not the good material. To increase the processing speed and to minimize the negative effects of large defects, an alternative method of PCA is being pursued when a fixed set of eigenvectors is used to process the thermal data from the RCC materials. These eigen vectors can be generated either from an analytic model of the thermal response of the material under examination, or from a large cross section of experimental data. This paper will provide the details of the analytic model; an overview of the PCA process; as well as a quantitative signal-to-noise comparison of the results of performing both embodiments of PCA on thermographic data from various RCC specimens. Details of a system that has been developed to allow insitu inspection of a majority of shuttle RCC components will be presented along with the acceptance test results for this system. Additionally, the results of applying this technology to the Space Shuttle Discovery after its return from flight will be presented.

Cramer, K. Elliott↗

Analysis and representation of complex structures in separated flows

We discuss our recent work on extraction and visualization of topological information in separated fluid flow data sets. As with scene analysis, an abstract representation of a large data set can greatly facilitate the understanding of complex, high-level structures. When studying flow topology, such a representation can be produced by locating and characterizing critical points in the velocity field and generating the associated stream surfaces. In 3D flows, the surface topology serves as the starting point. The 2D tangential velocity field near the surface of the body is examined for critical points. The tangential velocity field is integrated out along the principal directions of certain classes of critical points to produce curves depicting the topology of the flow near the body. The points and curves are linked to form a skeleton representing the 2D vector field topology. This skeleton provides a basis for analyzing the 3D structures associated with the flow separation. The points along the separation curves in the skeleton are used to start tangent curve integrations. Integration origins are successively refined to produce stream surfaces. The map of the global topology is completed by generating those stream surfaces associated with 3D critical points.

Helman, James↗

Arithmetic Data Cube as a Data Intensive Benchmark

Data movement across computational grids and across memory hierarchy of individual grid machines is known to be a limiting factor for application involving large data sets. In this paper we introduce the Data Cube Operator on an Arithmetic Data Set which we call Arithmetic Data Cube (ADC). We propose to use the ADC to benchmark grid capabilities to handle large distributed data sets. The ADC stresses all levels of grid memory by producing 2d views of an Arithmetic Data Set of d-tuples described by a small number of parameters. We control data intensity of the ADC by controlling the sizes of the views through choice of the tuple parameters.

Frumkin, Michael A.↗