Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Validating a large geophysical data set - Experiences with satellite-derived cloud parameters

The goal of this study is to validate the global cloud parameters derived from the satellite-borne HIRS2 and MSU atmospheric sounding instrument measurements, and to use the analysis of these data as one prototype for studying large geophysical data sets in general. The HIRS2/MSU data set contains a total of 40 physical parameters, filling 25 MB/day; raw HIRS2/MSU data are available for a period exceeding 10 years. Validation involves developing a quantitative sense for the physical meaning of the derived parameters over the range of environmental conditions sampled. This is accomplished by comparing the spatial and temporal distributions of the derived quantities with similar measurements made using other techniques, and with model results. The need to work with Level 2 (point) data, rather than Level 3 (gridded) data for validation purposes is discussed, and some techniques developed for charting the assumptions made in deriving an algorithm and generating a code to produce geophysical quantities from measured radiances are presented.

Kahn, Ralph↗

Envision: An interactive system for the management and visualization of large geophysical data sets

Envision is a software project at the University of Illinois and Texas A&M, funded by NASA's Applied Information Systems Research Project. It provides researchers in the geophysical sciences convenient ways to manage, browse, and visualize large observed or model data sets. Envision integrates data management, analysis, and visualization of geophysical data in an interactive environment. It employs commonly used standards in data formats, operating systems, networking, and graphics. It also attempts, wherever possible, to integrate with existing scientific visualization and analysis software. Envision has an easy-to-use graphical interface, distributed process components, and an extensible design. It is a public domain package, freely available to the scientific community.

Searight, K. R.↗

New Horizons Successful Completes the Historic First Flyby of Pluto and Its Moons

On July 14, 2015, after a 9.5 year trek across the solar system, NASA's New Horizons spacecraft flew by the dwarf planet Pluto and its system of moons, taking imagery, spectra and in-situ particle data. Data from New Horizons will address numerous outstanding questions on the geology and composition of Pluto and Charon, plus measurements of Pluto's atmosphere, and provide revised understanding of the formation and evolution of Pluto and Charon and its smaller moons. This data set is an invaluable glimpse into the outer Third Zone of the solar system. Data from the intense July 14th fly-by sequence will be downlinked to Earth over a period of 16 months, the duration set by the large data set (over 60 GBits) and the limited transmitted bandwidth rates (approx. 1-2 kbps) and sharing the three 70 m DSN assets with our missions. The small fraction (approx. 1%) of data downlinked during the early phase of the flyby has already revealed Pluto and Charon to be very different worlds, with increasing and dynamic complexity.

Pluto↗

The Rise of Neural Networks for Materials and Chemical Dynamics

Machine learning (ML) is quickly becoming a premier tool for modeling chemical processes and materials. ML-based force fields, trained on large data sets of high-quality electron structure calculations, are particularly attractive due their unique combination of computational efficiency and physical accuracy. This Perspective summarizes some recent advances in the development of neural network-based interatomic potentials. Designing high-quality training data sets is crucial to overall model accuracy. One strategy is active learning, in which new data are automatically collected for atomic configurations that produce large ML uncertainties. Another strategy is to use the highest levels of quantum theory possible. Transfer learning allows training to a data set of mixed fidelity. A model initially trained to a large data set of density functional theory calculations can be significantly improved by retraining to a relatively small data set of expensive coupled cluster theory calculations. These advances are exemplified by applications to molecules and materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Crack detection in fuel cell electrodes using a spatial filtering technique for overcoming noisy backgrounds

Image processing is a powerful tool that allows for rapid and automated data parsing in settings that occupy large variable spaces and require large data sets. Feature detection on difficultly discerned backgrounds is a subset of image processing that facilitates the extraction of quantitative metrics from otherwise subjective data. Crack detection and quantification is an important capability in polymer electrolyte membrane fuel cell quality control, failure analysis, and optimization. This work presents a technique to perform crack detection and quantification which overcomes challenges faced by commonly used image segmentation techniques. We demonstrate the use of a geometrically filtered noise‐level detection technique to select a binary threshold value from which we then quantify how cracked a sample is. Furthermore, we demonstrate the accuracy of our technique using programmatically generated test images of known crack amounts and their performance on real‐world fuel cell catalyst layer samples.

30 DIRECT ENERGY CONVERSION↗

Self-Adjusting Hash Tables for Embedded Flight Applications

A common practice in computer science to associate a value with a key is to use a class of algorithms called a hash-table. These algorithms enable rapid storage and retrieval of values based upon a key. This approach assumes that many keys will need to be stored immediately. A new set of hash-table algorithms optimally uses system resources to ideally represent keys and values in memory such that the information can be stored and retrieved with a minimal amount of time and space. These hash-tables support the efficient addition of new entries. Also, for large data sets, the look-up time for large data-set searches is independent of the number of items stored, i.e., O(1), provided that the chance of collision is low.

James, Mark↗

Visualization tools for the processing of airglow data from RAIDS

In anticipation of large data sets associated with a number of atmospheric imaging instruments being prepared for long term global coverage, NRL is developing graphical interfaces for all aspects of the program. For the first of these projects, RAIDS (the Remote Atmospheric and Ionospheric Detection System), a graphical approach to data handling, visualization, and analysis is envisioned and will set the stage for the satellites that follow. An overall system of hardware and a set of software 'tools,' that will allow for both the routine handling of all data and the analysis of large data sets assembled by scientists and instrument engineers, are currently being developed. The software for standard processing and visualization of instrument data is independent of computer platform and will allow for easy adaptation from one experiment to another. The processing will produce data sets that have similar characteristics, allowing for easy comparison of data obtained under similar circumstances. The visualization of both the engineering and scientific data is an important part of the system. By creating graphical environments for engineering evaluations and for scientific analysis data sets can be viewed and analyzed rapidly. This rapid analysis of data will contribute towards a greater portion of the RAIDS data being utilized.

Miller, Gordon J.↗

Efficient Implementation of an Optimal Interpolator for Large Spatial Data Sets

Interpolating scattered data points is a problem of wide ranging interest. A number of approaches for interpolation have been proposed both from theoretical domains such as computational geometry and in applications' fields such as geostatistics. Our motivation arises from geological and mining applications. In many instances data can be costly to compute and are available only at nonuniformly scattered positions. Because of the high cost of collecting measurements, high accuracy is required in the interpolants. One of the most popular interpolation methods in this field is called ordinary kriging. It is popular because it is a best linear unbiased estimator. The price for its statistical optimality is that the estimator is computationally very expensive. This is because the value of each interpolant is given by the solution of a large dense linear system. In practice, kriging problems have been solved approximately by restricting the domain to a small local neighborhood of points that lie near the query point. Determining the proper size for this neighborhood is a solved by ad hoc methods, and it has been shown that this approach leads to undesirable discontinuities in the interpolant. Recently a more principled approach to approximating kriging has been proposed based on a technique called covariance tapering. This process achieves its efficiency by replacing the large dense kriging system with a much sparser linear system. This technique has been applied to a restriction of our problem, called simple kriging, which is not unbiased for general data sets. In this paper we generalize these results by showing how to apply covariance tapering to the more general problem of ordinary kriging. Through experimentation we demonstrate the space and time efficiency and accuracy of approximating ordinary kriging through the use of covariance tapering combined with iterative methods for solving large sparse systems. We demonstrate our approach on large data sizes arising both from synthetic sources and from real applications.

Memarsadeghi, Nargess↗

Quantifying Atomically Dispersed Catalysts Using Deep Learning Assisted Microscopy

The catalytic performance of atomically dispersed catalysts (ADCs) is greatly influenced by their atomic configurations, such as atom–atom distances, clustering of atoms into dimers and trimers, and their distributions. Scanning transmission electron microscopy (STEM) is a powerful technique for imaging ADCs at the atomic scale; however, most STEM analyses of ADCs thus far have relied on human labeling, making it difficult to analyze large data sets. Here, we introduce a convolutional neural network (CNN)-based algorithm capable of quantifying the spatial arrangement of different adatom configurations. The algorithm was tested on different ADCs with varying support crystallinity and homogeneity. Results show that our algorithm can accurately identify atom positions and effectively analyze large data sets. Here, this work provides a robust method to overcome a major bottleneck in STEM analysis for ADC catalyst research. We highlight the potential of this method to serve as an on-the-fly analysis tool for catalysts in future in situ microscopy experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Fast Implementation of the ISOCLUS Algorithm

Unsupervised clustering is a fundamental building block in numerous image processing applications. One of the most popular and widely used clustering schemes for remote sensing applications is the ISOCLUS algorithm, which is based on the ISODATA method. The algorithm is given a set of n data points in d-dimensional space, an integer k indicating the initial number of clusters, and a number of additional parameters. The general goal is to compute the coordinates of a set of cluster centers in d-space, such that those centers minimize the mean squared distance from each data point to its nearest center. This clustering algorithm is similar to another well-known clustering method, called k-means. One significant feature of ISOCLUS over k-means is that the actual number of clusters reported might be fewer or more than the number supplied as part of the input. The algorithm uses different heuristics to determine whether to merge lor split clusters. As ISOCLUS can run very slowly, particularly on large data sets, there has been a growing .interest in the remote sensing community in computing it efficiently. We have developed a faster implementation of the ISOCLUS algorithm. Our improvement is based on a recent acceleration to the k-means algorithm of Kanungo, et al. They showed that, by using a kd-tree data structure for storing the data, it is possible to reduce the running time of k-means. We have adapted this method for the ISOCLUS algorithm, and we show that it is possible to achieve essentially the same results as ISOCLUS on large data sets, but with significantly lower running times. This adaptation involves computing a number of cluster statistics that are needed for ISOCLUS but not for k-means. Both the k-means and ISOCLUS algorithms are based on iterative schemes, in which nearest neighbors are calculated until some convergence criterion is satisfied. Each iteration requires that the nearest center for each data point be computed. Naively, this requires O(kn) time, where k denotes the current number of centers. Traditional techniques for accelerating nearest neighbor searching involve storing the k centers in a data structure. However, because of the iterative nature of the algorithm, this data structure would need to be rebuilt with each new iteration. Our approach is to store the data points in a kd-tree data structure. The assignment of points to nearest neighbors is carried out by a filtering process, which successively eliminates centers that can not possibly be the nearest neighbor for a given region of space. This algorithm is significantly faster, because large groups of data points can be assigned to their nearest center in a single operation. Preliminary results on a number of real Landsat datasets show that our revised ISOCLUS-like scheme runs about twice as fast.

Memarsadeghi, Nargess↗

Accessing and Visualizing scientific spatiotemporal data

This paper discusses work done by JPL 's Parallel Applications Technologies Group in helping scientists access and visualize very large data sets through the use of multiple computing resources, such as parallel supercomputers, clusters, and grids These tools do one or more of the following tasks visualize local data sets for local users, visualize local data sets for remote users, and access and visualize remote data sets The tools are used for various types of data, including remotely sensed image data, digital elevation models, astronomical surveys, etc The paper attempts to pull some common elements out of these tools that may be useful for others who have to work with similarly large data sets.

data sets↗

Benchmarking Memory Performance with the Data Cube Operator

Data movement across a computer memory hierarchy and across computational grids is known to be a limiting factor for applications processing large data sets. We use the Data Cube Operator on an Arithmetic Data Set, called ADC, to benchmark capabilities of computers and of computational grids to handle large distributed data sets. We present a prototype implementation of a parallel algorithm for computation of the operatol: The algorithm follows a known approach for computing views from the smallest parent. The ADC stresses all levels of grid memory and storage by producing some of 2d views of an Arithmetic Data Set of d-tuples described by a small number of integers. We control data intensity of the ADC by selecting the tuple parameters, the sizes of the views, and the number of realized views. Benchmarking results of memory performance of a number of computer architectures and of a small computational grid are presented.

Frumkin, Michael A.↗

Radiative properties of quantum emitters in boron nitride from excited state calculations and Bayesian analysis

Abstract Point defects in hexagonal boron nitride (hBN) have attracted growing attention as bright single-photon emitters. However, understanding of their atomic structure and radiative properties remains incomplete. Here we study the excited states and radiative lifetimes of over 20 native defects and carbon or oxygen impurities in hBN using ab initio density functional theory and GW plus Bethe-Salpeter equation calculations, generating a large data set of their emission energy, polarization and lifetime. We find a wide variability across quantum emitters, with exciton energies ranging from 0.3 to 4 eV and radiative lifetimes from ns to ms for different defect structures. Through a Bayesian statistical analysis, we identify various high-likelihood charge-neutral defect emitters, among which the native V N N B defect is predicted to possess emission energy and radiative lifetime in agreement with experiments. Our work advances the microscopic understanding of hBN single-photon emitters and introduces a computational framework to characterize and identify quantum emitters in 2D materials.

Chemistry↗

The System for Classification of Low-Pressure Systems (SyCLoPS): An All-In-One Objective Framework for Large-Scale Data Sets

We propose the first unified objective framework (SyCLoPS) for detecting and classifying all types of low-pressure systems (LPSs) in a given data set. We use the state-of-the-art automated feature tracking software TempestExtremes (TE) to detect and track LPS features globally in ERA5 and compute 16 parameters from commonly found atmospheric variables for classification. A Python classifier is implemented to classify all LPSs at once. The framework assigns 16 different labels (classes) to each LPS data point and designates four different types of high-impact LPS tracks, including tracks of tropical cyclone (TC), monsoonal system, subtropical storm and polar low. The classification process involves disentangling high-altitude and drier LPSs, differentiating tropical and non-tropical LPSs using novel criteria, and optimizing for the detection of the four types of high-impact LPS. A comparison of our labels with those in the International Best Track Archive for Climate Stewardship (IBTrACS) revealed an overall accuracy of 95% in distinguishing between tropical systems, extratropical cyclones, and disturbances. SyCLoPS produces a better TC detection skill compared to the previous algorithms, highlighted by an approximately 6% reduction in the false alarm rate compared to the previous TE algorithm. The vertical cross section composite of the four types of high-impact LPS we detect each shows distinct structural characteristics. Finally, we demonstrate that SyCLoPS is valuable for investigating various aspects of LPSs in climate data, such as the evolution of a single LPS track, patterns of LPS frequencies, and precipitation or wind influence associated with a particular LPS class.

54 ENVIRONMENTAL SCIENCES↗

CoreCruncher : Fast and Robust Construction of Core Genomes in Large Prokaryotic Data Sets

The core genome represents the set of genes shared by all, or nearly all, strains of a given population or species of prokaryotes. Inferring the core genome is integral to many genomic analyses, however, most methods rely on the comparison of all the pairs of genomes; a step that is becoming increasingly difficult given the massive accumulation of genomic data. Here, we present CoreCruncher; a program that robustly and rapidly constructs core genomes across hundreds or thousands of genomes. CoreCruncher does not compute all pairwise genome comparisons and uses a heuristic based on the distributions of identity scores to classify sequences as orthologs or paralogs/xenologs. Although it is much faster than current methods, our results indicate that our approach is more conservative than other tools and less sensitive to the presence of paralogs and xenologs. CoreCruncher is freely available from: https://github.com/lbobay/CoreCruncher. CoreCruncher is written in Python 3.7 and can also run on Python 2.7 without modification. It requires the python library Numpy and either Usearch or Blast. Certain options require the programs muscle or mafft.

59 BASIC BIOLOGICAL SCIENCES↗

Distributed Queries of Large Numerical Data Sets

We have extended a previously developed high-level data model, which combines numerical quantities and meta-data into a unified hybrid model, to distributed data. An elegant query language based on SQL is extended further to allow queries against such a distributed hybrid data base. The extension is realized by allowing statements in a non-SQL programming language to be embedded in SQL view definitions.

Nemes, Richard M.↗