Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A Robust Schema for Storing and Managing Machine Learning Data and Models

- Machine Learning (ML) has enabled models that can improve efficiency and decrease computational cost - ML models are crucial in enabling Integrated Computational Materials Engineering (ICME) - Large data sets require robust means of storing ML data and models

Brandon L. Hearley↗

Ocean and Earth System Modelling

Petascale supercomputing infrastructure + modelling and analysis capabilities + interdisciplinary upper-ocean expertise Multiscale ocean turbulence simulation Physical-biogeochemical interactions Analysis of large data sets from remote sensing and Earth system model ensembles

Ocean↗

Easy, Scalable Subsetting of GEDI Point Clouds

The GEDI Subsetter, a Python tool developed for NASA’s Multi-mission Algorithm and Analysis Platform (MAAP), optimizes the accessibility and visualization of GEDI point clouds by enabling users to efficiently subset data in a convenient, scalable manner. Complex science data often requires users to learn new software skills and handle many large files. Handling and cleaning large data sets is tedious and error-prone. These challenges significantly impede analysis. One of the goals of NASA's MAAP is to provide a platform that lowers the barrier to conducting research and analysis at scale. When a group of MAAP users wanted to conduct above-ground biomass estimation using GEDI data, we found that their existing workflow for leveraging GEDI data suffered from the barriers mentioned above. Furthermore, their workflow did not scale easily beyond a small number of granules. We found that existing tools related to GEDI data retrieval and subsetting were too limiting, so the GEDI Subsetter was written to support MAAP users’ needs. Being able to run many subsetting jobs simultaneously in the MAAP, and parallelizing the code itself, has led to significant speed improvements in obtaining relevant data, reducing subsetting time from hours to minutes. MAAP users can now more quickly and easily obtain only the data relevant to their research, by choosing which GEDI collection they want to work with (L1A, L2A, L2B, or L4A), and how they want to subset it, by specifying an area of interest, a temporal range, and relevant attributes. This has significantly reduced the feedback loop for users, allowing them to much more quickly subset GEDI data and begin their analysis. Although the GEDI Subsetter originally targeted users of the MAAP, it is generalized such that it can also be used outside of the MAAP and includes a command-line interface for convenience. Furthermore, with minor modifications, it should be possible to use it with non-GEDI data as the general pattern should be applicable to other sparse/track-based sensors.

Charles Daniels↗

Preliminary results from the flight of the Solar Array Module Plasma Interactions Experiment (SAMPIE)

SAMPIE, the Solar Array Module Plasma Interactions Experiment, flew in the Space Shuttle Columbia payload bay as part of the Office of Aeronautics and Space Technology-2 (OAST-2) mission on STS-62, March, 1994. SAMPIE biased samples of solar arrays and space power materials to varying potentials with respect to the surrounding space plasma, and recorded the plasma currents collected and the arcs which occurred, along with a set of plasma diagnostics data. A large set of high quality data was obtained on the behavior of solar arrays and space power materials in the space environment. This paper is the first report on the data SAMPIE telemetered to the ground during the mission. It will be seen that the flight data promise to help determine arcing thresholds, snapover potentials, and floating potentials for arrays and spacecraft in LEO.

Ferguson, Dale C.↗

Preliminary Results from the Flight of the Solar Array Module Plasma Interactions Experiment (SAMPIE)

SAMPIE, the Solar Array Module Plasma Interactions Experiment, flew in the Space Shuttle Columbia payload bay as part of the OAST-2 mission on STS-62, March, 1994. SAMPIE biased samples of solar arrays and space power materials to varying potentials with respect to the surrounding space plasma, and recorded the plasma currents collected and the arcs which occurred, along with a set of plasma diagnostics data. A large set of high quality data was obtained on the behavior of solar arrays and space power materials in the space environment. This paper is the first report on the data SAMPIE telemetered to the ground during the mission. It will be seen that the flight data promise to help determine arcing thresholds, snapover potentials and floating potentials for arrays and spacecraft in LEO.

Ferguson, Dale C.↗

Science Archives in the 21st Century: A NASA LAMBDA Report

Lambda is a thematic data center that focuses on serving the cosmic microwave background (CMB) research community. LAMBDA is an active archive for NASA's Cosmic Background Explorer (COBE) and Wilkinson Microwave Anisotropy Probe (WMAP) mission data sets. In addition, LAMBDA provides analysis software, on-line tools, relevant ancillary data and important web links. LAMBDA also tries to preserve the most important ground-based and suborbital CMB data sets. CMB data is unlike other astrophysical data, consisting of intrinsically diffuse surface brightness photometry with a signal contrast of the order 1 part in 100,000 relative to the uniform background. Because of the extremely faint signal levels, the signal-to-noise ratio is relatively low and detailed instrument-specific knowledge of the data is essential. While the number of data sets being produced is not especially large, those data sets are becoming large and complex. That tendency will increase when the many polarization experiments currently being deployed begin producing data. The LAMBDA experience supports many aspects of the NASA data archive model developed informally over the last ten years-that small focused data centers are often more effective than larger more ambitious collections, for example; that data centers are usually best run by active scientists; that it can be particularly advantageous if those scientists are leaders in the use of the archived data sets; etc. LAMBDA has done some things so well that they might provide lessons for other archives. A lot of effort has been devoted to developing a simple and consistent interface to data sets, for example; and serving all the documentation required via simple 'more' pages and longer explanatory supplements. Many of the problems faced by LAMBDA will also not surprise anyone trying to manage other space science data. These range from persuading mission scientists to provide their data as quickly as possible, to dealing with a high volume of nuisance (spam) messages. Because so many data center problems and solutions are common across individual data centers and disciplines it would be very valuable to establish some new systems of communication - such as informal email lists for administrators and developers. But resources are very limited, so new timeconsuming and inefficient mechanisms - like too-frequent and too-structured meetingsshould be avoided. Although there are great advantages to being small, agile and independent, there are also some areas where science data centers within and without NASA could be better coordinated - for the assignment of persistent identifiers; to encourage the early adoption of useful standards and technologies; etc. Some super-structure to facilitate such coordination might be beneficial as long as it doesn't begin to control the other work of the archives, and become a "methodology police". In this respect the CCSDS "Reference Model for an Open Archive Information System" is a little worrying. It may be that the closer a data center gets to following such a detailed prescription, the less effective it will become. It is much better to have an informal coordination process than a bureaucratic straight-jacket.

Butterworth, P.↗

An interactive environment for the analysis of large Earth observation and model data sets

We propose to develop an interactive environment for the analysis of large Earth science observation and model data sets. We will use a standard scientific data storage format and a large capacity (greater than 20 GB) optical disk system for data management; develop libraries for coordinate transformation and regridding of data sets; modify the NCSA X Image and X Data Slice software for typical Earth observation data sets by including map transformations and missing data handling; develop analysis tools for common mathematical and statistical operations; integrate the components described above into a system for the analysis and comparison of observations and model results; and distribute software and documentation to the scientific community.

Bowman, Kenneth P.↗

An interactive environment for the analysis of large Earth observation and model data sets

We propose to develop an interactive environment for the analysis of large Earth science observation and model data sets. We will use a standard scientific data storage format and a large capacity (greater than 20 GB) optical disk system for data management; develop libraries for coordinate transformation and regridding of data sets; modify the NCSA X Image and X DataSlice software for typical Earth observation data sets by including map transformations and missing data handling; develop analysis tools for common mathematical and statistical operations; integrate the components described above into a system for the analysis and comparison of observations and model results; and distribute software and documentation to the scientific community.

Bowman, Kenneth P.↗

An interactive environment for the analysis of large Earth observation and model data sets

Envision is an interactive environment that provides researchers in the earth sciences convenient ways to manage, browse, and visualize large observed or model data sets. Its main features are support for the netCDF and HDF file formats, an easy to use X/Motif user interface, a client-server configuration, and portability to many UNIX workstations. The Envision package also provides new ways to view and change metadata in a set of data files. It permits a scientist to conveniently and efficiently manage large data sets consisting of many data files. It also provides links to popular visualization tools so that data can be quickly browsed. Envision is a public domain package, freely available to the scientific community. Envision software (binaries and source code) and documentation can be obtained from either of these servers: ftp://vista.atmos.uiuc.edu/pub/envision/ and ftp://csrp.tamu.edu/pub/envision/. Detailed descriptions of Envision capabilities and operations can be found in the User's Guide and Reference Manuals distributed with Envision software.

Bowman, Kenneth P.↗

The Use of a Satellite Climatological Data Set to Infer Large Scale Three Dimensional Flow Characteristics

Ever since the first satellite image loops from the 6.3 micron water vapor channel on the METEOSAT-1 in 1978, there have been numerous efforts (many to a great degree of success) to relate the water vapor radiance patterns to familiar atmospheric dynamic quantities. The realization of these efforts is becoming evident with the merging of satellite derived winds into predictive models (Velden et al., 1997; Swadley and Goerss, 1989). Another parameter that has been quantified from satellite water vapor channel measurements is upper tropospheric relative humidity (UTH) (e.g., Soden and Bretherton, 1996; Schmetz and Turpeinen, 1988). These humidity measurements, in turn, can be used to quantify upper tropospheric water vapor and its transport to more accurately diagnose climate changes (Lerner et al., 1998; Schmetz et al. 1995a) and quantify radiative processes in the upper troposphere. Also apparent in water vapor imagery animations are regions of subsiding and ascending air flow. Indeed, a component of the translated motions we observe are due to vertical velocities. The few attempts at exploiting this information have been met with a fair degree of success. Picon and Desbois (1990) statistically related Meteosat monthly mean water vapor radiances to six standard pressure levels of the European Centre for Medium Range Weather Forecast (ECMWF) model vertical velocities and found correlation coefficients of about 0.50 or less. This paper presents some preliminary results of viewing climatological satellite water vapor data in a different fashion. Specifically, we attempt to infer the three dimensional flow characteristics of the mid- to upper troposphere as portrayed by GOES VAS during the warm ENSO event (1987) and a subsequent cold period in 1998.

Lerner, Jeffrey A.↗

A Comparison of Ffowcs Williams-Hawkings Solvers for Airframe Noise Applications

This paper presents a comparison between two implementations of the Ffowcs Williams and Hawkings equation for airframe noise applications. Airframe systems are generally moving at constant speed and not rotating, so these conditions are used in the current investigation. Efficient and easily implemented forms of the equations applicable to subsonic, rectilinear motion of all acoustic sources are used. The assumptions allow the derivation of a simple form of the equations in the frequency-domain, and the time-domain method uses the restrictions on the motion to reduce the work required to find the emission time. The comparison between the frequency domain method and the retarded time formulation reveals some of the advantages of the different approaches. Both methods are still capable of predicting the far-field noise from nonlinear near-field flow quantities. Because of the large input data sets and potentially large numbers of observer positions of interest in three-dimensional problems, both codes utilize the message passing interface to divide the problem among different processors. Example problems are used to demonstrate the usefulness and efficiency of the two schemes.

Lockard, David P.↗

Accelerating Large Data Analysis By Exploiting Regularities

We present techniques for discovering and exploiting regularity in large curvilinear data sets. The data can be based on a single mesh or a mesh composed of multiple submeshes (also known as zones). Multi-zone data are typical to Computational Fluid Dynamics (CFD) simulations. Regularities include axis-aligned rectilinear and cylindrical meshes as well as cases where one zone is equivalent to a rigid-body transformation of another. Our algorithms can also discover rigid-body motion of meshes in time-series data. Next, we describe a data model where we can utilize the results from the discovery process in order to accelerate large data visualizations. Where possible, we replace general curvilinear zones with rectilinear or cylindrical zones. In rigid-body motion cases we replace a time-series of meshes with a transformed mesh object where a reference mesh is dynamically transformed based on a given time value in order to satisfy geometry requests, on demand. The data model enables us to make these substitutions and dynamic transformations transparently with respect to the visualization algorithms. We present results with large data sets where we combine our mesh replacement and transformation techniques with out-of-core paging in order to achieve significant speed-ups in analysis.

Moran, Patrick J.↗

Tools for Analysis and Visualization of Large Time-Varying CFD Data Sets

In the second year, we continued to built upon and improve our scanline-based direct volume renderer that we developed in the first year of this grant. This extremely general rendering approach can handle regular or irregular grids, including overlapping multiple grids, and polygon mesh surfaces. It runs in parallel on multi-processors. It can also be used in conjunction with a k-d tree hierarchy, where approximate models and error terms are stored in the nodes of the tree, and approximate fast renderings can be created. We have extended our software to handle time-varying data where the data changes but the grid does not. We are now working on extending it to handle more general time-varying data. We have also developed a new extension of our direct volume renderer that uses automatic decimation of the 3D grid, as opposed to an explicit hierarchy. We explored this alternative approach as being more appropriate for very large data sets, where the extra expense of a tree may be unacceptable. We also describe a new approach to direct volume rendering using hardware 3D textures and incorporates lighting effects. Volume rendering using hardware 3D textures is extremely fast, and machines capable of using this technique are becoming more moderately priced. While this technique, at present, is limited to use with regular grids, we are pursuing possible algorithms extending the approach to more general grid types. We have also begun to explore a new method for determining the accuracy of approximate models based on the light field method described at ACM SIGGRAPH '96. In our initial implementation, we automatically image the volume from 32 equi-distant positions on the surface of an enclosing tessellated sphere. We then calculate differences between these images under different conditions of volume approximation or decimation. We are studying whether this will give a quantitative measure of the effects of approximation. We have created new tools for exploring the differences between images produced by various rendering methods. Images created by our software can be stored in the SGI RGB format. Our idtools software reads in pair of images and compares them using various metrics. The differences of the images using the RGB, HSV, and HSL color models can be calculated and shown. We can also calculate the auto-correlation function and the Fourier transform of the image and image differences. We will explore how these image differences compare in order to find useful metrics for quantifying the success of various visualization approaches. In general, progress was consistent with our research plan for the second year of the grant.

Wilhelms, Jane↗

"Tools For Analysis and Visualization of Large Time- Varying CFD Data Sets"

During the four years of this grant (including the one year extension), we have explored many aspects of the visualization of large CFD (Computational Fluid Dynamics) datasets. These have included new direct volume rendering approaches, hierarchical methods, volume decimation, error metrics, parallelization, hardware texture mapping, and methods for analyzing and comparing images. First, we implemented an extremely general direct volume rendering approach that can be used to render rectilinear, curvilinear, or tetrahedral grids, including overlapping multiple zone grids, and time-varying grids. Next, we developed techniques for associating the sample data with a k-d tree, a simple hierarchial data model to approximate samples in the regions covered by each node of the tree, and an error metric for the accuracy of the model. We also explored a new method for determining the accuracy of approximate models based on the light field method described at ACM SIGGRAPH (Association for Computing Machinery Special Interest Group on Computer Graphics) '96. In our initial implementation, we automatically image the volume from 32 approximately evenly distributed positions on the surface of an enclosing tessellated sphere. We then calculate differences between these images under different conditions of volume approximation or decimation.

Wilhelms, Jane↗

Application-Controlled Demand Paging for Out-of-Core Visualization

In the area of scientific visualization, input data sets are often very large. In visualization of Computational Fluid Dynamics (CFD) in particular, input data sets today can surpass 100 Gbytes, and are expected to scale with the ability of supercomputers to generate them. Some visualization tools already partition large data sets into segments, and load appropriate segments as they are needed. However, this does not remove the problem for two reasons: 1) there are data sets for which even the individual segments are too large for the largest graphics workstations, 2) many practitioners do not have access to workstations with the memory capacity required to load even a segment, especially since the state-of-the-art visualization tools tend to be developed by researchers with much more powerful machines. When the size of the data that must be accessed is larger than the size of memory, some form of virtual memory is simply required. This may be by segmentation, paging, or by paged segments. In this paper we demonstrate that complete reliance on operating system virtual memory for out-of-core visualization leads to poor performance. We then describe a paged segment system that we have implemented, and explore the principles of memory management that can be employed by the application for out-of-core visualization. We show that application control over some of these can significantly improve performance. We show that sparse traversal can be exploited by loading only those data actually required. We show also that application control over data loading can be exploited by 1) loading data from alternative storage format (in particular 3-dimensional data stored in sub-cubes), 2) controlling the page size. Both of these techniques effectively reduce the total memory required by visualization at run-time. We also describe experiments we have done on remote out-of-core visualization (when pages are read by demand from remote disk) whose results are promising.

Cox, Michael↗