Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Network issues for large mass storage requirements

File Servers and Supercomputing environments need high performance networks to balance the I/O requirements seen in today's demanding computing scenarios. UltraNet is one solution which permits both high aggregate transfer rates and high task-to-task transfer rates as demonstrated in actual tests. UltraNet provides this capability as both a Server-to-Server and Server-to-Client access network giving the supercomputing center the following advantages highest performance Transport Level connections (to 40 MBytes/sec effective rates); matches the throughput of the emerging high performance disk technologies, such as RAID, parallel head transfer devices and software striping; supports standard network and file system applications using SOCKET's based application program interface such as FTP, rcp, rdump, etc.; supports access to the Network File System (NFS) and LARGE aggregate bandwidth for large NFS usage; provides access to a distributed, hierarchical data server capability using DISCOS UniTree product; supports file server solutions available from multiple vendors, including Cray, Convex, Alliant, FPS, IBM, and others.

Perdue, James↗

Incorporation of a hierarchical grid component structure into GRIDGEN

The underlying framework of the GRIDGEN multiple block grid generation system has been refined so that grid components are now stored within a hierarchical data structure. This restructuring has enhanced the usability of the software by allowing grids to be generated on a more intuitive level. This new framework also provides a means by which the multiple block system can be edited at most any level in the grid generation process. Editing tools are currently being added to GRIDGEN so that a change to the grid can be propagated backward and forward in the data hierarchy. The new data structure, the editing tools, and other recent GRIDGEN improvements are described in this paper.

Steinbrenner, John P.↗

Crystal ball gazing

Over the last seven years, the CPU on my desk has increased speed by two orders of magnitude, from around 1 MIP to more than 100 MIPS; more important is that it is about as fast as any uniprocessor of any type available for any price, for compute bound problems. Memory on the system is also about 100 times as big, while disk is only about 10 times as big. Local network and I/O performance have increased greatly, though not quite at the same rate as processor speed. More important, I will argue, is that the CPU's address space is 64 bits, rather than 32 bits, allowing us to rethink some time honored presumptions. The Internet has gone from a few hundred machines to a million, and now have grown to span the entire globe, and wide area networks have now becoming commercial services. 'PC's' are now real computers, bringing what was top of the line computing capability to the masses only a few years behind the leading edge. So even a year or two from now, we can anticipate commonplace desktop machines running at speeds hundreds of MIPS, with main memories in the hundreds of megabytes to a gigabyte, able to draw millions of vectors/second, and all capable of some reasonable 3D graphics. And only a few years later, this will be the $1500 PC. So the 1990's certainly brings: 64 bit processors becoming standard; BIP/BFLOP class uniprocessors; large scale multiprocessors for special purpose applications; I/O as the most significant computer engineering problem; Hierarchical data servers in everyday use; routine access to archived data around the world; and what else? What do systems such as those we will have this decade imply to those building data analysis systems today? Many of the presumptions of the 1970's and 1980's need to be reexamined in the light of 1990's technology.

Gettys, Jim↗

PATHFINDER: Probing Atmospheric Flows in an Integrated and Distributed Environment

PATHFINDER is a software effort to create a flexible, modular, collaborative, and distributed environment for studying atmospheric, astrophysical, and other fluid flows in the evolving networked metacomputer environment of the 1990s. It uses existing software, such as HDF (Hierarchical Data Format), DTM (Data Transfer Mechanism), GEMPAK (General Meteorological Package), AVS, SGI Explorer, and Inventor to provide the researcher with the ability to harness the latest in desktop to teraflop computing. Software modules developed during the project are available in the public domain via anonymous FTP from the National Center for Supercomputing Applications (NCSA). The address is ftp.ncsa.uiuc.edu, and the directory is /SGI/PATHFINDER.

Wilhelmson, R. B.↗

Parallel Finite Element Solution of 3D Rayleigh-Benard-Marangoni Flows

A domain decomposition strategy and parallel gradient-type iterative solution scheme have been developed and implemented for computation of complex 3D viscous flow problems involving heat transfer and surface tension effects. Details of the implementation issues are described together with associated performance and scalability studies. Representative Rayleigh-Benard and microgravity Marangoni flow calculations and performance results on the Cray T3D and T3E are presented. The work is currently being extended to tightly-coupled parallel "Beowulf-type" PC clusters and we present some preliminary performance results on this platform. We also describe progress on related work on hierarchic data extraction for visualization.

Carey, G. F.↗

Instrument Remote Control via the Astronomical Instrument Markup Language

The Instrument Remote Control (IRC) project ongoing at NASA's Goddard Space Flight Center's (GSFC) Information Systems Center (ISC) supports NASA's mission by defining an adaptive intranet-based framework that provides robust interactive and distributed control and monitoring of remote instruments. An astronomical IRC architecture that combines the platform-independent processing capabilities of Java with the power of Extensible Markup Language (XML) to express hierarchical data in an equally platform-independent, as well as human readable manner, has been developed. This architecture is implemented using a variety of XML support tools and Application Programming Interfaces (API) written in Java. IRC will enable trusted astronomers from around the world to easily access infrared instruments (e.g., telescopes, cameras, and spectrometers) located in remote, inhospitable environments, such as the South Pole, a high Chilean mountaintop, or an airborne observatory aboard a Boeing 747. Using IRC's frameworks, an astronomer or other scientist can easily define the type of onboard instrument, control the instrument remotely, and return monitoring data all through the intranet. The Astronomical Instrument Markup Language (AIML) is the first implementation of the more general Instrument Markup Language (IML). The key aspects of our approach to instrument description and control applies to many domains, from medical instruments to machine assembly lines. The concepts behind AIML apply equally well to the description and control of instruments in general. IRC enables us to apply our techniques to several instruments, preferably from different observatories.

Sall, Ken↗

TRMM .25 deg x .25 deg Gridded Precipitation Text Product

Since the launch of the Tropical Rainfall Measuring Mission (TRMM), the Precipitation Measurement Missions science team has endeavored to provide TRMM precipitation retrievals in a variety of formats that are more easily usable by the broad science community than the standard Hierarchical Data Format (HDF) in which TRMM data is produced and archived. At the request of users, the Precipitation Processing System (PPS) has developed a .25 x .25 gridded product in an easily used ASCII text format. The entire TRMM mission data has been made available in this format. The paper provides the details of this new precipitation product that is designated with the TRMM designator 3G68.25. The format is packaged into daily files. It provides hourly precipitation information from the TRMM microwave imager (TMI), precipitation radar (PR), and TMI/PR combined rain retrievals. A major advantage of this approach is the inclusion only of rain data, compression when a particular grid has no rain from the PR or combined, and its direct ASCII text format. For those interested only in rain retrievals and whether rain is convection or stratiform, these products provide a huge reduction in the data volume inherent in the standard TRMM products. This paper provides examples of the 3G68 data products and their uses. It also provides information about C tools that can be used to aggregate daily files into larger time samples. In addition, it describes the possibilities inherent in the spatial sampling which allows resampling into coarser spatial sampling. The paper concludes with information about downloading the gridded text data products.

Stocker, Erich↗

Processing TES Level-1B Data

TES L1B Subsystem is a computer program that performs several functions for the Tropospheric Emission Spectrometer (TES). The term "L1B" (an abbreviation of "level 1B"), refers to data, specific to the TES, on radiometric calibrated spectral radiances and their corresponding noise equivalent spectral radiances (NESRs), plus ancillary geolocation, quality, and engineering data. The functions performed by TES L1B Subsystem include shear analysis, monitoring of signal levels, detection of ice build-up, and phase correction and radiometric and spectral calibration of TES target data. Also, the program computes NESRs for target spectra, writes scientific TES level-1B data to hierarchical- data-format (HDF) files for public distribution, computes brightness temperatures, and quantifies interpixel signal variability for the purpose of first-order cloud and heterogeneous land screening by the level-2 software summarized in the immediately following article. This program uses an in-house-developed algorithm, called "NUSRT," to correct instrument line-shape factors.

DeBaca, Richard C.↗

Guided Tour of Pythonian Museum

At http:hdfeos.orgzoo, we have a large collection of Python examples of dealing with NASA HDF (Hierarchical Data Format) products. During this hands-on Python tutorial session, we'll present a few common hacks to access and visualize local NASA HDF data. We'll also cover how to access remote data served by OPeNDAP (Open-source Project for a Network Data Access Protocol). As a glue language, we will demonstrate how you can use Python for your data workflow - from searching data to analyzing data with machine learning.

hdf↗

Task 28: Web Accessible APIs in the Cloud Trade Study

This study explored three candidate architectures for serving NASA Earth Science Hierarchical Data Format Version 5 (HDF5) data via Hyrax running on Amazon Web Services (AWS). We studied the cost and performance for each architecture using several representative Use-Cases. The objectives of the project are: Conduct a trade study to identify one or more high performance integrated solutions for storing and retrieving NASA HDF5 and Network Common Data Format Version 4 (netCDF4) data in a cloud (web object store) environment. The target environment is Amazon Web Services (AWS) Simple Storage Service (S3).Conduct needed level of software development to properly evaluate solutions in the trade study and to obtain required benchmarking metrics for input into government decision of potential follow-on prototyping. Develop a cloud cost model for the preferred data storage solution (or solutions) that accounts for different granulation and aggregation schemes as well as cost and performance trades.

cost model↗

Additional Metadata Guidelines to Improve the Structure and Usability of HDF and NetCDF Files

The Hierarchical Data Format (HDF) and Network Common Data Form (NetCDF) are data file formats created to aid users in the creation or use of scientific data. These file formats are useful for handling large data volumes and hosting extensive metadata as global attributes or variables and are popular with the modeling community. HDF and NetCDF files are largely used with remote sensing data and have been used to support measurements from numerous campaigns, from satellite to aircraft or ground and mobile based measurements. The files from airborne field studies, however, vary greatly in terms of the file structure and the amount and content of metadata. Information relevant to the file that can be useful to the user such as the data producer, location where data was taken, variable descriptions, or information about the instrument might not be included in the file. This metadata might be present in another file in the dataset containing the same data using the International Consortium for Atmospheric Research on Transport and Transformation (ICARTT) format. Recently, the Aerosols, Clouds, and their Interactions for Earth System Models (MACIE) group started a grassroots effort to develop a set of requirements for the HDF and NetCDF files for field studies, aiming to make the data products more interoperable and usable. Particularly, these requirements seek to make the files more compliant to Climate and Forecast (CF) metadata conventions and to standardize the file structure and the global and variable attributes. These requirements would help to ensure that HDF and NetCDF files contain adequate metadata to better support their use for research, e.g., the modeling community, and to enhance the usability and interoperability of data for research communities at large. To be presented are the details of the MACIE requirements as well as examples of the implementation of these requirements for merge files and lidar observation data files.

Sean Leavor↗

Modified Recursive Hierarchical Segmentation of Data

An algorithm and a computer program that implements the algorithm that performs recursive hierarchical segmentation (RHSEG) of data have been developed. While the current implementation is for two-dimensional data having spatial characteristics (e.g., image, spectral, or spectral-image data), the generalized algorithm also applies to three-dimensional or higher dimensional data and also to data with no spatial characteristics. The algorithm and software are modified versions of a prior RHSEG algorithm and software, the outputs of which often contain processing-window artifacts including, for example, spurious segmentation-image regions along the boundaries of processing-window edges.

Tilton, James C.↗

Utilizing Hierarchical Segmentation to Generate Water and Snow Masks to Facilitate Monitoring Change with Remotely Sensed Image Data

The hierarchical segmentation (HSEG) algorithm is a hybrid of hierarchical step-wise optimization and constrained spectral clustering that produces a hierarchical set of image segmentations. This segmentation hierarchy organizes image data in a manner that makes the image's information content more accessible for analysis by enabling region-based analysis. This paper discusses data analysis with HSEG and describes several measures of region characteristics that may be useful analyzing segmentation hierarchies for various applications. Segmentation hierarchy analysis for generating landwater and snow/ice masks from MODIS (Moderate Resolution Imaging Spectroradiometer) data was demonstrated and compared with the corresponding MODIS standard products. The masks based on HSEG segmentation hierarchies compare very favorably to the MODIS standard products. Further, the HSEG based landwater mask was specifically tailored to the MODIS data and the HSEG snow/ice mask did not require the setting of a critical threshold as required in the production of the corresponding MODIS standard product.

Tilton, James C.↗

A physics-informed and hierarchically regularized data-driven model for predicting fluid flow through porous media

This paper presents a new deep learning data-driven model for predicting structure dependent pore-fluid velocity fields in rock. The model is based on a Convolutional Auto-Encoder (CAE) artificial neural network capable of learning from image data generated by direct numerical simulations of fluid flow through pore-structures, such as by Lattice Boltzmann or molecular dynamics methods. The main novelty of the model in comparison to previous CAE-based data-driven approaches consists of three parts. The first is a methodology for decomposing the full-domain of the porous media into sub-regions, or “sub-domains”, in order to reduce the overall size of the CAE, batch process the sub-domains in parallel, and enable the CAE to learn local and generalizable nonlinear mappings of pore-fluid velocities. The second consists of embedding the finite difference solutions of the incompressible Navier-Stokes and continuity equations into convolutional layers prior to the CAE in order to provide the CAE with knowledge of fluid dynamics physics (PhyFlow). The third main novelty is that the training of the CAE is regularized with a hierarchical loss function that encourages the learning of fluid flow patterns (in a way similar to ranked modes in principal component analysis), ranking from most to least important. This is shown to increase the stability in learning, reduce over-fitting, and promote interpretability of the CAE neural network layers (HierCAE). The comprehensive new data-driven model, which we call the PhyFlow-HierCAE model, is shown to exhibit improved accuracy and generalizability of flow field predictions over conventional CAE models, attributable to the embedded physical knowledge and the hierarchical regularization, as well as realize orders of magnitude speed-ups in computation times as a surrogate for the direct numerical simulations. Examples of training and forward predictions on unseen pore-structures are provided and evaluated for data from Lattice Boltzmann and molecular dynamics simulations of pore-fluid flow. The model is shown to be a fast and accurate emulator (or “surrogate”) for predicting effective permeability of unseen pore-structures based on learning from relatively small direct numerical simulation datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Deep Hierarchical Super Resolution for Scientific Data

We present a novel technique for hierarchical super resolution (SR) with neural networks (NNs), which upscales volumetric data represented with an octree data structure to a high-resolution uniform gridwith minimal seam artifacts on octree node boundaries. Our method uses existing state-of-the-art SR models and adds flexibility to upscale input data with varying levels of detail across the domain, instead of only uniform grid data that are supported in previous approaches.The key is to use a hierarchy of SR NNs, each trained to perform 2x SR between two levels of detail, with a hierarchical SR algorithm that minimizes seam artifacts by starting from the coarsest level of detail and working up.We show that our hierarchical approach outperforms baseline interpolation and hierarchical upscaling methods, and demonstrate the usefulness of our proposed approach across three use cases including data reduction using hierarchical downsampling+SR instead of uniform downsampling+SR, computation savings for hierarchical finite-time Lyapunov exponent field calculation, and super-resolving low-resolution simulation results for a high-resolution approximation visualization.

97 MATHEMATICS AND COMPUTING↗

DDStore: Distributed Data Store for Scalable Training of Graph Neural Networks on Large Atomistic Modeling Datasets

Graph neural networks (GNNs) are a class of Deep Learning models used in designing atomistic materials for effective screening of large chemical spaces. To ensure robust prediction, GNN models must be trained on large volumes of atomistic data on leadership class supercomputers. Even with the advent of modern architectures that consist of multiple storage layers that include node-local NVMe devices in addition to device memory for caching large datasets, extreme-scale model training faces I/O challenges at scale.We present DDStore, an in-memory distributed data store designed for GNN training on large-scale graph data. DDStore provides a hierarchical, distributed, data caching technique that combines data chunking, replication, low-latency random access, and high throughput communication. DDStore achieves near-linear scaling for training a GNN model using up to 1000 GPUs on the Summit and Perlmutter supercomputers, and reaches up to a 6.15x reduction in GNN training time compared to state-of-the-art methodologies.

Choi, Jong Youl↗

Spectral analysis for automated exploration and sample acquisition

Future space exploration missions will rely heavily on the use of complex instrument data for determining the geologic, chemical, and elemental character of planetary surfaces. One important instrument is the imaging spectrometer, which collects complete images in multiple discrete wavelengths in the visible and infrared regions of the spectrum. Extensive computational effort is required to extract information from such high-dimensional data. A hierarchical classification scheme allows multispectral data to be analyzed for purposes of mineral classification while limiting the overall computational requirements. The hierarchical classifier exploits the tunability of a new type of imaging spectrometer which is based on an acousto-optic tunable filter. This spectrometer collects a complete image in each wavelength passband without spatial scanning. It may be programmed to scan through a range of wavelengths or to collect only specific bands for data analysis. Spectral classification activities employ artificial neural networks, trained to recognize a number of mineral classes. Analysis of the trained networks has proven useful in determining which subsets of spectral bands should be employed at each step of the hierarchical classifier. The network classifiers are capable of recognizing all mineral types which were included in the training set. In addition, the major components of many mineral mixtures can also be recognized. This capability may prove useful for a system designed to evaluate data in a strange environment where details of the mineral composition are not known in advance.

Eberlein, Susan↗

The Mars 2020 Ground Data System Architecture

The Mars 2020 Mission’s primary objective is to collect 20 geographically unique samples during its prime mission of one and a quarter Martian years, or just over 2 Earth years. Mission planners determined the project needed to develop a system that would enable the operations team to analyze engineering and science data, make science decisions, select viable rover targets at a millimeter resolution and validate an uplink bundle for a car sized rover with more complex science instruments than any previous Mars surface mission. All this had to be done within a five hour time frame. Doing this with a small team would be a challenge, but this had to be accomplished by a large team of engineers and scientists located across North America and Europe. Achieving this level of operational efficiency was unheard of in the prime mission. In addition, the mission had another set of requirements that had nothing to do with surface operations; the Mars 2020 Ground Data System (GDS) was also expected to comply with a new set of security requirements to keep up with the ever changing cybersecurity landscape. The Mars 2020 Ground Data System (GDS) is a re-architected version of the Mars Science Laboratory GDS. The primary goal was to integrate the lessons learned from previous Mars surface missions, accommodate a set of new requirements and capabilities required to ensure mission success, and comply with a new set of cybersecurity controls. The new architecture includes several unique qualities including a data lake, language-agnostic system-wide event-based operations, containerization, automated deployment, network segmentation, infrastructure-as-code, API-driven interfaces, and the first Mars surface GDS to operate primarily in the cloud. The new architecture enabled greater access to the system’s data, tighter integration with the operations team, and a higher level of traceability. The availability of the data also enabled a new set of capabilities previously not possible on surface missions. These new capabilities include an autonomous data to information, pipeline for downlink analysis, horizontal scaling of science data processing capabilities, autonomous round trip data tracking of science and engineering data, integration of flight system state into the tactical planning cycle, high fidelity targeting utilizing kinematic data, and hierarchical image and 3d meshes data representations. This paper will introduce the requirements for the Mars 2020 Mission, the heritage architecture, and the rationale for the changes to achieve the new architecture. The paper will continue to describe the fundamental changes made to the GDS architecture, how these changes enabled a more tightly integrated GDS, and the new capabilities that were enabled by the new architecture. The paper will conclude with the lessons learned from the process of rearchitecting a heritage GDS system and from the first 200 days of operations supporting over 800 users from around the world.

Lopez-Roig, Reynaldo↗