Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

SAR image data compression for an on-line archive system

This paper summarizes the investigation of SAR image data compression for an on-line archive data distribution system. This system is planned for the ground processing system of Alaska SAR Facility (ASF) and Shuttle Imaging Radar (SIR-C). The objective of the SAR image data compression is to enable the data archive system to provide the remote users a large data base with good image quality, short response time, low transfer cost, and minimal decoding complexity. The requirements and limitations of the on-line archive data distribution system are presented. The effects of SAR image data characteristics on data compression are addressed. The users' survey results suggest that compression ratios between 10:1 and 20:1 appear suitable. Based on the algorithm evaluation results, the two-level tree-searched vector quantization technique has been recommended as the SAR image data compression algorithm for the on-line archive data distribution system.

Chang, C. Y.↗

Magnetic pair distribution function data using polarized neutrons and ad hoc corrections

Here, we report the first example of magnetic pair distribution function (mPDF) data obtained through the use of neutron polarization analysis. Using the antiferromagnetic semiconductor MnTe as a test case, we present high-quality mPDF data collected on the HYSPEC instrument at the Spallation Neutron Source using longitudinal polarization analysis to isolate the magnetic scattering cross section. Clean mPDF patterns are obtained for MnTe in both the magnetically ordered state and the correlated paramagnet state, where only short-range magnetic order is present. We also demonstrate significant improvement in the quality of high-resolution mPDF data through the application of ad hoc corrections that require only minimal human input, minimizing potential sources of error in the data processing procedure. We briefly discuss the current limitations and future outlook of mPDF analysis using polarized neutrons. Overall, this work provides a useful benchmark for mPDF analysis using polarized neutrons and provides an encouraging picture of the potential for routine collection of high-quality mPDF data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Extracting Resilience Metrics From Distribution Utility Data Using Outage and Restore Process Statistics

Resilience curves track the accumulation and restoration of outages during an event on an electric distribution grid. We show that a resilience curve generated from utility data can always be decomposed into an outage process and a restore process and that these processes generally overlap in time. We use many events in real utility data to characterize the statistics of these processes, and derive formulas based on these statistics for resilience metrics such as restore duration, customer hours not served, and outage and restore rates. The formulas express the mean value of these metrics as a function of the number of outages in the event. We also give a formula for the variability of restore duration, which allows us to predict a maximum restore duration with 95% confidence. Overall, we give a simple and general way to decompose resilience curves into outage and restore processes and then show how to use these processes to extract resilience metrics from standard distribution system data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ARENA: Asynchronous Reconfigurable Accelerator Ring to Enable Data-Centric Parallel Computing

The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications. In this work, we propose ARENA – an asynchronous reconfigurable accelerator ring architecture as a potential scenario on how the future HPC and data centers will be like. Despite using the coarse-grained reconfigurable arrays (CGRAs) as the substrate platform, our key contribution is not only the CGRA-cluster design itself, but also the ensemble of a new architecture and programming model that enables asynchronous tasking across a cluster of reconfigurable nodes, so as to bring specialized computation to the data rather than the reverse. We presume distributed data storage without asserting any prior knowledge on the data distribution. Hardware specialization occurs at runtime when a task finds the majority of data it requires are available at the present node. In other words, we dynamically generate specialized CGRA accelerators where the data reside. The asynchronous tasking for bringing computation to data is achieved by circulating the task token, which describes the dataflow graphs to be executed for a task, among the CGRA cluster connected by a fast ring network. Evaluations on a set of HPC and data-driven applications across different domains show that ARENA can provide better parallel scalability with reduced data movement (53.9 percent). Compared with contemporary compute-centric parallel models, ARENA can bring on average 4.37× speedup. The synthesized CGRAs and their task-dispatchers only occupy 2.93mm 2 chip area under 45nm process technology and can run at 800MHz with on average 759.8mW power consumption. ARENA also supports the concurrent execution of multi-applications, offering ideal architectural support for future high-performance parallel computing and data analytics systems.

97 MATHEMATICS AND COMPUTING↗

Seasat-A ASVT: Commercial demonstration experiments. Results analysis methodology for the Seasat-A case studies

The SEASAT-A commercial demonstration program ASVT is described. The program consists of a set of experiments involving the evaluation of a real time data distributions system, the SEASAT-A user data distribution system, that provides the capability for near real time dissemination of ocean conditions and weather data products from the U.S. Navy Fleet Numerical Weather Central to a selected set of commercial and industrial users and case studies, performed by commercial and industrial users, using the data gathered by SEASAT-A during its operational life. The impact of the SEASAT-A data on business operations is evaluated by the commercial and industrial users. The approach followed in the performance of the case studies, and the methodology used in the analysis and integration of the case study results to estimate the actual and potential economic benefits of improved ocean condition and weather forecast data are described.

Source record↗

Data-Driven Distributed Algorithms for Estimating Eigenvalues and Eigenvectors of Interconnected Dynamical Systems

Here, the paper presents data-driven algorithms to estimate in a distributed manner the eigenvalues, right and left eigenvectors of an unknown linear (or linearized) interconnected dynamic system. In particular, the proposed algorithms do not require the identification of the system model in advance before performing the estimation. As a first step, we consider interconnected dynamical system with distinct eigenvalues. The proposed strategy first estimates the eigenvalues using the well-known Prony method. The right and left eigenvectors are then estimated by solving distributively a set of linear equations. One important feature of the proposed algorithms is that the topology of communication network used to perform the distributed estimation can be chosen arbitrarily, given that it is connected, and is also independent of the structure or sparsity of the system (state) matrix. The proposed distributed algorithms are demonstrated via a numerical example.

97 MATHEMATICS AND COMPUTING↗

Validation of non-negative matrix factorization for rapid assessment of large sets of atomic pair distribution function data

The use of the non-negative matrix factorization (NMF) technique is validated for automatically extracting physically relevant components from atomic pair distribution function (PDF) data from time-series data such as in situ experiments. The use of two matrix-factorization techniques, principal component analysis and NMF, on PDF data is compared in the context of a chemical synthesis reaction taking place in a synchrotron beam, applying the approach to synthetic data where the correct composition is known and on measured PDFs from previously published experimental data. The NMF approach yields mathematical components that are very close to the PDFs of the chemical components of the system and a time evolution of the weights that closely follows the ground truth. Lastly, it is discussed how this would appear in a streaming context if the analysis were being carried out at the beamline as the experiment progressed.

36 MATERIALS SCIENCE↗

Distributed Lunar Data Platform with Advanced Machine Learning Capabilities in Support of Lunar Science and Exploration

The United States 2020 Space Policy directive declares that NASA, in cooperation with private industry, will “extend human economic activity into deep space by establishing a permanent human presence on the Moon”. This goal will require advanced data management, as well as analysis, modeling and representation of lunar information in order to prepare for Artemis human missions, lunar science investigations and exploration. To meet this requirement, we conceptualize and present an implementation strategy for a distributed platform for lunar data retrieval, inferencing and analysis, which will be based on federated learning and the NASA Celestial Mapping System (CMS). In addition to demonstrating the imperative of enabling lunar-borne data to remain in-situ but still accessible, this presentation will also include examples of how third parties could contribute both datasets and new functionality into this platform using an AI-based data import pipeline and a plug-in architecture respectively.

Artificial Intelligence↗

Transfer-AE: A novel autoencoder-based impact detection model for structural digital twin

Accurately detecting the location and intensity of impacts is crucial for ensuring structural safety. Currently, AI-based structural impact detection methods are widely used for their excellent detection accuracy. However, their generalization capability is limited by the scenarios present in the training data. Many complex and dangerous impact scenarios are difficult to conduct real-world experiments on to collect sufficient samples. To capture all impact scenarios and fully leverage the advantages of AI-based detection technologies, advanced methods involve combining real-world structural monitoring data with corresponding numerical models to construct digital twins. These methods continuously refine the created numerical models with limited real-world data and provide diverse impact scenarios through numerical model simulations. However, there are inevitable differences between digital models and physical models that are challenging to correct through mechanical means. This discrepancy in data distribution between the two models significantly hinders the application of digital twin technology in impact/event identification tasks. To address this challenge, this study proposes a novel model based on autoencoders, named Transfer-AE. Transfer-AE encodes the common features of digital twins in the latent space to bridge the uncertainty gap at a macro scale between numerical models and physical models and synchronously fits the magnitude and location of the impact load in the decoder. This enables consistent detection results for the same impact event, whether the sample comes from the numerical model or the physical model. Transfer-AE includes two operating modes: Mode 1 has a fixed computational complexity with stable inference speed, but the training cost and difficulty increase with data distribution. Mode 2's computational complexity increases with data distribution, but it has a fixed training cost and speed. In both cases involving the geodesic dome structure simulating a deep space habitat and the IASC-ASCE benchmark structure, Transfer-AE demonstrated the best performance in impact localization and quantification tasks compared to mainstream domain-adaptive transfer models.

Chengjia Han↗

Bayesian inference for the seismic moment tensor using regional waveforms and a data-derived distribution of velocity models

The largest source of uncertainty in any source inversion is the velocity model used to construct the transfer function employed in the forward model that relates observed ground motion to the seismic moment tensor. However, standard inverse procedures often does not quantify uncertainty in the seismic moment tensor due to error in the Green’s functions from uncertain event location and Earth structure. We attempt to incorporate this uncertainty into an estimation of the seismic moment tensor using a distribution of velocity models calculated in a prior effort based on different and complementary data sets. The posterior distribution of velocity models is then used to construct Green’s functions for use in Bayesian inference of an unknown seismic moment tensor using regional waveform data. The combined likelihood is estimated using data-specific error models and the posterior of the seismic moment tensor is estimated and can be interpreted in terms of most-probable source-type.

58 GEOSCIENCES↗

A Future Vision of a Data Acquisition: Distributed Sensing, Processing, and Health Monitoring

This paper presents a vision fo a highly enhanced data acquisition and health monitoring system at NASA Stennis Space Center (SSC) rocket engine test facility. This vision includes the use of advanced processing capabilities in conjunction with highly autonomous distributed sensing and intelligence, to monitor and evaluate the health of data in the context of it's associated process. This method is expected to significantly reduce data acquisitions costs and improve system reliability. A Universal Signal Conditioning Amplifier (USCA) based system, under development at Kennedy Space Center, is being evaluated for adaptation to the SSC testing infrastructure. Kennedy's USCA architecture offers many advantages including flexible and auto-configuring data acquisition with improved calibration and verifiability. Possible enhancements at SSC may include multiplexing the distributed USCAs to reduce per channel cost, and the use of IEEE-485 to Allen-Bradley Control Net Gateways for interfacing with the resident control systems.

Figueroa, Fernando↗

TuckerMPI: A Parallel C++/MPI Software Package for Large-scale Data Compression via the Tucker Tensor Decomposition

With this study, our goal is compression of massive-scale grid-structured data, such as the multi-terabyte output of a high-fidelity computational simulation. For such data sets, we have developed a new software package called TuckerMPI, a parallel C++/MPI software package for compressing distributed data. The approach is based on treating the data as a tensor, i.e., a multidimensional array, and computing its truncated Tucker decomposition, a higher-order analogue to the truncated singular value decomposition of a matrix. The result is a low-rank approximation of the original tensor-structured data. Compression efficiency is achieved by detecting latent global structure within the data, which we contrast to most compression methods that are focused on local structure. In this work, we describe TuckerMPI, our implementation of the truncated Tucker decomposition, including details of the data distribution and in-memory layouts, the parallel and serial implementations of the key kernels, and analysis of the storage, communication, and computational costs. We test the software on 4.5 and 6.7 terabyte data sets distributed across 100 s of nodes (1,000 s of MPI processes), achieving compression ratios between 100 and 200,000×, which equates to 99--99.999% compression (depending on the desired accuracy) in substantially less time than it would take to even read the same dataset from a parallel file system. Moreover, we show that our method also allows for reconstruction of partial or down-sampled data on a single node, without a parallel computer so long as the reconstructed portion is small enough to fit on a single machine, e.g., in the instance of reconstructing/visualizing a single down-sampled time step or computing summary statistics. The code is available at https://gitlab.com/tensors/TuckerMPI.

97 MATHEMATICS AND COMPUTING↗

Multimission Telemetry Visualization (MTV) system: A mission applications project from JPL's Multimedia Communications Laboratory

This paper describes the Multimission Telemetry Visualization (MTV) data acquisition/distribution system. MTV was developed by JPL's Multimedia Communications Laboratory (MCL) and designed to process and display digital, real-time, science and engineering data from JPL's Mission Control Center. The MTV system can be accessed using UNIX workstations and PC's over common datacom and telecom networks from worldwide locations. It is designed to lower data distribution costs while increasing data analysis functionality by integrating low-cost, off-the-shelf desktop hardware and software. MTV is expected to significantly lower the cost of real-time data display, processing, distribution, and allow for greater spacecraft safety and mission data access.

Koeberlein, Ernest, III↗

Reliable Measures of Spread in High Dimensional Latent Spaces

Understanding geometric properties of the latent spaces of natural language processing models allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model’s latent space, or how fully the available latent space is being used. We demonstrate that the commonly used measures of data spread, average cosine similarity and a partition function min/max ratio I (V), do not provide reliable metrics to compare the use of latent space across data distributions. We propose and examine six alternative measures of data spread, all of which improve over these current metrics when applied to seven synthetic data distributions. Of our proposed measures, we recommend one principal component-based measure and one entropy-based measure that provide reliable, relative measures of spread and can be used to compare models of different sizes and dimensionalities.

97 MATHEMATICS AND COMPUTING↗

Airborne Aerosol Closure Studies During PRIDE

The Puerto Rico Dust Experiment (PRIDE) was conducted during June/July of 2000 to study the properties of Saharan dust aerosols transported across the Atlantic Ocean to the Caribbean Islands. During PRIDE, the NASA Ames Research Center six-channel (380 - 1020 nm) airborne autotracking sunphotometer (AATS-6) was operated aboard a Piper Navajo airplane alongside a suite of in situ aerosol instruments. The in situ aerosol instrumentation relevant to this paper included a Forward Scattering Spectrometer Probe (FSSP-100) and a Passive Cavity Aerosol Spectrometer Probe (PCASP), covering the radius range of approx. 0.05 to 10 microns. The simultaneous and collocated measurement of multi-spectral aerosol optical depth and in situ particle size distribution data permits a variety of closure studies. For example, vertical profiles of aerosol optical depth obtained during local aircraft ascents and descents can be differentiated with respect to altitude and compared to extinction profiles calculated using the in situ particle size distribution data (and reasonable estimates of the aerosol index of refraction). Additionally, aerosol extinction (optical depth) spectra can be inverted to retrieve estimates of the particle size distributions, which can be compared directly to the in situ size distributions. In this paper we will report on such closure studies using data from a select number of vertical profiles at Cabras Island, Puerto Rico, including measurements in distinct Saharan Dust Layers. Preliminary results show good agreement to within 30% between mid-visible aerosol extinction derived from the AATS-6 optical depth profiles and extinction profiles forward calculated using 60s-average in situ particle size distributions and standard Saharan dust aerosol refractive indices published in the literature. In agreement with tendencies observed in previous studies, our initial results show an underestimate of aerosol extinction calculated based on the in situ size distributions relative to the extinction obtained from the sunphotometer measurements. However, a more extensive analysis of all available AATS-6 and in situ size distribution data is necessary to ascertain whether the preliminary results regarding the degree of extinction closure is representative of the entire range of dust conditions encountered in PRIDE. Finally, we will compare the spectral extinction measurements obtained in PRIDE to similar data obtained in Saharan dust layers encountered above the Canary Islands during ACE-2 (Aerosol Characterization Experiment) in July 1997. Thus, the evolution of Saharan dust spectral properties during its transport across the Atlantic can be investigated, provided the dust origin and microphysical properties are found to be comparable.

Redemann, Jens↗

Automatic generation of efficient array redistribution routines for distributed memory multicomputers

Appropriate data distribution has been found to be critical for obtaining good performance on Distributed Memory Multicomputers like the CM-5, Intel Paragon and IBM SP-1. It has also been found that some programs need to change their distributions during execution for better performance (redistribution). This work focuses on automatically generating efficient routines for redistribution. We present a new mathematical representation for regular distributions called PITFALLS and then discuss algorithms for redistribution based on this representation. One of the significant contributions of this work is being able to handle arbitrary source and target processor sets while performing redistribution. Another important contribution is the ability to handle an arbitrary number of dimensions for the array involved in the redistribution in a scalable manner. Our implementation of these techniques is based on an MPI-like communication library. The results presented show the low overheads for our redistribution algorithm as compared to naive runtime methods.

Ramaswamy, Shankar↗