Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

An Update on the Geothermal Data Repository's Data Standards and Pipelines: Geospatial Data and Distributed Acoustic Sensing Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.

cloud-optimized↗

Optimal PV Inverter Control in Distribution Systems via Data-Driven Distributionally Robust Optimization

Distribution systems with high penetration of uncertain solar generation call for advanced control strategies of photovoltaics (PVs) inverters. This paper proposes a data-driven distributionally robust optimization (DDDRO) approach to optimally controlling the PV inverters to improve the system operation performance under solar power uncertainties. In the proposed DDDRO approach, a Wasserstein ball-based method is proposed to construct the distributional ambiguity set to model the uncertainties of PV generation through partial observations of historical data without knowing exact probability distributions. We further reformulate the computationally intractable DDDRO model to a mixed integer second order cone programming (MISOCP) problem. The effectiveness and out-of-sample performance of the proposed approach have been demonstrated on a modified IEEE 33-node system. We conduct a comparative study to compare the proposed method with traditional chance constrained programming (CCP). It shows that the proposed DDDRO approach can provide a less conservative yet robust solution to minimize the worse-case expectation of the total network loss while maintaining nodal voltages in a secure range.

Xue, Yaosuo↗

Packaging and distributing ecological data from multisite studies

Studies of global change and other regional issues depend on ecological data collected at multiple study areas or sites. An information system model is proposed for compiling diverse data from dispersed sources so that the data are consistent, complete, and readily available. The model includes investigators who collect and analyze field measurements, science teams that synthesize data, a project information system that collates data, a data archive center that distributes data to secondary users, and a master data directory that provides broader searching opportunities. Special attention to format consistency is required, such as units of measure, spatial coordinates, dates, and notation for missing values. Often data may need to be enhanced by estimating missing values, aggregating to common temporal units, or adding other related data such as climatic and soils data. Full documentation, an efficient data distribution mechanism, and an equitable way to acknowledge the original source of data are also required.

Information Systems↗

An Offload NIC for NASA, NLR, and Grid Computing

This work addresses distributed data management and access dynamically configurable high-speed access to data distributed and shared over wide-area high-speed network environments. An offload engine NIC (network interface card) is proposed that scales at nX10-Gbps increments through 100-Gbps full duplex. The Globus de facto standard was used in projects requiring secure, robust, high-speed bulk data transport. Novel extension mechanisms were derived that will combine these technologies for use by GridFTP, bandwidth management resources, and host CPU (central processing unit) acceleration. The result will be wire-rate encrypted Globus grid data transactions through offload for splintering, encryption, and compression. As the need for greater network bandwidth increases, there is an inherent need for faster CPUs. The best way to accelerate CPUs is through a network acceleration engine. Grid computing data transfers for the Globus tool set did not have wire-rate encryption or compression. Existing technology cannot keep pace with the greater bandwidths of backplane and network connections. Present offload engines with ports to Ethernet are 32 to 40 Gbps f-d at best. The best of ultra-high-speed offload engines use expensive ASICs (application specific integrated circuits) or NPUs (network processing units). The present state of the art also includes bonding and the use of multiple NICs that are also in the planning stages for future portability to ASICs and software to accommodate data rates at 100 Gbps. The remaining industry solutions are for carrier-grade equipment manufacturers, with costly line cards having multiples of 10-Gbps ports, or 100-Gbps ports such as CFP modules that interface to costly ASICs and related circuitry. All of the existing solutions vary in configuration based on requirements of the host, motherboard, or carriergrade equipment. The purpose of the innovation is to eliminate data bottlenecks within cluster, grid, and cloud computing systems, and to add several more capabilities while reducing space consumption and cost. Provisions were designed for interoperability with systems used in the NASA HEC (High-End Computing) program. The new acceleration engine consists of state-ofthe- art FPGA (field-programmable gate array) core IP, C, and Verilog code; novel communication protocol; and extensions to the Globus structure. The engine provides the functions of network acceleration, encryption, compression, packet-ordering, and security added to Globus grid or for cloud data transfer. This system is scalable in nX10-Gbps increments through 100-Gbps f-d. It can be interfaced to industry-standard system-side or network-side devices or core IP in increments of 10 GigE, scaling to provide IEEE 40/100 GigE compliance.

Awrach, James↗

Ground System for Solar Dynamics Observatory (SDO) Mission

NASA s Goddard Space Flight Center (GSFC) has recently completed its Critical Design Review (CDR) of a new dual Ka and S-band ground system for the Solar Dynamics Observatory (SDO) Mission. SDO, the flagship mission under the new Living with a Star Program Office, is one of GSFC s most recent large-scale in-house missions. The observatory is scheduled for launch in August 2008 from the Kennedy Space Center aboard an Atlas-5 expendable launch vehicle. Unique to this mission is an extremely challenging science data capture requirement. The mission is required to capture 99.99% of available science over 95% of all observation opportunities. Due to the continuous, high volume (150 Mbps) science data rate, no on-board storage of science data will be implemented on this mission. With the observatory placed in a geo-synchronous orbit at 36,000 kilometers within view of dedicated ground stations, the ground system will in effect implement a "real-time" science data pipeline with appropriate data accounting, data storage, data distribution, data recovery, and automated system failure detection and correction to keep the science data flowing continuously to three separate Science Operations Centers (SOCs). Data storage rates of approx. 45 Tera-bytes per month are expected. The Mission Operations Center (MOC) will be based at GSFC and is designed to be highly automated. Three SOCs will share in the observatory operations, each operating their own instrument. Remote operations of a multi-antenna ground station in White Sands, New Mexico from the MOC is part of the design baseline.

Tann, Hun K.↗

Distributed Data-Driven Power Iteration for Strongly Connected Networks

Here, this paper presents data-driven power iteration to distributively estimate the dominant eigenvalues of an unknown linear time-invariant system. The proposed strategy only requires a single trajectory data or measurements. Furthermore, in order to perform the distributed estimation, the communication network topology can be chosen to be any strongly connected directed graphs. The proposed data-driven power iteration is demonstrated using several numerical examples and is then applied to estimate the generalized algebraic connectivity of cooperative systems and to control the epidemic spreading.

Gusrialdi, Azwirman↗

Coarrars for Parallel Processing

The design of the Coarray feature of Fortran 2008 was guided by answering the question "What is the smallest change required to convert Fortran to a robust and efficient parallel language." Two fundamental issues that any parallel programming model must address are work distribution and data distribution. In order to coordinate work distribution and data distribution, methods for communication and synchronization must be provided. Although originally designed for Fortran, the Coarray paradigm has stimulated development in other languages. X10, Chapel, UPC, Titanium, and class libraries being developed for C++ have the same conceptual framework.

Fortran↗

Symmetry-mode analysis for local structure investigations using pair distribution function data

Symmetry-adapted distortion modes provide a natural way of describing distorted structures derived from higher-symmetry parent phases. Structural refinements using symmetry-mode amplitudes as fit variables have been used for at least ten years in Rietveld refinements of the average crystal structure from diffraction data; more recently, this approach has also been used for investigations of the local structure using real-space pair distribution function (PDF) data. Here, the value of performing symmetry-mode fits to PDF data is further demonstrated through the successful application of this method to two topical materials: TiSe2, where a subtle but long-range structural distortion driven by the formation of a charge-density wave is detected, and MnTe, where a large but highly localized structural distortion is characterized in terms of symmetry-lowering displacements of the Te atoms. Here, the analysis is performed using fully open-source code within the DiffPy framework via two packages developed for this work: isopydistort, which provides a scriptable interface to the ISODISTORT web application for group theoretical calculations, and isopytools, which converts the ISODISTORT output into a DiffPy-compatible format for subsequent fitting and analysis. These developments expand the potential impact of symmetry-adapted PDF analysis by enabling high-throughput analysis and removing the need for any commercial software.

36 MATERIALS SCIENCE↗

Distributed Data-Driven Optimization for Voltage Regulation in Distribution Systems

Here, this paper proposes a distributed data-driven optimization framework for voltage regulation in distribution systems. The recursive kernel regression and alternating direction method of multipliers (ADMM) are selected to cover the system learning and distributed optimization tasks. The proposed distributed data-driven framework is capable of having a rapid response to system or load changes while considering the operation optimality. Besides, the distributed algorithm parallels the computation tasks and reduces the computational expense of a single agent. To validate the performance of the proposed method, a hypothetical 7-Bus system and the IEEE 123-Bus system are selected to show the effectiveness of the proposed data-driven framework. According to the numerical study results, the proposed method offers great flexibility for selecting customized kernel models for different regions and can effectively improve the system voltage profile in a distributed manner.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

A Web Service and Android Application for the Distribution of Rainfall Estimates and Earth Observation Data

The full potential of Satellite Rainfall Estimates (SRE) can only be realized if timely access to the datasets is possible. Existing data distribution web portals are often focused on global products and offer limited customization options, especially for the purpose of routine regional monitoring. Furthermore, most online systems are designed to meet the needs of desktop users, limiting the compatibility with mobile devices. In response to the growing demand for SRE and to address the current limitations of available web portals a project was devised to create a set of freely available applications and services, available at a common portal that can: (1) simplify cross-platform access to Tropical Rainfall Measuring Mission Online Visualization and Analysis System (TOVAS) data (including from Android mobile devices), (2) provide customized and continuous monitoring of SRE in response to user demands and (3) combine data from different online data distribution services, including rainfall estimates, river gauge measurements or imagery from Earth Observation missions at a single portal, known as the Tropical Rainfall Measuring Mission (TRMM) Explorer. The TRMM Explorer project suite includes a Python-based web service and Android applications capable of providing SRE and ancillary data in different intuitive formats with the focus on regional and continuous analysis. The outputs include dynamic plots, tables and data files that can also be used to feed downstream applications and services. A case study in Southern Angola is used to describe the potential of the TRMM Explorer for SRE distribution and analysis in the context of ungauged watersheds. The development of a collection of data distribution instances helped to validate the concept and identify the limitations of the program, in a real context and based on user feedback. The TRMM Explorer can successfully supplement existing web portals distributing SRE and provide a cost-efficient resource to small and medium-sized organizations with specific SRE monitoring needs, namely in developing and transition countries.

precipitation↗

A relational data-knowledge base system and its potential in developing a distributed data-knowledge system

A new approach used in constructing a rational data knowledge base system is described. The relational database is well suited for distribution due to its property of allowing data fragmentation and fragmentation transparency. An example is formulated of a simple relational data knowledge base which may be generalized for use in developing a relational distributed data knowledge base system. The efficiency and ease of application of such a data knowledge base management system is briefly discussed. Also discussed are the potentials of the developed model for sharing the data knowledge base as well as the possible areas of difficulty in implementing the relational data knowledge base management system.

Rahimian, Eric N.↗

A general purpose subroutine for fast fourier transform on a distributed memory parallel machine

One issue which is central in developing a general purpose Fast Fourier Transform (FFT) subroutine on a distributed memory parallel machine is the data distribution. It is possible that different users would like to use the FFT routine with different data distributions. Thus, there is a need to design FFT schemes on distributed memory parallel machines which can support a variety of data distributions. An FFT implementation on a distributed memory parallel machine which works for a number of data distributions commonly encountered in scientific applications is presented. The problem of rearranging the data after computing the FFT is also addressed. The performance of the implementation on a distributed memory parallel machine Intel iPSC/860 is evaluated.

Dubey, A.↗

High-speed data duplication/data distribution: An adjunct to the mass storage equation

The term 'mass storage' invokes the image of large on-site disk and tape farms which contain huge quantities of low- to medium-access data. Although the cost of such bulk storage is recognized, the cost of the bulk distribution of this data rarely is given much attention. Mass data distribution becomes an even more acute problem if the bulk data is part of a national or international system. If the bulk data distribution is to travel from one large data center to another large data center then fiber-optic cables or the use of satellite channels is feasible. However, if the distribution must be disseminated from a central site to a number of much smaller, and, perhaps varying sites, then cost prohibits the use of fiber-optic cable or satellite communication. Given these cost constraints much of the bulk distribution of data will continue to be disseminated via inexpensive magnetic tape using the various next day postal service options. For non-transmitted bulk data, our working hypotheses are that the desired duplication efficiency of the total bulk data should be established before selecting any particular data duplication system; and, that the data duplication algorithm should be determined before any bulk data duplication method is selected.

Howard, Kevin↗

OhioView: Distribution of Remote Sensing Data Across Geographically Distributed Environments

Various issues associated with the distribution of remote sensing data across geographically distributed environments are presented in viewgraph form. Specific topics include: 1) NASA education program background; 2) High level architectures, technologies and applications; 3) LeRC internal architecture and role; 4) Potential GIBN interconnect; 5) Potential areas of network investigation and research; 6) Draft of OhioView data model; and 7) the LeRC strategy and roadmap.

Ramos, Calvin T.↗

Deliverable D12 – Distributed Wind Data Catalog Development Guide and Instruction Manual

Pacific Northwest National Laboratory and Technical University of Denmark completed this deliverable as part of Work Package 2: Data Information Catalog for Distributed Wind Research (WP2) for the International Energy Agency Wind Technology Collaboration Programme Task 41: Enabling Wind to Contribute to a Distributed Energy Future (IEA Wind Task 41). As the final deliverable for WP2, Deliverable D12 includes a data instruction guide for the IEA Wind Task 41 distributed wind data catalog. As such, this document includes: a step-by-step explanation of how the IEA Wind Task 41 data catalog was created, how it was populated, and how to use it; future options for the IEA Wind Task 41 data catalog, and a summary with recommendations for future work.

17 WIND ENERGY↗

A data and information system for processing, archival, and distribution of data for global change research

Work on this project was focused on information management techniques for Marshall Space Flight Center's EOSDIS Version 0 Distributed Active Archive Center (DAAC). The centerpiece of this effort has been participation in EOSDIS catalog interoperability research, the result of which is a distributed Information Management System (IMS) allowing the user to query the inventories of all the DAAC's from a single user interface. UAH has provided the MSFC DAAC database server for the distributed IMS, and has contributed to definition and development of the browse image display capabilities in the system's user interface. Another important area of research has been in generating value-based metadata through data mining. In addition, information management applications for local inventory and archive management, and for tracking data orders were provided.

Graves, Sara J.↗

Measuring the effects of distributed database models on transaction availability measures

Data distribution, data replication, and system reliability are key factors in determining the availability measures for transactions in distributed database systems. In order to simplify the evaluation of these measures, database designers and researchers tend to make unrealistic assumptions about these factors. Here, the effect of such assumptions on the computational complexity and accuracy of such evaluations is investigated. A database system is represented with five parameters related to the above factors. Probabilistic analysis is employed to evaluate the availability of read-one and read-write transactions. Both the read-one/write-all and the majority-read/majority-write replication control policies are considered. It is concluded that transaction availability is more sensitive to variations in degrees of replication, less sensitive to data distribution, and insensitive to reliability variations in a heterogeneous system. The computational complexity of the evaluations is found to be mainly determined by the chosen distributed database model, while the accuracy of the results are not so much dependent on the models.

Mukkamala, Ravi↗