Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Research in Structures and Dynamics, 1984

A symposium on advanced and trends in structures and dynamics was held to communicate new insights into physical behavior and to identify trends in the solution procedures for structures and dynamics problems. Pertinent areas of concern were (1) multiprocessors, parallel computation, and database management systems, (2) advances in finite element technology, (3) interactive computing and optimization, (4) mechanics of materials, (5) structural stability, (6) dynamic response of structures, and (7) advanced computer applications.

Robert J Hayduk↗

Maximum-Likelihood Decoder on a Hypercube Multiprocessor

Efficient parallel processing used to implement complex decoders. Hypercube multiprocessor connection scheme practical to decode long convolutional codes with efficient use of hardware. Hypercube design reduces both communication time among processors and space needed for interconnection. Decoding concept applicable to concurrent processing of digital signals using convolutional codes for error correction.

Pollara, F.↗

Iterative finite element solver on transputer networks

The parallelism inherent in the Conjugate Gradient method is described. The initial results of a parallel implementation on a network of twelve transputers are discussed. The high efficiencies obtained indicate that significant speedup can be obtained with larger transputer arrays if communication overhead can be kept low. To this end, a method of communication that allows large, dynamically reconfigurable transputer arrays to exchange data in log sub 4 N steps for N processors is suggested.

Danial, Albert↗

A Parallel Rendering Algorithm for MIMD Architectures

Applications such as animation and scientific visualization demand high performance rendering of complex three dimensional scenes. To deliver the necessary rendering rates, highly parallel hardware architectures are required. The challenge is then to design algorithms and software which effectively use the hardware parallelism. A rendering algorithm targeted to distributed memory MIMD architectures is described. For maximum performance, the algorithm exploits both object-level and pixel-level parallelism. The behavior of the algorithm is examined both analytically and experimentally. Its performance for large numbers of processors is found to be limited primarily by communication overheads. An experimental implementation for the Intel iPSC/860 shows increasing performance from 1 to 128 processors across a wide range of scene complexities. It is shown that minimal modifications to the algorithm will adapt it for use on shared memory architectures as well.

Crockett, Thomas W.↗

Acquisition and tracking performance measurements for a high speed area array detector system

A proof-of-concept (POC) demonstration system has been developed which demonstrates acquisition, tracking and point-ahead angle sensing for a space optical communications terminal utilizing a single high speed area array detector. The detector is the 128 x 128 pixel Kodak HS-40 photodiode array. It has 64 parallel readout channels and can operate at frames rates up to 40,000 frames/sec with rms readout noise of 20 photoelectrons. A windowing scheme and special purpose digital signal processing electronics are employed to implement acquisition and tracking algorithms. The system operates at greater than 1 kHz sample (frame) rates. Acquisition can be performed in as little as 30 milliseconds with less than 1 picowatt of 0.85 micron beacon power on the detector. At the same power level, the rms tracking accuracy is approximately 1/16 pixel. Results of system analysis and measurements using the POC system are presented.

Short, R. C.↗

Dynamic Load Balancing for Adaptive Meshes using Symmetric Broadcast Networks

Many scientific applications involve grids that lack a uniform underlying structure. These applications are often dynamic in the sense that the grid structure significantly changes between successive phases of execution. In parallel computing environments, mesh adaptation of grids through selective refinement/coarsening has proven to be an effective approach. However, achieving load balance while minimizing inter-processor communication and redistribution costs is a difficult problem. Traditional dynamic load balancers are mostly inadequate because they lack a global view across processors. In this paper, we compare a novel load balancer that utilizes symmetric broadcast networks (SBN) to a successful global load balancing environment (PLUM) created to handle adaptive unstructured applications. Our experimental results on the IBM SP2 demonstrate that performance of the proposed SBN load balancer is comparable to results achieved under PLUM.

Das, Sajal K.↗

Advanced-to-Revolutionary Space Technology Options - The Responsibly Imaginable

Paper summarizes a spectrum of low TRL, high risk technologies and systems approaches which could massively change the cost and safety of space exploration/exploitation/industrialization. These technologies and approaches could be studied in a triage fashion, the method of evaluation wherein several prospective solutions are investigated in parallel to address the innate risk of each, with resources concentrated on the more successful as more is learned. Technology areas addressed include Fabrication, Materials, Energetics, Communications, Propulsion, Radiation Protection, ISRU and LEO access. Overall and conceptually it should be possible with serious research to enable human space exploration beyond LEO both safe and affordable with a design process having sizable positive margins. Revolutionary goals require, generally, revolutionary technologies. By far, Revolutionary Energetics is the most important, has the most leverage, of any advanced technology for space exploration applications.

Bushnell, Dennis M.↗

Connector system for photovoltaic array

A photovoltaic assembly comprising; (a) at least two photovoltaic components that are adjacent to each other in a first direction, each photovoltaic component comprising (i) a partial recess in communication with the partial recess in an adjacent photovoltaic component and (ii) one or more connector receptors aligned in a second direction which is non-parallel to the first direction; (b) a connector located at feast partially in the partial recess of the photovoltaic component and at least partially in the partial recess of the adjacent photovoltaic component so that the connector connects the photovoltaic component to the adjacent photovoltaic component, the connector comprising: (i) a flexible housing having a first end and a second end; (ii) one or more connection ports at the first end; (iii) one or more connection ports at the second end; and (iv) one more flexible electrical conductors that extend from the one or more connection ports at the first end to the one or more connection ports at the second end; wherein the connector is flexible so that the first end and the second end are movable relative to each other in a plane, out of the plane, or both; wherein the one or more connection ports at the first end and the one or more connection ports at the second end form a connection with the one or more connector receptors of the photovoltaic component and the adjacent photovoltaic component so that the connector electrically connects the photovoltaic component to the adjacent photovoltaic component.

14 SOLAR ENERGY↗

Compiling global name-space parallel loops for distributed execution

Distributed memory machines do not provide hardware support for a global address space. Thus programmers are forced to partition the data across the memories of the architecture and use explicit message passing to communicate data between processors. The compiler support required to allow programmers to express their algorithms using a global name-space is examined. A general method is presented for analysis of a high level source program and its translation into a set of independently executing tasks communicating via messages. If the compiler has enough information, this translation can be carried out at compile time. Otherwise, run-time code is generated to implement the required data movement. The analysis required in both situations is described and the performance of the generated code on the Intel iPSC/2 is presented.

Koelbel, Charles↗

Stable parallel training of Wasserstein conditional generative adversarial neural networks

In this work, we propose a stable, parallel approach to train Wasserstein conditional generative adversarial neural networks (W-CGANs) under the constraint of a fixed computational budget. Differently from previous distributed GANs training techniques, our approach avoids inter-process communications, reduces the risk of mode collapse and enhances scalability by using multiple generators, each one of them concurrently trained on a single data label. The use of the Wasserstein metric also reduces the risk of cycling by stabilizing the training of each generator. We illustrate the approach on the CIFAR10, CIFAR100, and ImageNet1k datasets, three standard benchmark image datasets, maintaining the original resolution of the images for each dataset. Performance is assessed in terms of scalability and final accuracy within a limited fixed computational time and computational resources. To measure accuracy, we use the inception score, the Fréchet inception distance, and image quality. An improvement in inception score and Fréchet inception distance is shown in comparison to previous results obtained by performing the parallel approach on deep convolutional conditional generative adversarial neural networks as well as an improvement of image quality of the new images created by the GANs approach. Weak scaling is attained on both datasets using up to 2000 NVIDIA V100 GPUs on the OLCF supercomputer Summit.

97 MATHEMATICS AND COMPUTING↗

Cyber-Power Co-Simulation for End-to-End Synchrophasor Network Analysis and Applications

The resiliency, reliability and security of the next generation cyber-power smart grid depend upon efficiently leveraging advanced communication and computing technologies. Also, developing real-time data-driven applications is critical to enable wide-area monitoring and control of the cyber-power grid given high-resolution data from Phasor Measurement Units (PMUs). North American Synchrophasor Initiative Network (NASPlnet) provides guidance for PMU data exchanges. With the advancement in networking and grid operation, it is necessary to evaluate the performance of different data flow architectures suggested by NASPInet and analyze the impact on applications. Therefore, we need a cyber-power co-simulation framework that supports very large-scale co-simulation capable of running in parallel, high-performance computing platforms and capturing real-life network behavior. This work presents an end-to-end automated and user-driven cyber-power co-simulation using NS3 to model communication networks, GridPACK to model the power grid, and HELICS as a co-simulation engine. Comparative analysis of latency in synchrophasor networks and a performance evaluation of a power system stabilizer application utilizing PMU data in an IEEE 39 bus test system is presented using this cosimulation testbed.

Mustafa, Hussain M.↗

Root-Raised Cosine Filter Implementation That Uses Canonical Signed Digits for High-Speed Digital Filter Applications

NASA Lewis Research Center's Space Communications Division has been investigating high-speed digital filters that can operate at a higher speed than those in current use for a digital modulator and demodulator (modem). Using the Canonical Signed Digits (CSD) number representation for filter coefficients is a very effective way to increase the filter's speed while reducing complexity in the digital filter hardware design. This approach is a good alternative to using an expensive parallel-processing design technique or custom, application-specific integrated circuits. Such integrated circuits may not be suitable for applications that require filter speeds faster than what application-specific integrated circuits digital signal processors can offer for a dedicated channel. When a communication channel is a dedicated, multiplication process--a costly, time-consuming process--it can be greatly simplified by a replacement of the filter coefficients with CSD numbers. A computer code written with the MATLAB software package runs the program and generates CSD-represented filter coefficients that are based on minimizing minimum mean square errors. Also, the Alta Group of Cadence's Signal Processing Workstation is used to simulate and analyze the CSD filter responses. The impulse response of the root-raised cosine filter that is used as a base model is defined. From this filter, a set of coefficients is sampled and stored in a file. For the all coefficients, the optimal CSD number for each coefficient is searched on the basis of the minimum-mean-square-errors criterion. Because the distribution of CSD numbers is not uniform, quantization errors tend to be bigger for coefficients greater than 1/2. To offset errors that occur in a region of coefficients between 1/2 to 1 and to better represent fractions with CSD numbers, an extra nonzero digit is allowed for any coefficients exceeding 1/2. This will greatly improve frequency response as well as intersymbol interference at the receiver. The frequency response of a set of collected CSD-represented filter coefficients was compared with the same filter that was conventionally implemented. Analyses show CSD-implemented filters perform as well as conventional filters. Comparison of eye diagrams and bit-error-rate curves between CSD filters and traditionally implemented filters are almost indistinguishable. However, filter complexity was reduced from almost 3.5 to 1 for CSD filters. Complete computer simulation results are available. In the near future, work will focus on building actual working digital filter hardware in a field programmable gate array (FPGA).

Kim, Heechul↗

PaRSEC: Scalability, flexibility, and hybrid architecture support for task-based applications in ECP

This paper highlights the most significant enhancements made to PaRSEC, a scalable task-based runtime system designed for hybrid machines, during the Exascale Computing Project (ECP). The enhancements focus on expanding the capabilities of PaRSEC to address the evolving landscape of parallel computing. Notable achievements include the integration of support for three major types of accelerators (NVIDIA, AMD, and Intel GPUs), the refinement and increased flexibility of the communication subsystem, and the introduction of new programming interfaces tailored for irregular applications. Additionally, the project resulted in the development of powerful debugging and performance analysis tools aimed at assisting users in understanding and optimizing their applications. We present a comprehensive demonstration of these advancements through a series of benchmarks and applications within ECP and beyond, thereby showcasing the enhanced capabilities of PaRSEC across the diverse architectures within the ECP, providing valuable insights into the runtime system’s adaptability and performance across varied computing environments.

Bouteiller, Aurelien↗

Optimizing Distributed Training on Frontier for Large Language Models

Large language models (LLMs) have demonstrated remarkable success as foundational models, benefiting various downstream applications through fine-tuning. Loss scaling studies have demonstrated the superior performance of larger LLMs compared to their smaller counterparts. Nevertheless, training LLMs with billions of parameters poses significant challenges and requires considerable computational resources. For example, training a one trillion parameter GPT-style model on 20 trillion tokens requires a staggering 120 million exaflops. This research explores efficient distributed training strategies to extract this computation from Frontier, the world's first exascale supercomputer. We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning. For 22 Billion, 175 Billion, and 1 Trillion parameters, we achieved GPU throughputs of 38.38%, 36.14%, and 31.96%, respectively. For the training of the 175 Billion parameter model and the 1 Trillion parameter model, we achieved 100% weak scaling efficiency on 1024 and 3072 Mi250X GPUs, respectively. We also achieved strong scaling efficiencies of 89% and 87% for these two models. We trained these models only tens of iterations instead of training till completion.

Yin, Junqi↗

Eight microprocessor-based instrument data systems in the Galileo Orbiter spacecraft

Instrument data systems consist of a microprocessor, 3K bytes of Read Only Memory and 3K bytes of Random Access Memory. It interfaces with the spacecraft data bus through an isolated user interface with a direct memory access bus adaptor, and/or parallel data from instrument devices such as registers, buffers, analog to digital converters, multiplexers, and solid state sensors. These data systems support the spacecraft hardware and software communication protocol, decode and process instrument commands, generate continuous instrument operating modes, control the instrument mechanisms, acquire, process, format, and output instrument science data.

Barry, R. C.↗

Real-Time Reed-Solomon Decoder

RS decoder uses dedicated hardware and data pipelining for high-speed operation. Parallel processing techniques provide equivalent of over one billion operations per second at one step in decoding. Decoder finds commercial application in data encoding/decoding, telemetry, and radio communications.

Lahmeyer, C. R.↗

Initial operating capability for the hypercluster parallel-processing test bed

The NASA Lewis Research Center is investigating the benefits of parallel processing to applications in computational fluid and structural mechanics. To aid this investigation, NASA Lewis is developing the Hypercluster, a multi-architecture, parallel-processing test bed. The initial operating capability (IOC) being developed for the Hypercluster is described. The IOC will provide a user with a programming/operating environment that is interactive, responsive, and easy to use. The IOC effort includes the development of the Hypercluster Operating System (HYCLOPS). HYCLOPS runs in conjunction with a vendor-supplied disk operating system on a Front-End Processor (FEP) to provide interactive, run-time operations such as program loading, execution, memory editing, and data retrieval. Run-time libraries, that augment the FEP FORTRAN libraries, are being developed to support parallel and vector processing on the Hypercluster. Special utilities are being provided to enable passage of information about application programs and their mapping to the operating system. Communications between the FEP and the Hypercluster are being handled by dedicated processors, each running a Message-Passing Kernel, (MPK). A shared-memory interface allows rapid data exchange between HYCLOPS and the communications processors. Input/output handlers are built into the HYCLOPS-MPK interface, eliminating the need for the user to supply separate I/O support programs on the FEP.

Cole, Gary L.↗

Parallel grid generation algorithm for distributed memory computers

A parallel grid-generation algorithm and its implementation on the Intel iPSC/860 computer are described. The grid-generation scheme is based on an algebraic formulation of homotopic relations. Methods for utilizing the inherent parallelism of the grid-generation scheme are described, and implementation of multiple levELs of parallelism on multiple instruction multiple data machines are indicated. The algorithm is capable of providing near orthogonality and spacing control at solid boundaries while requiring minimal interprocessor communications. Results obtained on the Intel hypercube for a blended wing-body configuration are used to demonstrate the effectiveness of the algorithm. Fortran implementations bAsed on the native programming model of the iPSC/860 computer and the Express system of software tools are reported. Computational gains in execution time speed-up ratios are given.

Moitra, Stuti↗