Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “communication and synchronization reducing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Unified Communication Optimization Strategies for Sparse Triangular Solver on CPU and GPU Clusters

This paper presents a unified communication optimization framework for sparse triangular solve (SpTRSV) algorithms on CPU and GPU clusters. The framework builds upon a 3D communication-avoiding (CA) layout of Px × Py × Pz processes that divides a sparse matrix into Pz submatrices, each handled by a Px × Py 2D grid with block-cyclic distribution. We propose three communication optimization strategies: First, a new 3D SpTRSV algorithm is developed, which trades the inter-grid communication and synchronization with replicated computation. This design requires only one inter-grid synchronization, and the inter-grid communication is efficiently implemented with sparse allreduce operations. Second, broadcast and reduction communication trees are used to reduce message latency of the intra-grid 2D communication on CPU clusters. Finally, we leverage GPU-initiated one-sided communication to implement the communication trees on GPU clusters. With these nested inter- and intra-grid communication optimization strategies, the proposed 3D SpTRSV algorithm can attain up to 3.45x speedups compared to the baseline 3D SpTRSV algorithm using up to 2048 Cori Haswell CPU cores. In addition, the proposed GPU 3D SpTRSV algorithm can achieve up to 6.5x speedups compared to the proposed CPU 3D SpTRSV algorithm with Pz up to 64. Finally it is remarkable that the proposed GPU 3D SpTRSV can scale to 256 GPUs using the Perlmutter system while the existing 2D SpTRSV algorithm can only scale up to 4 GPUs.

Sao, Piyush↗

Technologies for unattended network operations

The necessary network management functions for a telecommunications, navigation and information management (TNIM) system in the framework of an extension of the ISO model for communications network management are described. Various technologies that could substantially reduce the need for TNIM network management, automate manpower intensive functions, and deal with synchronization and control at interplanetary distances are presented. Specific technologies addressed include the use of the ISO Common Management Interface Protocol, distributed artificial intelligence for network synchronization and fault management, and fault-tolerant systems engineering.

Jaworski, Allan↗

Asynchronous interactive control systems

A class of interactive control systems is derived by generalizing interactive manipulator control systems. The general structural properties of such systems are discussed and an appropriate general software implementation is proposed. This is based on the fact that tasks of interactive control systems can be represented as a network of a finite set of actions which have specific operational characteristics and specific resource requirements, and which are of limited duration. This has enabled the decomposition of the overall control algorithm into a set of subalgorithms, called subcontrollers, which can operate simultaneously and asynchronously. Coordinate transformations of sensor feedback data and actuator set-points have enabled the further simplification of the subcontrollers and have reduced their conflicting resource requirements. The modules of the decomposed control system are implemented as parallel processes with disjoint memory space communicating only by I/O. The synchronization mechanisms for dynamic resource allocation among subcontrollers and other synchronization mechanisms are also discussed in this paper. Such a software organization is suitable for the general form of multiprocessing using computer networks with distributed storage.

Vuskovic, M. I.↗

Anticipating decoherence in quantum systems

Large-scale quantum technologies require coherence across distant nodes, necessitating indistinguishable quantum states. However, environmental disorder, including dephasing, spectral diffusion, and spin-bath interactions, undermines coherence. Using statistical methods, we uncover correlations in decoherence channels induced by slowly varying environments. Spectral diffusion serves as a representative demonstration case that can be extended to other remote, disordered systems such as spins in nitrogen-vacancy centers and quantum-dot spin qubits, as well as flux noise in superconducting qubits. In this work, we employ replica-theory-inspired trajectory analysis to reveal predictable temporal structures in decoherence dynamics, and validate these through an anticipatory systems framework with internal prediction of unseen spectral dynamics in multiple quantum systems, showing that this framework could, if implemented, reduce spectral shift by average factors of approximately 2 to 19, depending on emitter stability, thereby enabling enhanced coherence and multi-node synchronization for scalable quantum communication, computation, imaging, and sensing.

Maan, Pranshu [Purdue University]↗

Joint Synchronization Of Viterbi And Reed-Solomon Decoders

Synchronization times reduced to reduce loss of data. Scheme for decoding received doubly encoded binary-data signal provides for joint synchronization of two decoders. Applies to concatenated error-correcting channel coding communication system in which, at transmitter, data first encoded by interleaved Reed-Solomon code (block code), then by convolutional code.

Statman, Joseph I.↗

Integrated testing and verification system for research flight software design document

The NASA Langley Research Center is developing the MUST (Multipurpose User-oriented Software Technology) program to cut the cost of producing research flight software through a system of software support tools. The HAL/S language is the primary subject of the design. Boeing Computer Services Company (BCS) has designed an integrated verification and testing capability as part of MUST. Documentation, verification and test options are provided with special attention on real time, multiprocessing issues. The needs of the entire software production cycle have been considered, with effective management and reduced lifecycle costs as foremost goals. Capabilities have been included in the design for static detection of data flow anomalies involving communicating concurrent processes. Some types of ill formed process synchronization and deadlock also are detected statically.

Taylor, R. N.↗

Sample-Clock Phase-Control Feedback

To demodulate a communication signal, a receiver must recover and synchronize to the symbol timing of a received waveform. In a system that utilizes digital sampling, the fidelity of synchronization is limited by the time between the symbol boundary and closest sample time location. To reduce this error, one typically uses a sample clock in excess of the symbol rate in order to provide multiple samples per symbol, thereby lowering the error limit to a fraction of a symbol time. For systems with a large modulation bandwidth, the required sample clock rate is prohibitive due to current technological barriers and processing complexity. With precise control of the phase of the sample clock, one can sample the received signal at times arbitrarily close to the symbol boundary, thus obviating the need, from a synchronization perspective, for multiple samples per symbol. Sample-clock phase-control feedback was developed for use in the demodulation of an optical communication signal, where multi-GHz modulation bandwidths would require prohibitively large sample clock frequencies for rates in excess of the symbol rate. A custom mixedsignal (RF/digital) offset phase-locked loop circuit was developed to control the phase of the 6.4-GHz clock that samples the photon-counting detector output. The offset phase-locked loop is driven by a feedback mechanism that continuously corrects for variation in the symbol time due to motion between the transmitter and receiver as well as oscillator instability. This innovation will allow significant improvements in receiver throughput; for example, the throughput of a pulse-position modulation (PPM) with 16 slots can increase from 188 Mb/s to 1.5 Gb/s.

Quirk, Kevin J.↗

A Fully Decentralized Modulation Scheme for Modular Multilevel Converters

As the number of submodules rapidly increase, control architecture of Modular Multilevel Converters (MMCs) is transforming from centralized to distributed in order to address the challenges of tremendous computational burden and massive wiring. In this digest, we propose a modulation scheme for the MMCs with distributed controls, where each submodule is able to handle both controls and modulation in a decentralized manner. It enables the distributed controllers to synchronize in terms of fast-switching pulse width modulation (PWM). Aiming at this key technical challenge, this research develops a physics-informed PWM control strategy of mirroring the spontaneous synchronization feature of coupled oscillator’s behavior at switching frequency. Thanks to less reliance on wirings and communications, the proposed approach has the potential of i) improving reliability of MMC systems, and ii) reducing cost and space requirements for MMCs. The proposed approach has been preliminarily verified through a simulation tool for power electronics systems.

Lu, Minghui [BATTELLE (PACIFIC NW LAB)]↗

Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency

Synchronous SGD with data parallelism, the most popular parallelization strategy for CNN training, suffers from the expensive communication cost of averaging gradients among all workers. The iterative parameter updates of SGD cause frequent communications and it becomes the performance bottleneck. In this paper, we propose a lazy parameter update algorithm that adaptively adjusts the parameter update frequency to address the expensive communication cost issue. Our algorithm accumulates the gradients if the difference of the accumulated gradients and the latest gradients is sufficiently small. Here, the less frequent parameter updates reduce the per-iteration communication cost while maintaining the model accuracy. Our experimental results demonstrate that the lazy update method remarkably improves the scalability while maintaining the model accuracy. For ResNet50 training on ImageNet, the proposed algorithm achieves a significantly higher speedup (739.6 on 2048 Cori KNL nodes) as compared to the vanilla synchronous SGD (276.6) while the model accuracy is almost not affected (<0.2% difference).

97 MATHEMATICS AND COMPUTING↗

Energy–Performance Trade-offs in Privacy-Preserving Federated Learning on SmartNIC-Enabled HPC Systems

Federated learning (FL) is increasingly deployed on accelerator-rich high-performance computing (HPC) systems, yet the system-level energy cost of privacy-aware FL remains poorly understood, particularly across heterogeneous networking and server-placement options. We present a measurement-driven study of energy–performance trade-offs for FL on GH200-class nodes across three deployment configurations: CPU-Ethernet, CPU-InfiniBand (RDMA-capable), and a DPU-hosted FL server over InfiniBand using a BlueField-3 SmartNIC/DPU. Using NVIDIA FLARE (NVFLARE), we align node-level power telemetry with per-round timing extracted from NVFLARE logs to quantify time-to-solution (TTS), energy-to-solution (ETS), energy-delay product (EDP), and synchronization behavior for three transformer models (ALBERT, DistilBERT, BERT), trained with and without differential privacy (DP). We find that interconnect choice is the dominant driver of runtime and energy: host-managed InfiniBand consistently reduces communication overhead versus Ethernet, yielding lower TTS/ETS/EDP. In contrast, in our NVFLARE deployment, placing the FL server on the DPU does not consistently match CPU-InfiniBand performance and can be slower—especially for larger models—highlighting that server placement alone is not sufficient to guarantee end-to-end gains. Finally, under our fixed-round protocol, DP increases per-round cost and runtime variance; ETS increases largely in proportion to TTS because average node power remains relatively stable across configurations.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Characterization of a Photon Counting Test Bed for Space to Ground Optical Pulse Position Modulation Communications Links

The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) has developed a laboratory transmitter and receiver prototype of a space-to-ground optical communications link. The system is meant to emulate future deep space optical communication links, such as the first crewed flight of Orion, in which the transmitted laser is modulated using pulse position modulation and the receiver is capable of detecting single photons. The transmitter prototype consists of a software defined radio, a high extinction ratio electro-optic modulator system, and a 1550 nm laser. The receiver is a scalable concept and utilizes a single-pixel array of fiber coupled superconducting nanowire single photon detectors. The transmit and receive waveforms follow the Consultative Committee for Space Data Systems (CCSDS) Optical Communications Coding and Synchronization Standard. A software model of the optical transmitter and receiver has also been implemented to predict performance of the optical test bed. This paper describes the transmitter and receiver prototypes as well as the system test configuration. System level tests results are presented and shown to align with predictions from software simulations. The validated software model can be used to in the future to reduce the design cycle of optical communications systems.

software defined radio↗

Synchronizing Data-Bus Messages

Adapter allows communications among as many as 30 data processors without central bus controller. Adapter improves reliability of multiprocessor system by eliminating point of failure that causes entire system to fail. Scheme prevents data collisions and eliminates nonessential polling, thereby reducing power consumption.

Harris, L. H.↗

Smart and Intelligent Sensors

John C. Stennis Space Center (SSC) provides rocket engine propulsion testing for NASA's space programs. Since the development of the Space Shuttle, every Space Shuttle Main Engine (SSME) has undergone acceptance testing at SSC before going to Kennedy Space Center (KSC) for integration into the Space Shuttle. The SSME is a large cryogenic rocket engine that uses Liquid Hydrogen (LH2) as the fuel. As NASA moves to the new ARES V launch system, the main engines on the new vehicle, as well as the upper stage engine, are currently base lined to be cryogenic rocket engines that will also use LH2. The main rocket engines for the ARES V will be larger than the SSME, while the upper stage engine will be approximately half that size. As a result, significant quantities of hydrogen will be required during the development, testing, and operation of these rocket engines.Better approaches are needed to simplify sensor integration and help reduce life-cycle costs. 1.Smarter sensors. Sensor integration should be a matter of "plug-and-play" making sensors easier to add to a system. Sensors that implement new standards can help address this problem; for example, IEEE STD 1451.4 defines transducer electronic data sheet (TEDS) templates for commonly used sensors such as bridge elements and thermocouples. When a 1451.4 compliant smart sensor is connected to a system that can read the TEDS memory, all information needed to configure the data acquisition system can be uploaded. This reduces the amount of labor required and helps minimize configuration errors. 2.Intelligent sensors. Data received from a sensor be scaled, linearized; and converted to engineering units. Methods to reduce sensor processing overhead at the application node are needed. Smart sensors using low-cost microprocessors with integral data acquisition and communication support offer the means to add these capabilities. Once a processor is embedded, other features can be added; for example, intelligent sensors can make a health assessment to inform the data acquisition client when sensor performance is suspect. 3.Distributed sample synchronization. Networks of sensors require new ways for synchronizing samples. Standards that address the distributed timing problem (for example, IEEE STD 1588) provide the means to aggregate samples from many distributed smart sensors with sub-microsecond accuracy. 4. Reduction in interconnect. Alternative means are needed to reduce the frequent problems associated with cabling and connectors. Wireless technologies offer the promise of reducing interconnects and simultaneously making it easy to quickly add a sensor to a system.

Lansaw, John↗

Low-Cost Telemetry System for Small/Micro Satellites

A Software Defined Radio (SDR) concept uses a minimum amount of analog/radio frequency components to up/downconvert the RF signal to/from a digital format. Once in the digital domain, all other processing (filtering, modulation, demodulation, etc.) is done in software. The project will leverage existing designs and enhance capabilities in the commercial sector to provide a path to a radiation-hardened SDR transponder. The SDR transponder would incorporate baseline technologies dealing with improved Forward Error Correcting (FEC) codes to be deployed to all Near Earth Network (NEN) ground stations. By incorporating this FEC, at least a tenfold increase in data throughput can be achieved. A family of transponder products can be implemented using common platform architecture, allowing new products to be more quickly introduced into the market. Software can be reused across products, reducing software/hardware costs dramatically. New features and capabilities, such as encoding and decoding algorithms, filters, and bit synchronizers, can be added to the existing infrastructure without requiring major new capital expenditures, allowing implementation of advanced features in the communication systems. As new telecommunication technologies emerge, incorporating them into the SDR fabric will be easily accomplished with little or no requirements for new hardware. There are no preferred flight platforms for the SDR technology, so it can be used on any type of orbital or sub-orbital platform, all within a fully radiation hardened design.

Sims, William↗

Computational Architecture For Control Of Remote Manipulator

Synchronization done by hardware to reduce software overhead. Computing resources located at both master-arm node and slave-arm node. This architecture provides for effective control while reducing computational burden on host computer and reducing and balancing load on communication channel.

Szakaly, Zoltan F.↗

Definition Study for Space Shuttle Experiments Involving Large, Steerable Millimeter-Wave Antenna Arrays

The potential uses and techniques for the shuttle spacelab Millimeter Wave Large Aperture Antenna Experiment (MWLAE) are documented. Potential uses are identified: applications to radio astronomy, the sensing of atmospheric turbulence by its effect on water vapor line emissions, and the monitoring of oil spills by multifrequency radiometry. IF combining is preferable to RF combining with respect to signal to noise ratio for communications receiving antennas of the size proposed for MWLAE. A design approach using arrays of subapertures is proposed to reduce the number of phase shifters and mixers for uses which require a filled aperture. Correlation radiometry and a scheme utilizing synchronous Dicke switches and IF combining are proposed as potential solutions.

Levis, C. A.↗

Development of a Crosslink Channel Simulator for Simulation of Formation Flying Satellite Systems

Multi-vehicle missions are an integral part of NASA s and other space agencies current and future business. These multi-vehicle missions generally involve collectively utilizing the array of instrumentation dispersed throughout the system of space vehicles, and communicating via crosslinks to achieve mission goals such as formation flying, autonomous operation, and collective data gathering. NASA s Goddard Space Flight Center (GSFC) is developing the Formation Flying Test Bed (FFTB) to provide hardware-in- the-loop simulation of these crosslink-based systems. The goal of the FFTB is to reduce mission risk, assist in mission planning and analysis, and provide a technology development platform that allows algorithms to be developed for mission hctions such as precision formation flying, synchronization, and inter-vehicle data synthesis. The FFTB will provide a medium in which the various crosslink transponders being used in multi-vehicle missions can be plugged in for development and test. An integral part of the FFTB is the Crosslink Channel Simulator (CCS),which is placed into the communications channel between the crosslinks under test, and is used to simulate on-orbit effects to the communications channel due to relative vehicle motion or antenna misalignment. The CCS is based on the Starlight software programmable platform developed at General Dynamics Decision Systems which provides the CCS with the ability to be modified on the fly to adapt to new crosslink formats or mission parameters.

Hart, Roger↗

Update on Development of SiC Multi-Chip Power Modules

Progress has been made in a continuing effort to develop multi-chip power modules (SiC MCPMs). This effort at an earlier stage was reported in 'SiC Multi-Chip Power Modules as Power-System Building Blocks' (LEW-18008-1), NASA Tech Briefs, Vol. 31, No. 2 (February 2007), page 28. The following recapitulation of information from the cited prior article is prerequisite to a meaningful summary of the progress made since then: 1) SiC MCPMs are, more specifically, electronic power-supply modules containing multiple silicon carbide power integrated-circuit chips and silicon-on-insulator (SOI) control integrated-circuit chips. SiC MCPMs are being developed as building blocks of advanced expandable, reconfigurable, fault-tolerant power-supply systems. Exploiting the ability of SiC semiconductor devices to operate at temperatures, breakdown voltages, and current densities significantly greater than those of conventional Si devices, the designs of SiC MCPMs and of systems comprising multiple SiC MCPMs are expected to afford a greater degree of miniaturization through stacking of modules with reduced requirements for heat sinking; 2) The stacked SiC MCPMs in a given system can be electrically connected in series, parallel, or a series/parallel combination to increase the overall power-handling capability of the system. In addition to power connections, the modules have communication connections. The SOI controllers in the modules communicate with each other as nodes of a decentralized control network, in which no single controller exerts overall command of the system. Control functions effected via the network include synchronization of switching of power devices and rapid reconfiguration of power connections to enable the power system to continue to supply power to a load in the event of failure of one of the modules; and, 3) In addition to serving as building blocks of reliable power-supply systems, SiC MCPMs could be augmented with external control circuitry to make them perform additional power-handling functions as needed for specific applications. Because identical SiC MCPM building blocks could be utilized in such a variety of ways, the cost and difficulty of designing new, highly reliable power systems would be reduced considerably. This concludes the information from the cited prior article. The main activity since the previously reported stage of development was the design, fabrication, and testing a 120- VDC-to-28-VDC modular power-converter system composed of eight SiC MCPMs in a 4 (parallel)-by-2 (series) matrix configuration, with normally-off controllable power switches. The SiC MCPM power modules include closed-loop control subsystems and are capable of operating at high power density or high temperature. The system was tested under various configurations, load conditions, load-transient conditions, and failure-recovery conditions. Planned future work includes refinement of the demonstrated modular system concept and development of a new converter hardware topology that would enable sharing of currents without the need for communication among modules. Toward these ends, it is also planned to develop a new converter control algorithm that would provide for improved sharing of current and power under all conditions, and to implement advanced packaging concepts that would enable operation at higher power density.

Lostetter, Alexander↗