Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Modulation and coding for satellite and space communications

Several modulation and coding advances supported by NASA are summarized. To support long-constraint-length convolutional code, a VLSI maximum-likelihood decoder, utilizing parallel processing techniques, which is being developed to decode convolutional codes of constraint length 15 and a code rate as low as 1/6 is discussed. A VLSI high-speed 8-b Reed-Solomon decoder which is being developed for advanced tracking and data relay satellite (ATDRS) applications is discussed. A 300-Mb/s modem with continuous phase modulation (CPM) and codings which is being developed for ATDRS is discussed. Trellis-coded modulation (TCM) techniques are discussed for satellite-based mobile communication applications.

Yuen, Joseph H.↗

On the utility of threads for data parallel programming

Threads provide a useful programming model for asynchronous behavior because of their ability to encapsulate units of work that can then be scheduled for execution at runtime, based on the dynamic state of a system. Recently, the threaded model has been applied to the domain of data parallel scientific codes, and initial reports indicate that the threaded model can produce performance gains over non-threaded approaches, primarily through the use of overlapping useful computation with communication latency. However, overlapping computation with communication is possible without the benefit of threads if the communication system supports asynchronous primitives, and this comparison has not been made in previous papers. This paper provides a critical look at the utility of lightweight threads as applied to data parallel scientific programming.

Fahringer, Thomas↗

A scalable exponential-DG approach for nonlinear conservation laws: With application to Burger and Euler equations

In this work, we propose an Exponential DG framework for partial differential equations. We decompose 7 governing equations into linear and nonlinear parts to which we apply the discontinuous Galerkin 8 (DG) spatial discretization. In particular, we construct the linear part using Jacobian that effectively 9 capture stiff characteristics in the system. The former is integrated analytically, whereas the latter 10 is approximated. This approach i) is stable with a large Courant number (Cr > 1); ii) supports 11 high-order solutions both in time and space; iii) is computationally favorable compared to IMEX 12 DG methods with no preconditioner; iv) becomes comparable to explicit RKDG methods on uniform 13 mesh and beneficial on non-uniform grid for Euler equations; v) is scalable in a modern massively 14 parallel computing architecture due to its explicit nature of exponential time integrators and com15 pact communication stencil of DG method. Numerical results demonstrate the performance of our 16 proposed methods through various examples. We also discuss the stability and convergence analysis 17 for our exponential DG scheme in the context of Burgers equation.

42 ENGINEERING↗

Record acceleration of the two-dimensional Ising model using a high-performance wafer-scale engine

The versatility and wide-ranging applicability of the Ising model, originally introduced to study phase transitions in magnetic materials, have made it a cornerstone in statistical physics and a valuable tool for evaluating the performance of emerging computer hardware. Here, we present a novel implementation of the two-dimensional Ising model on Cerebras Wafer-Scale Engine (WSE) – a revolutionary processor that is opening new frontiers in computing. In our deployment of the checkerboard algorithm, we optimized the Ising model to take advantage of the unique WSE architecture. Specifically, we employed a compressed bit representation storing 16 spins on each int16 word, and efficiently distributed the spins over the processing units enabling seamless weak scaling and limiting communications to only immediate neighboring units. Our implementation can handle up to 754 simulations in parallel, achieving an aggregate of over 61.8 trillion flip attempts per second for Ising models with up to 200 million spins. This represents a gain of up to 148 times over previously reported single-devices with a highly optimized implementation on NVIDIA V100 and up to 88 times in productivity compared to NVIDIA H100. Our findings highlight the significant potential of the WSE in scientific computing, particularly in the field of materials modeling.

Ising model↗

Simple circuit produces high-speed, fixed duration pulses

Circuit generates an output pulse of fixed width from a variable width input pulse. The circuit consists of a tunnel diode in parallel with an inductance driven by a constant current generator. It is used for pulsed communication equipment design.

Garrahan, N. M.↗

Multiscale Simulations of Magnetic Island Coalescence

We describe a new interactive parallel Adaptive Mesh Refinement (AMR) framework written in the Python programming language. This new framework, PyAMR, hides the details of parallel AMR data structures and algorithms (e.g., domain decomposition, grid partition, and inter-process communication), allowing the user to focus on the development of algorithms for advancing the solution of a systems of partial differential equations on a single uniform mesh. We demonstrate the use of PyAMR by simulating the pairwise coalescence of magnetic islands using the resistive Hall MHD equations. Techniques for coupling different physics models on different levels of the AMR grid hierarchy are discussed.

Dorelli, John C.↗

The role of optimization in structural model refinement

To evaluate the role that optimization can play in structural model refinement, it is necessary to examine the existing environment for the structural design/structural modification process. The traditional approach to design, analysis, and modification is illustrated. Typically, a cyclical path is followed in evaluating and refining a structural system, with parallel paths existing between the real system and the analytical model of the system. The major failing of the existing approach is the rather weak link of communication between the cycle for the real system and the cycle for the analytical model. Only at the expense of much human effort can data sharing and comparative evaluation be enhanced for the two parallel cycles. Much of the difficulty can be traced to the lack of a user-friendly, rapidly reconfigurable engineering software environment for facilitating data and information exchange. Until this type of software environment becomes readily available to the majority of the engineering community, the role of optimization will not be able to reach its full potential and engineering productivity will continue to suffer. A key issue in current engineering design, analysis, and test is the definition and development of an integrated engineering software support capability. The data and solution flow for this type of integrated engineering analysis/refinement system is shown.

Lehman, L. L.↗

Distributed Outage Detection in Power Distribution Networks

Real time topology knowledge is essential for situational awareness of power distribution networks. Line outages change the topology of a distribution network. Hence, outage detection is an important task. Most of the existing outage detection algorithms are centralized, in which sensors communicate their data to a control center which performs outage detection using the received data. However, with the increasing size of the distribution network and with different areas of the network being monitored by different operators, communication is a bottleneck and scalability is a major concern. To address these issues, we propose a novel outage detection algorithm using a divide and conquer approach. First, we divide a distribution network into sub-networks, such that outage detection can be run in parallel in each sub-network independently ensuring scalability to large networks. Further, to reduce the latency, bandwidth and attenuation challenges associated with communications in a large network, we divide each sub-network into multiple control areas which communicate only with their neighbors. We employ a distributed iterative load estimation across the control areas of each sub-network and then use the load estimate for local outage detection in each control area. Here, the performance of our algorithm is evaluated for multiple feeder models and compared against traditional centralized outage detection algorithms.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Spatiotemporal parallelization of an analytical heat conduction model for additive manufacturing via a hybrid OpenMP + MPI approach

The ability to do thermal simulations for entire additive manufacturing builds is a key computational problem facing the additive manufacturing community; however, complex numerical models considering multiple physical phenomena currently do not have the capacity for simulations at this scale. To this end, conduction only analytic models offer a viable approach due to the massive drop in computational expense. In this work, we extend an existing implementation which uses a governing equation which can be evaluated at any point in space and time. This implementation already utilizes OpenMP with a spatial decompositions scheme stemming from a melt pool tracking algorithm. Furthermore, we then combine this with a parallel in time (PinT) approach to make the problem highly parallelizable. The new scheme, which uses MPI for internode communication and OpenMP for intranode communication, is shown to scale very well across multiple computational nodes. This approach results in the ability to simulate the 3D solidification conditions for entire layers of additively manufactured parts in minutes making part scale thermal simulations more practical.

36 MATERIALS SCIENCE↗

Distributed-Memory Parallel JointNMF

Joint Nonnegative Matrix Factorization (JointNMF) is a hybrid method for mining information from datasets that contain both feature and connection information. We propose distributed-memory parallelizations of three algorithms for solving the JointNMF problem based on Alternating Nonnegative Least Squares, Projected Gradient Descent, and Projected Gauss-Newton. We extend well-known communication-avoiding algorithms using a single processor grid case to our coupled case on two processor grids. We demonstrate the scalability of the algorithms on up to 960 cores (40 nodes) with 60% parallel efficiency. The more sophisticated Alternating Nonnegative Least Squares (ANLS) and Gauss-Newton variants outperform the first-order gradient descent method in reducing the objective on large-scale problems. We perform a topic modelling task on a large corpus of academic papers that consists of over 37 million paper abstracts and nearly a billion citation relationships, demonstrating the utility and scalability of the methods.

Eswar, Srinivas↗

Parallel computation of manipulator inverse dynamics

In this article, parallel computation of manipulator inverse dynamics is investigated. A hierarchical graph-based mapping approach is devised to analyze the inherent parallelism in the Newton-Euler formulation at several computational levels, and to derive the features of an abstract architecture for exploitation of parallelism. At each level, a parallel algorithm represents the application of a parallel model of computation that transforms the computation into a graph whose structure defines the features of an abstract architecture, i.e., number of processors, communication structure, etc. Data-flow analysis is employed to derive the time lower bound in the computation as well as the sequencing of the abstract architecture. The features of the target architecture are defined by optimization of the abstract architecture to exploit maximum parallelism while minimizing architectural complexity. An architecture is designed and implemented that is capable of efficient exploitation of parallelism at several computational levels. The computation time of the Newton-Euler formulation for a 6-degree-of-freedom (dof) general manipulator is measured as 187 microsec. The increase in computation time for each additional dof is 23 microsec, which leads to a computation time of less than 500 microsec, even for a 12-dof redundant arm.

Fijany, Amir↗

Improved load distribution in parallel sparse Cholesky factorization

Compared to the customary column-oriented approaches, block-oriented, distributed-memory sparse Cholesky factorization benefits from an asymptotic reduction in interprocessor communication volume and an asymptotic increase in the amount of concurrency that is exposed in the problem. Unfortunately, block-oriented approaches (specifically, the block fan-out method) have suffered from poor balance of the computational load. As a result, achieved performance can be quite low. This paper investigates the reasons for this load imbalance and proposes simple block mapping heuristics that dramatically improve it. The result is a roughly 20% increase in realized parallel factorization performance, as demonstrated by performance results from an Intel Paragon system. We have achieved performance of nearly 3.2 billion floating point operations per second with this technique on a 196-node Paragon system.

Rothberg, Edward↗

Secure Network-Centric Aviation Communication (SNAC)

The existing National Airspace System (NAS) communications capabilities are largely unsecured, are not designed for efficient use of spectrum and collectively are not capable of servicing the future needs of the NAS with the inclusion of new operators in Unmanned Aviation Systems (UAS) or On Demand Mobility (ODM). SNAC will provide a ubiquitous secure, network-based communications architecture that will provide new service capabilities and allow for the migration of current communications to SNAC over time. The necessary change in communication technologies to digital domains will allow for the adoption of security mechanisms, sharing of link technologies, large increase in spectrum utilization, new forms of resilience and redundancy and the possibly of spectrum reuse. SNAC consists of a long term open architectural approach with increasingly capable designs used to steer research and development and enable operating capabilities that run in parallel with current NAS systems.

Communications↗

Spaceborne Processor Array

A Spaceborne Processor Array in Multifunctional Structure (SPAMS) can lower the total mass of the electronic and structural overhead of spacecraft, resulting in reduced launch costs, while increasing the science return through dynamic onboard computing. SPAMS integrates the multifunctional structure (MFS) and the Gilgamesh Memory, Intelligence, and Network Device (MIND) multi-core in-memory computer architecture into a single-system super-architecture. This transforms every inch of a spacecraft into a sharable, interconnected, smart computing element to increase computing performance while simultaneously reducing mass. The MIND in-memory architecture provides a foundation for high-performance, low-power, and fault-tolerant computing. The MIND chip has an internal structure that includes memory, processing, and communication functionality. The Gilgamesh is a scalable system comprising multiple MIND chips interconnected to operate as a single, tightly coupled, parallel computer. The array of MIND components shares a global, virtual name space for program variables and tasks that are allocated at run time to the distributed physical memory and processing resources. Individual processor- memory nodes can be activated or powered down at run time to provide active power management and to configure around faults. A SPAMS system is comprised of a distributed Gilgamesh array built into MFS, interfaces into instrument and communication subsystems, a mass storage interface, and a radiation-hardened flight computer.

Chow, Edward T.↗

Supercomputing systems - A projection to 2000

Advances in computer architecture, computer science, computational methods, and constituent technologies are expected to lead to significant advances in the performance of scientific supercomputing system capabilities over the next decade. By the year 2000, single 1-in-sq dies are projected to incorporate four processors, each of which would be operating faster than 750 million instructions per second (MIPS) for a total on-chip processing performance in excess of 2000 MIPS. Scalable parallel processors can be expected to contain thousands of such multiple processor chips. In general, semiconductor performance advances appear to change about one order of magnitude every five years. Rotating magnetic memory and communications technology are not advancing as rapidly, with the result that the allocation of functions within the system configurations fo future supercomputer systems will require important changes. Availability of massively parallel heterogeneous processing capabilities should be a catalyst leading to new approaches for applications.

Lundstrom, S. F.↗

A GPU ‐Accelerated 3D Unstructured Mesh Based Particle Tracking Code for Multi‐Species Impurity Transport Simulation in Fusion Tokamaks

ABSTRACT This paper presents the multi‐species global impurity transport capability developed in a GPU‐accelerated fully 3D unstructured mesh‐based code, GITRm, to simultaneously track multiple impurity species and handle interactions of these impurities with mixed‐material surfaces. Different computational approaches to model particle‐surface interaction or surface response have been developed and compared. Sheath electric field is taken into account by employing a fast distance‐to‐boundary calculation, which is carried out in parallel on distributed or partitioned meshes on multiple GPUs without the need for any inter‐process communication during the simulation. Several example cases, including two for the DIII‐D tokamak, that is, one with the SAS‐V divertor and the other with the collector probes, are used to demonstrate the utility of the current multi‐species capability. For the DIII‐D probe case, the capability of GITRm to resolve the spatial distribution of particles in localized regions, such as diagnostic probes, within non‐axisymmetric tokamak geometries is demonstrated. These simulations involve up to 320 million particles and utilize up to 48 GPUs.

Nath, Dhyanjyoti D. [Scientific Computation Resear↗

Cyber Security Analysis for Nuclear Reactor Control Systems (Final Technical Report)

This project investigated the cyber-security impacts of moving from an all analog, point-to-point, instrumentation and control (I&C) system to a digital I&C system based on Modbus and a shared communication medium. A formalism called a hybrid attack graph was expanded to support the nuclear research reactor system. The hybrid attack graph allows one to check a system for vulnerabilities, in this case cyber-security vulnerabilities, and to document the attack vectors (scenarios) causing those vulnerabilities. In parallel, a simulation of the system was developed to model both the physical reactor parameters and operations, as well as the network interconnects and communications. This simulation platform was modeled on the nuclear research reactor located at Washington State University. The simulation platform provided a sandbox to evaluate and quantify the impact of identified and proposed vulnerabilities in the system and to determine the effectiveness of countermeasures at stopping these attacks. The simulation and hybrid attack graph tools were integrated to provide a streamlined process of generating attack scenarios, playing those scenarios out in the simulation, and then analyzing the results to correlate system state to states in the hybrid attack graph. This process was used to (1) quantify the impact of attack scenarios and (2) to determine if the system moved through the hybrid attack graph as anticipated. The hybrid attack graph tool was extended and customized to produce a tool to automatically identify critical assets (CAs) and critical digital assets (CDAs) as defined by NRC Regulatory Guide 5.71. This tool was verified using the nuclear research reactor at Washington State University. Finally, a series of educational modules covering the findings of the different aspects of this research have been created.

97 MATHEMATICS AND COMPUTING↗

Wireless Temperature-Monitoring System

A relatively inexpensive instrumentation system that includes units that are connected to thermocouples and that are parts of a radio-communication network has been developed to enable monitoring of temperatures at multiple locations. Because there is no need to string wires or cables for communication, the system is well suited for monitoring temperatures at remote locations and for applications in which frequent changes of monitored or monitoring locations are needed. The system can also be adapted to monitoring of slowly varying physical quantities, other than temperature, that can be transduced by solid-state electronic sensors. electronic sensors. The system comprises any number of transmitting units and a single receiving unit. Each transmitting unit includes connections for as many as four external thermocouples, a signal-conditioning module, a control module, and a radio-communication module. The signal-conditioning module acts as an interface between the thermocouples and the rest of the transmitting unit and includes a built-in solid ambient temperature sensor that is in addition to the external thermocouples. The control module is a system-on-chip embedded processor that includes analog-to-digital converters, serial and parallel data ports, and an interface for local connection to an analog meter that is used during installation to verify correct operation. The radio-communication module contains a commercial spread-spectrum transceiver that operates in the 900-MHz industrial, scientific, and medical (ISM) frequency band. This transceiver transmits data to the receiving unit at a rate of 19,200 baud. The receiving unit includes a transceiver like that of a transmitting unit, plus a control module that contains a system-on-chip processor that includes serial data port for output to a computer that runs monitoring and/or control software, a parallel data port for output to a printer, and a seven-segment light-emitting-diode display. Each transmitting unit is battery-powered and can operate for at least seven days continuously while reporting temperatures every half hour. The receiving unit is powered by a wall-mounted transformer source. The receiving unit responds to each transmitting unit and reports the readings of each of the four thermocouples and of the ambient-temperature sensor of the transmitting unit. The end-to-end accuracy of the system is plus or minus 0.2 C over the temperature range from 0 to 100 C. The radio-communication range between the receiving and transmitting units is approximately equal to 0.5 mile (approximately equal to 0.8 km).

Solano, Wanda↗