Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

GPU Acceleration of Large-Scale Full-Frequency GW Calculations

Many-body perturbation theory is a powerful method to simulate electronic excitations in molecules and materials starting from the output of density functional theory calculations. By implementing the theory efficiently so as to run at scale on the latest leadership high-performance computing systems it is possible to extend the scope of GW calculations. Here, we present a GPU acceleration study of the full-frequency GW method as implemented in the WEST code. Excellent performance is achieved through the use of (i) optimized GPU libraries, e.g., cuFFT and cuBLAS, (ii) a hierarchical parallelization strategy that minimizes CPU-CPU, CPU-GPU, and GPU-GPU data transfer operations, (iii) nonblocking MPI communications that overlap with GPU computations, and (iv) mixed precision in selected portions of the code. A series of performance benchmarks has been carried out on leadership high-performance computing systems, showing a substantial speedup of the GPU-accelerated version of WEST with respect to its CPU version. Good strong and weak scaling is demonstrated using up to 25 920 GPUs. Finally, we showcase the capability of the GPU version of WEST for large-scale, full-frequency GW calculations of realistic systems, e.g., a nanostructure, an interface, and a defect, comprising up to 10 368 valence electrons.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ARM-IRL: Adaptive Resilience Metric Quantification Using Inverse Reinforcement Learning

The resilience of safety-critical systems is gaining importance due to the rise in cyber and physical threats, especially within critical infrastructure. Traditional static resilience metrics may not capture dynamic system states, leading to inaccurate assessments and ineffective responses to cyber threats. This work aims to develop a data-driven, adaptive method for resilience metric learning. We propose a data-driven approach using inverse reinforcement learning (IRL) to learn a single, adaptive resilience metric. The method infers a reward function from expert control actions. Unlike previous approaches using static weights or fuzzy logic, this work applies adversarial inverse reinforcement learning (AIRL), training a generator and discriminator in parallel to learn the reward structure and derive an optimal policy. The proposed approach is evaluated on multiple scenarios: optimal communication network rerouting, power distribution network reconfiguration, and cyber–physical restoration of critical loads using the IEEE 123-bus system. The adaptive, learned resilience metric enables faster critical load restoration in comparison to conventional RL approaches.

97 MATHEMATICS AND COMPUTING↗

Umbilical Deployment Device

The landing scheme for NASA's next-generation Mars rover will encompass a novel landing technique (see figure). The rover will be lowered from a rocket-powered descent stage and then placed onto the surface while hanging from three bridles. Communication between the rover and descent stage will be maintained through an electrical umbilical cable, which will be deployed in parallel with structural bridles. The -inch (13-mm) umbilical cable contains a Kevlar rope core, around which wires are wrapped to create a cable. This cable is helically coiled between two concentric truncated cones. It is deployed by pulling one end of the cable from the cone. A retractable mechanism maintains tension on the cable after deployment. A break-tie tethers the umbilical end attached to the rover even after the cable is cut after touchdown. This break-tie allows the descent stage to develop some velocity away from the rover prior to the cable releasing from the rover deck, then breaks away once the cable is fully extended. The descent stage pulls the cable up so that recontact is not made. The packaging and deployment technique can store a long length of cable in a relatively small volume while maintaining compliance with the minimum bend radius requirement for the cable being deployed. While the packaging technique could be implemented without the use of break-ties, they were needed in this design due to the vibratory environment and the retraction required by the cable. The break-ties used created a series of load-spikes in the deployment signature. The load spikes during the deployment of the initial three coils of umbilical showed no increase between the different temperature trials. The cold deployment did show an increased load requirement for cable extraction in the region where no break-ties were used. This increase in cable drag was superimposed on the loads required to rupture the last set of break-ties, and as such, these loads saw significant increase when compared to their ambient counterparts. While the loads showed spikes of high magnitude, they were of short duration. Because of this, neither the deployment of the rover, nor the motion of the descent stage, would be adversely affected. In addition, the umbilical was found to have a maximum of 1.2 percent chance for recontact with the ultra-high frequency antenna due to the large margin of safety built in.

Shafer, Michael W.↗

Systems and methods for rapid processing and storage of data

Systems and methods of building massively parallel computing systems using low power computing complexes in accordance with embodiments of the invention are disclosed. A massively parallel computing system in accordance with one embodiment of the invention includes at least one Solid State Blade configured to communicate via a high performance network fabric. In addition, each Solid State Blade includes a processor configured to communicate with a plurality of low power computing complexes interconnected by a router, and each low power computing complex includes at least one general processing core, an accelerator, an I/O interface, and cache memory and is configured to communicate with non-volatile solid state memory.

Stalzer, Mark A.↗

Range Information Systems Management (RISM) Phase 1 Report

RISM investigated alternative approaches, technologies, and communication network architectures to facilitate building the Spaceports and Ranges of the future. RISM started by document most existing US ranges and their capabilities. In parallel, RISM obtained inputs from the following: 1) NASA and NASA-contractor engineers and managers, and; 2) Aerospace leaders from Government, Academia, and Industry, participating through the Space Based Range Distributed System Working Group (SBRDSWG), many of whom are also; 3) Members of the Advanced Range Technology Working Group (ARTWG) subgroups, and; 4) Members of the Advanced Spaceport Technology Working Group (ASTWG). These diverse inputs helped to envision advanced technologies for implementing future Ranges and Range systems that builds on today s cabled and wireless legacy infrastructures while seamlessly integrating both today s emerging and tomorrow s building-block communication techniques. The fundamental key is to envision a transition to a Space Based Range Distributed Subsystem. The enabling concept is to identify the specific needs of Range users that can be solved through applying emerging communication tech

Bastin, Gary L.↗

Scalability of OpenFOAM Density-Based Solver with Runge–Kutta Temporal Discretization Scheme

Compressible density-based solvers are widely used in OpenFOAM, and the parallel scalability of these solvers is crucial for large-scale simulations. In this paper, we report our experiences with the scalability of OpenFOAM’s native rhoCentralFoam solver, and by making a small number of modifications to it, we show the degree to which the scalability of the solver can be improved. The main modification made is to replace the first-order accurate Euler scheme in rhoCentralFoam with a third-order accurate, four-stage Runge-Kutta or RK4 scheme for the time integration. The scaling test we used is the transonic flow over the ONERA M6 wing. This is a common validation test for compressible flows solvers in aerospace and other engineering applications. Numerical experiments show that our modified solver, referred to as rhoCentralRK4Foam, for the same spatial discretization, achieves as much as a 123.2% improvement in scalability over the rhoCentralFoam solver. As expected, the better time resolution of the Runge–Kutta scheme makes it more suitable for unsteady problems such as the Taylor–Green vortex decay where the new solver showed a 50% decrease in the overall time-to-solution compared to rhoCentralFoam to get to the final solution with the same numerical accuracy. Finally, the improved scalability can be traced to the improvement of the computation to communication ratio obtained by substituting the RK4 scheme in place of the Euler scheme. All numerical tests were conducted on a Cray XC40 parallel system, Theta, at Argonne National Laboratory.

Li, Sibo↗

Structural Simluation Toolkit (SST) v.12.0

The Structural Simulation Toolkit (SST) was developed to explore innovations in highly concurrent computing systems where the instruction set architecture (ISA), micro-architecture, and memory interact with the programming model and communications system. The package provides a fully modular design for extensive exploration of an individual system parameter as well as a parallel simulation environment based on message passing interface (MPI) which enable a high level of performance as well as the ability to look at large systems. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodrigues, ArunF.↗

Microcontroller-based aerosol jet printer control software

This software is used with a microcontroller to control an aerosol jet printing system. It was written for a 32-bit ARM microcontroller, providing greater processing speed and complexity relative to more conventional low-cost printer controllers. The microcontroller software was written using a real-time operating system (FreeRTOS), which simplifies parallel operation of multiple functions and further customization. It supports simultaneous motion planning, pulse generation for stepper motors, reading encoder feedback, communicating with multiple mass flow controllers, controlling solid state relays, handling a software-based emergency stop, and sending data to the client PC. The program is written in C and communicates with the client PC via USB connection. Sandia National Laboratories is a multi-mission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-0436 O

Secor, EthanBenjamin↗

Robust Event Simulation Variants Endowed to Ns-3 for General Exploration

Sandia's ns-3 contributions are modifications to the open source ns-3 network simulator that enable simulation speedup. One improvement, for example, removes the need for simulating the transmission and receipt of packets between nodes that are too distant to actually be able to communicate. In large-scale simulations this optimization has shown significant gains in performance. Additional contributions will pave the way for parallel discrete simulation (PDES) in ns-3. SAND2020-12456 O Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dickson, Joseph↗

A transputer based finite element solver

The feasibility of performing FEM structural-mechanics analyses on transputer systems is investigated experimentally. Transputers are programmable microprocessors equipped with local memory and point-to-point communication links; they can be joined in a large concurrent system via a programming language which supports distributed processing; this permits parallel processing at relatively low hardware cost. The computational tasks required by FEM programs are reviewed; the hardware (one PC, one master transputer, and 12 slave transputers) employed in the test calculations is described; and results demonstrating the speed and efficiency of the transputer array in assembling a global stiffness matrix and performing Gauss-Jordan matrix inversion are presented in graphs. It is predicted that larger transputer networks could approach the power of supercomputers at minicomputer costs.

Favenesi, J. A.↗

Electrooptical adaptive switching network for the hypercube computer

An all-optical network design for the hyperswitch network using regular free-space interconnects between electronic processor nodes is presented. The adaptive routing model used is described, and an adaptive routing control example is presented. The design demonstrates that existing electrooptical techniques are sufficient for implementing efficient parallel architectures without the need for more complex means of implementing arbitrary interconnection schemes. The electrooptical hyperswitch network significantly improves the communication performance of the hypercube computer.

Chow, E.↗

Fast, Massively Parallel Data Processors

Proposed fast, massively parallel data processor contains 8x16 array of processing elements with efficient interconnection scheme and options for flexible local control. Processing elements communicate with each other on "X" interconnection grid with external memory via high-capacity input/output bus. This approach to conditional operation nearly doubles speed of various arithmetic operations.

Heaton, Robert A.↗

Fast Multilevel Implementation of Recursive Spectral Bisection for Partitioning Unstructured Problems

If problems involving unstructured meshes are to be solved efficiently on distributed-memory parallel computers, the meshes must be partitioned and distributed across processors in a way that balances tile computational load and minimizes communication. The recursive spectral bisection method (RSB) has been shown to be very effective for such partitioning problems compared to alternative methods, but RSB in its simplest form is expensive. Here a multilevel version of RSB is introduced that attains about an order-of-magnitude improvement in run time on typical examples.

Barnard, Stephen T.↗

High-Payoff Space Transportation Design Approach with a Technology Integration Strategy

A general architectural design sequence is described to create a highly efficient, operable, and supportable design that achieves an affordable, repeatable, and sustainable transportation function. The paper covers the following aspects of this approach in more detail: (1) vehicle architectural concept considerations (including important strategies for greater reusability); (2) vehicle element propulsion system packaging considerations; (3) vehicle element functional definition; (4) external ground servicing and access considerations; and, (5) simplified guidance, navigation, flight control and avionics communications considerations. Additionally, a technology integration strategy is forwarded that includes: (a) ground and flight test prior to production commitments; (b) parallel stage propellant storage, such as concentric-nested tanks; (c) high thrust, LOX-rich, LOX-cooled first stage earth-to-orbit main engine; (d) non-toxic, day-of-launch-loaded propellants for upper stages and in-space propulsion; (e) electric propulsion and aero stage control.

McCleskey, C. M.↗

Real Time Photon-Counting Receiver for High Photon Efficiency Optical Communications

We present a scalable design for a photon-counting ground receiver based on superconducting nanowire single photon detectors (SNSPDs) and field programmable gate array (FPGA) real-time processing for applications to space-to-ground photon starved links, such as the Orion EM-2 Optical Communication Demonstration (O2O), and future deep space or low transmitter power missions. The receiver is designed to receive a serially concatenated pulse position modulation (SCPPM) waveform, which follows the Consultative Committee for Space Data Systems (CCSDS) Optical Communications Coding and Synchronization Red Book standard. The receiver design uses multiple individually fiber coupled, 80% detection efficiency commercial SNSPDs in parallel to scale to a required data rate, and is capable of achieving data rates up to 528 Mbps. For efficient fiber coupling from the telescope to the array of parallel detectors that can be scaled both to telescope aperture size and the number of detectors, we use either a single mode fiber (SMF) photonic lantern or a few-mode fiber (FMF) photonic lantern. In this paper we give an overview of the receiver system design, the characteristics of the photonic lanterns, the performance of the SNSPDs, and system level tests. We show that 40 Mbps can be received using a single SNSPD, and discuss aspects for scaling to higher data rates.

Vyhnalek, Brian E.↗

Real Time Photon-Counting Receiver for High Photon Efficiency Optical Communications

We present a scalable design for a photon-counting ground receiver based on superconducting nanowire single photon detectors (SNSPDs) and field programmable gate array (FPGA) real-time processing for applications to space-to-ground photon starved links, such as the Orion EM-2 Optical Communication Demonstration (O2O), and future deep space or low transmitter power missions. The receiver is designed to receive a serially concatenated pulse position modulation (SCPPM) waveform, which follows the Consultative Committee for Space Data Systems (CCSDS) Optical Communications Coding and Synchronization Red Book standard. The receiver design uses multiple individually fiber coupled, 80% detection efficiency commercial SNSPDs in parallel to scale to a required data rate, and is capable of achieving data rates up to 528 Mbps. For efficient fiber coupling from the telescope to the array of parallel detectors that can be scaled both to telescope aperture size and the number of detectors, we use either a single mode fiber (SMF) photonic lantern or a few-mode fiber (FMF) photonic lantern. In this paper we give an overview of the receiver system design, the characteristics of the photonic lanterns, the performance of the SNSPDs, and system level tests. We show that 40 Mbps can be received using a single SNSPD, and discuss aspects for scaling to higher data rates.

Vyhnalek, Brian E.↗

Mapping unstructured grid computations to massively parallel computers

Investigated here is this mapping problem: assign the tasks of a parallel program to the processors of a parallel computer such that the execution time is minimized. First, a taxonomy of objective functions and heuristics used to solve the mapping problem is presented. Next, we develop a highly parallel heuristic mapping algorithm, called Cyclic Pairwise Exchange (CPE), and discuss its place in the taxonomy. CPE uses local pairwise exchanges of processor assignments to iteratively improve an initial mapping. A variety of initial mapping schemes are tested and recursive spectral bipartitioning (RSB) followed by CPE is shown to result in the best mappings. For the test cases studied here, problems arising in computational fluid dynamics and structural mechanics on unstructured triangular and tetrahedral meshes, RSB and CPE outperform methods based on simulated annealing. Much less time is required to do the mapping and the results obtained are better. Compared with random and naive mappings, RSB and CPE reduce the communication time two fold for the test problems used. Finally, we use CPE in two applications on a CM-2. The first application is a data parallel mesh-vertex upwind finite volume scheme for solving the Euler equations on 2-D triangular unstructured meshes. CPE is used to map grid points to processors. The performance of this code is compared with a similar code on a Cray-YMP and an Intel iPSC/860. The second application is parallel sparse matrix-vector multiplication used in the iterative solution of large sparse linear systems of equations. We map rows of the matrix to processors and use an inner-product based matrix-vector multiplication. We demonstrate that this method is an order of magnitude faster than methods based on scan operations for our test cases.

Hammond, Steven Warren↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗