Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Coarrars for Parallel Processing

The design of the Coarray feature of Fortran 2008 was guided by answering the question "What is the smallest change required to convert Fortran to a robust and efficient parallel language." Two fundamental issues that any parallel programming model must address are work distribution and data distribution. In order to coordinate work distribution and data distribution, methods for communication and synchronization must be provided. Although originally designed for Fortran, the Coarray paradigm has stimulated development in other languages. X10, Chapel, UPC, Titanium, and class libraries being developed for C++ have the same conceptual framework.

Fortran↗

Parallel physics-informed neural networks via domain decomposition

Here we develop a distributed framework for the physics-informed neural networks (PINNs) based on two recent extensions, namely conservative PINNs (cPINNs) and extended PINNs (XPINNs), which employ domain decomposition in space and in time-space, respectively. This domain decomposition endows cPINNs and XPINNs with several advantages over the vanilla PINNs, such as parallelization capacity, large representation capacity, efficient hyperparameter tuning, and is particularly effective for multi-scale and multi-physics problems. Here, we present a parallel algorithm for cPINNs and XPINNs constructed with a hybrid programming model described by MPI + X, where X ∈ {CPUs, GPUs}. The main advantage of cPINN and XPINN over the more classical data and model parallel approaches is the flexibility of optimizing all hyperparameters of each neural network separately in each subdomain. We compare the performance of distributed cPINNs and XPINNs for various forward problems, using both weak and strong scalings. Our results indicate that for space domain decomposition, cPINNs are more efficient in terms of communication cost but XPINNs provide greater flexibility as they can also handle time-domain decomposition for any differential equations, and can deal with any arbitrarily shaped complex subdomains. To this end, we also present an application of the parallel XPINN method for solving an inverse diffusion problem with variable conductivity on the United States map, using ten regions as subdomains.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Data Structure and Parallel Decomposition Considerations on a Fibonacci Grid

The Fibonacci grid, proposed by Swinbank and Purser (see companion abstract), provides attractive properties for global numerical atmospheric prediction by offering an optimally homogeneous, geometrically regular, and approximately isotropic discretization, with only the polar regions requiring special numerical treatment. It is a mathematical idealization, applied to the sphere, of the multi-spiral patterns often found in botanical structures, such as in pine cones and sunflower heads. Computationally, it is natural to organize the domain, into zones, in each of which the same pair, or triple, of "Fibonacci spirals" dominate. But the further subdivision of such zones into "tiles" of a shape and size suitable for distribution to the processors of a massively parallel computer requires very careful consideration if the subsequent spatial computations along the respective spirals, especially those computations (such as compact differencing schemes) that involve recursion, can be implemented in an efficient "load-balanced "manner without requiring excessive amounts of inter-processor communications. In this paper we show how certain "number theoretic" properties of the Fibonacci sequence (whose numbers prescribe the multiplicity of successive spirals) may be exploited in the decomposition of grid zones into tidy arrangements of triangular grid tiles, each tile possessing one side approximately parallel to the constant-latitude zone boundary. We also describe how the spatially recursive processes may be decomposed across such a tiling, and the directionality of the recursions reversed on alternate grid lines, to ensure a very high degree of load balancing throughout the execution of the computations required for one time step of a global model.

Michalakes, John↗

Utilizing Gaps and Key Performance Parameters to Inform NASA Environmental Control and Life Support and Human Health and Performance Capability Technology Decisions

Human spaceflight is a complex endeavor requiring a multitude of capabilities for transportation, crew health, scientific goals, and safe return to Earth. The difference between spaceflight proven capabilities and those needed for a particular mission is defined as a capability gap. Capability gaps are not technology specific. Each capability gap is approachable with a wide array of technologies that have unique benefits and challenges. Determining what a capability’s relevant and distinguishing key performance parameters (KPPs) are for a mission is critical. Mass, power, and volume are always constrained and important, but defining these in a way normalized by performance is challenging. Additionally, KPP definition for reliability, dormancy, and integration needs are very important and still evolving. This paper provides the approach of the Environmental Control and Life Support – Crew Health and Performance (ECLSS-CHP) System Capability Leadership Team (SCLT) to defining gaps and KPPs in support of the NASA’s Capabilities Integration Team data call objectives. The nine ECLSS-CHP capability areas are decomposed to capabilities, gaps, and KPPs. Rather than defining very detailed gaps, ECLSS-CHP defines high-level gaps to be technology agnostic. Within a gap, detailed KPPs are defined to both compare technologies and measure progress within a technology over time. Ideally, KPPs are clearly defined, widely communicated both internally and externally, and provide a common nomenclature to describe the state of the art and the degree of improvement required for exploration missions. KPPs help define when the gap is closed and the core mission objectives can be accomplished. Further technology improvements to enhance the capability, as measured by improved KPPs, must then be weighed against investments in open capability gaps that prevent NASA from achieving its exploration missions. It is uncommon that a technology maturation to improve all the relevant KPPs simultaneously but using KPPs is a critical technology investment decision making component. In addition to traditional technology selections, KPPs are informing how investments in ground testing prior to and in parallel with ISS technology demonstrations are required to improve reliability KPPs. The collection of all major technology activities within a capability area are captured on technology roadmaps to communicate how diverse program activities are coordinated to close gaps and infuse into exploration mission needs. A selection of ECLSS-CHP gaps and KPPs and their formulation, current state, and how they inform capability roadmap planning are discussed. The paper will contain a summary of the approximately 60 gaps. Gaps are classified as to their type (architecture, knowledge, technology, developmental, or engineering) depending on the magnitude of the gap. The paper will provide brief overviews of a few major technology challenges and the technologies being considered, but will reference detailed papers for a more thorough treatment of the challenges and state of the art. Data analysis of the gaps is in work and results are not currently available for this abstract. It is anticipated the paper will include examples of select KPPs with descriptions as to why these are the relevant measures. Additionally some KPPs will be graphically presented over time to show progress to date and when performance targets need to be achieved to support exploration missions. Graphical summaries of how gaps closures with near term mission elements support follow-on mission elements will be provided.

Life Support↗

Utilizing Gaps and Key Performance Parameters to Inform NASA Environmental Control and Life Support and Human Health and Performance Capability Technology Decisions

Human spaceflight is a complex endeavor requiring multiple capabilities for transportation, crew health, scientific goals, and safe return to Earth. The difference between spaceflight proven capabilities and those needed for future exploration architectures is defined as a capability gap. Capability gaps are not technology specific. Each capability gap is approachable with a wide array of technologies that have unique benefits and challenges. Determining what a capability’s relevant and distinguishing key performance parameters (KPPs) are for a mission is critical. Mass, power, and volume are always constrained and important, but defining these in a way normalized by performance is challenging. Additionally, KPP definition for reliability, dormancy, and integration needs are very important and still evolving. This paper provides the approach of the Environmental Control and Life Support – Crew Health and Performance (ECLSS-CHP) System Capability Leadership Team (SCLT) has used to define gaps and KPPs in support of the NASA’s Capabilities Integration Team data call objectives. The nine ECLSS-CHP capability areas are decomposed to capabilities with ~76 gaps and supported with KPPs. Rather than defining very detailed gaps, ECLSS-CHP defines high-level gaps to be technology agnostic. Within a gap, detailed KPPs are defined to both compare technologies and measure progress within a technology over time. Ideally, KPPs are clearly defined, widely communicated both internally and externally, and provide a common nomenclature to describe the state of the art and the degree of improvement required for exploration missions. KPPs help define when the gap is closed, and the core mission objectives can be accomplished. Further technology improvements to enhance the capability, as measured by improved KPPs, must then be weighed against investments in open capability gaps that prevent NASA from achieving its exploration missions. It is uncommon that a technology maturation to improve all the relevant KPPs simultaneously but using KPPs is a critical technology investment decision making component. In addition to traditional technology selections, KPPs are informing how investments in ground testing prior to and in parallel with ISS technology demonstrations are required to improve reliability KPPs. The collection of all major technology activities within a capability area are captured on technology roadmaps to communicate how diverse program activities are coordinated to close gaps and infuse into exploration mission needs. A selection of ECLSS-CHP gaps and KPPs and their formulation, current state, and how they inform capability roadmap planning are discussed.

Life Support↗

Impact of Blockchain Delay on Grid-Tied Solar Inverter Performance

This paper investigates the impact of the delay resulting from a blockchain, a promising security measure, for a hierarchical control system of inverters connected to the grid. The blockchain communication network is designed at the secondary control layer for resilience against cyberattacks. To represent the latency in the communication channel, a model is developed based on the complexity of the blockchain framework. Taking this model into account, this work evaluates the plant’s performance subject to communication delays, introduced by the blockchain, among the hierarchical control agents. In addition, this article considers an optimal model-based control strategy that performs the system’s internal control loop. The work shows that the blockchain’s delay size influences the convergence of the power supplied by the inverter to the reference at the point of common coupling. In the results section, real-time simulations on OPAL-RT are performed to test the resilience of two parallel inverters with increasing blockchain complexity.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Caffeine v0.1.0

Caffeine is the CoArray Fortran Framework of Efficient Interfaces to Network Environments. Caffeine aims to produce a parallel runtime library that will support Fortran compilers with a programming-model-agnostic application binary interface (ABI) to various lower-level communication libraries. The current version of Caffeine uses the GASNet-EX networking middleware, also developed at Berkeley Lab. On many combinations of applications and platforms, GASNet-EX outperforms the widely used Message Passing Interface (MPI). Through GASNet-EX's support for communicating between graphics processing units (GPU), GASNet-EX has features that specifically target the emerging, leading-edge exascale computing platforms.

Rouson, Damian↗

Moving formal methods into practice. Verifying the FTPP Scoreboard: Results, phase 1

This report documents the Phase 1 results of an effort aimed at formally verifying a key hardware component, called Scoreboard, of a Fault-Tolerant Parallel Processor (FTPP) being built at Charles Stark Draper Laboratory (CSDL). The Scoreboard is part of the FTPP virtual bus that guarantees reliable communication between processors in the presence of Byzantine faults in the system. The Scoreboard implements a piece of control logic that approves and validates a message before it can be transmitted. The goal of Phase 1 was to lay the foundation of the Scoreboard verification. A formal specification of the functional requirements and a high-level hardware design for the Scoreboard were developed. The hardware design was based on a preliminary Scoreboard design developed at CSDL. A main correctness theorem, from which the functional requirements can be established as corollaries, was proved for the Scoreboard design. The goal of Phase 2 is to verify the final detailed design of Scoreboard. This task is being conducted as part of a NASA-sponsored effort to explore integration of formal methods in the development cycle of current fault-tolerant architectures being built in the aerospace industry.

Srivas, Mandayam↗

A Lagrange multiplier based divide and conquer finite element algorithm

A novel domain decomposition method based on a hybrid variational principle is presented. Prior to any computation, a given finite element mesh is torn into a set of totally disconnected submeshes. First, an incomplete solution is computed in each subdomain. Next, the compatibility of the displacement field at the interface nodes is enforced via discrete, polynomial and/or piecewise polynomial Lagrange multipliers. In the static case, each floating subdomain induces a local singularity that is resolved very efficiently. The interface problem associated with this domain decomposition method is, in general, indefinite and of variable size. A dedicated conjugate projected gradient algorithm is developed for solving the latter problem when it is not feasible to explicitly assemble the interface operator. When implemented on local memory multiprocessors, the proposed methodology requires less interprocessor communication than the classical method of substructuring. It is also suitable for parallel/vector computers with shared memory and compares favorably with factorization based parallel direct methods.

Farhat, C.↗

Cell contact as an independent factor modulating cardiac myocyte hypertrophy and survival in long-term primary culture

Cardiac myocytes maintained in cell culture develop hypertrophy both in response to mechanical loading as well as to receptor-mediated signaling mechanisms. However, it has been shown that the hypertrophic response to these stimuli may be modulated through effects of intercellular contact achieved by maintaining cells at different plating densities. In this study, we show that the myocyte plating density affects not only the hypertrophic response and features of the differentiated phenotype of isolated adult myocytes, but also plays a significant role influencing myocyte survival in vitro. The native rod-shaped phenotype of freshly isolated adult myocytes persists in an environment which minimizes myocyte attachment and spreading on the substratum. However, these conditions are not optimal for long-term maintenance of cultured adult cardiac myocytes. Conditions which promote myocyte attachment and spreading on the substratum, on the other hand, also promote the re-establishment of new intercellular contacts between myocytes. These contacts appear to play a significant role in the development of spontaneous activity, which enhances the redevelopment of highly differentiated contractile, junctional, and sarcoplasmic reticulum structures in the cultured adult cardiomyocyte. Although it has previously been shown that adult cardiac myocytes are typically quiescent in culture, the addition of beta-adrenergic agonists stimulates beating and myocyte hypertrophy, and thereby serves to increase the level of intercellular contact as well. However, in densely-plated cultures with intrinsically high levels of intercellular contact, spontaneous contractile activity develops without the addition of beta-adrenergic agonists. In this study, we compare the function, morphology, and natural history of adult feline cardiomyocytes which have been maintained in cultures with different levels of intercellular contact, with and without the addition of beta-adrenergic agonists. Intercellular contact, communication, and transmission of contractile forces between myocytes appears to play a primary role in remodeling the 2-dimensional cell layer into a parallel alignment of elongated myocytes with highly developed intercalated disk-like junctions. This highly differentiated state is very stable, and cultures which achieve this state exhibit significantly greater longevity than more sparsely plated myocytes. These myocytes typically continue beating, and survive from 6 to more than 12 weeks in culture. When this level of contact and differentiation are not achieved, even among beta-adrenergic stimulated myocytes, contractile activity is not sustained, myofibrils atrophy, there is little or no development of junctional complexes, and the period of myocyte viability is typically no more than 5 weeks in vitro.

Non-NASA Center↗

Lunar Terrain Coverage Analysis Data Delivery Workflow

In this work, we are developing a lunar terrain database to enable fast rendering of sun illumination and earth visibility for a proposed coverage analysis tool. This development will advance lunar mission design and formulation for current and future communications architectures, and will aid in lunar surface mission planning and communications/navigation operations. Our effort can be described in three steps: (1) we parallelize a brute force algorithm, which computes elevation masks from laser altimetry data acquired by the Lunar Reconnaissance Orbiter’s (LRO) Lunar Orbiter Laser Altimeter (LOLA); (2) we investigate parallel I/O methods to store terrain mask information from step (1) into a parallel file system; and (3) we finally deliver data to the terrain coverage analysis tool.

Michels, Dominik↗

Impacts of Hybrid Parallelism and Vectorization on the Performance of Newton-Krylov Methods in Computational Aerodynamics

Finding the numerical solution of moderate and high-fidelity aerodynamics problems on modern computer architectures involves, 1) decomposing the domain into smaller regions of nearly equal size, and 2) allocating computational resources for calculations on each domain and communication between domains. Modern computer clusters are composed from hierarchies of processing, memory, and communication resources with varying capabilities and latencies.This paper focuses on the combination of domain decomposition provided by ParMETIS [1]and Newton-Krylov Methods [2–5] for the solution of Computational Aerodynamics problems of interest to NASA. Herein, trade-offs encountered when mapping aerodynamics problems to modern computer architectures are explored through examples and discussions of trade-offs in parallelism from MPI [6], Open MP [7], and vectorization as partition sizes and computational resources are varied. An example of the impact that domain decomposition and MPI+OpenMPresource allocation can have on an adjoint calculation is presented in this abstract. The full paper will include more detailed examples, discussions of difficulties and potential methods to overcome them, and topics identified for future study.

Computational Aerodynamics, Hybrid Parallelism, Ve↗

Novel Highly Parallel and Systolic Architectures Using Quantum Dot-Based Hardware

VLSI technology has made possible the integration of massive number of components (processors, memory, etc.) into a single chip. In VLSI design, memory and processing power are relatively cheap and the main emphasis of the design is on reducing the overall interconnection complexity since data routing costs dominate the power, time, and area required to implement a computation. Communication is costly because wires occupy the most space on a circuit and it can also degrade clock time. In fact, much of the complexity (and hence the cost) of VLSI design results from minimization of data routing. The main difficulty in VLSI routing is due to the fact that crossing of the lines carrying data, instruction, control, etc. is not possible in a plane. Thus, in order to meet this constraint, the VLSI design aims at keeping the architecture highly regular with local and short interconnection. As a result, while the high level of integration has opened the way for massively parallel computation, practical and full exploitation of such a capability in many applications of interest has been hindered by the constraints on interconnection pattern. More precisely. the use of only localized communication significantly simplifies the design of interconnection architecture but at the expense of somewhat restricted class of applications. For example, there are currently commercially available products integrating; hundreds of simple processor elements within a single chip. However, the lack of adequate interconnection pattern among these processing elements make them inefficient for exploiting a large degree of parallelism in many applications.

Fijany, Amir↗

End-to-end Analytics for Grid Arch Design & All-hazard Assessment

Resiliency, reliability, and security of the next-generation smart grid depend upon leveraging advanced communication and computing technologies, integrating them with physical power systems, and developing real-time, fast, data-based applications to help in wide-area monitoring and control of the grid. Using a high sampling data rate from phasor measurement units (PMUs) to develop applications has opened the door to achieving the next-generation grid requirements. The North American Synchrophasor Initiative Network (NASPlnet) was developed in 2007-09 to create a standard and guide for PMU data exchanges. With the advancement in both networking and grid requirements, it is necessary to evaluate the performance of different NASPInet versions and their impact on applications. Therefore, we need a cyber-power cosimulation framework that supports very large-scale co-simulation capable of running in parallel, high-performance computing platforms and capturing real-life network behavior. This work presents a cyber-physical co-simulation testbed using NS3 to model the communication network, GridPACK to model the power grid, and HELICS as a co-simulation engine. Comparative analysis of latency in synchrophasor networks and a performance evaluation of a power system stabilizer application based on PMU data in an Institute of Electrical and Electronics Engineers 39-bus test system is presented using this co-simulation testbed.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Navier-Strokes Chimera Code on the Connection Machine CM-5: Design and Performance

We have implemented a three-dimensional compressible Navier-Stokes code on the Connection Machine CM-5. The code is set up for implicit time-stepping on single or multiple structured grids. For multiple grids and geometrically complex problems, we follow the 'chimera' approach, where flow data on one zone is interpolated onto another in the region of overlap. We will describe our design philosophy and give some timing results for the current code. A parallel machine like the CM-5 is well-suited for finite-difference methods on structured grids. The regular pattern of connections of a structured mesh maps well onto the architecture of the machine. So the first design choice, finite differences on a structured mesh, is natural. We use centered differences in space, with added artificial dissipation terms. When numerically solving the Navier-Stokes equations, there are liable to be some mesh cells near a solid body that are small in at least one direction. This mesh cell geometry can impose a very severe CFL (Courant-Friedrichs-Lewy) condition on the time step for explicit time-stepping methods. Thus, though explicit time-stepping is well-suited to the architecture of the machine, we have adopted implicit time-stepping. We have further taken the approximate factorization approach. This creates the need to solve large banded linear systems and creates the first possible barrier to an efficient algorithm. To overcome this first possible barrier we have considered two options. The first is just to solve the banded linear systems with data spread over the whole machine, using whatever fast method is available. This option is adequate for solving scalar tridiagonal systems, but for scalar pentadiagonal or block tridiagonal systems it is somewhat slower than desired. The second option is to 'transpose' the flow and geometry variables as part of the time-stepping process: Start with x-lines of data in-processor. Form explicit terms in x, then transpose so y-lines of data are in-processor. Form explicit terms in y, then transpose so z-lines are in processor. Form explicit terms in z, then solve linear systems in the z-direction. Transpose to the y-direction, then solve linear systems in the y-direction. Finally transpose to the x direction and solve linear systems in the x-direction. This strategy avoids inter-processor communication when differencing and solving linear systems, but requires a large amount of communication when doing the transposes. The transpose method is more efficient than the non-transpose strategy when dealing with scalar pentadiagonal or block tridiagonal systems. For handling geometrically complex problems the chimera strategy was adopted. For multiple zone cases we compute on each zone sequentially (using the whole parallel machine), then send the chimera interpolation data to a distributed data structure (array) laid out over the whole machine. This information transfer implies an irregular communication pattern, and is the second possible barrier to an efficient algorithm. We have implemented these ideas on the CM-5 using CMF (Connection Machine Fortran), a data parallel language which combines elements of Fortran 90 and certain extensions, and which bears a strong similarity to High Performance Fortran. We make use of the Connection Machine Scientific Software Library (CMSSL) for the linear solver and array transpose operations.

Jespersen, Dennis C.↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

A highly reliable, autonomous data communication subsystem for an advanced information processing system

The need to meet the stringent performance and reliability requirements of advanced avionics systems has frequently led to implementations which are tailored to a specific application and are therefore difficult to modify or extend. Furthermore, many integrated flight critical systems are input/output intensive. By using a design methodology which customizes the input/output mechanism for each new application, the cost of implementing new systems becomes prohibitively expensive. One solution to this dilemma is to design computer systems and input/output subsystems which are general purpose, but which can be easily configured to support the needs of a specific application. The Advanced Information Processing System (AIPS), currently under development has these characteristics. The design and implementation of the prototype I/O communication system for AIPS is described. AIPS addresses reliability issues related to data communications by the use of reconfigurable I/O networks. When a fault or damage event occurs, communication is restored to functioning parts of the network and the failed or damage components are isolated. Performance issues are addressed by using a parallelized computer architecture which decouples Input/Output (I/O) redundancy management and I/O processing from the computational stream of an application. The autonomous nature of the system derives from the highly automated and independent manner in which I/O transactions are conducted for the application as well as from the fact that the hardware redundancy management is entirely transparent to the application.

Nagle, Gail↗

High-Performance, Low-Complexity Codes Researched for Communication Channels

NASA Lewis Research Center s Communications Technology Division has an ongoing program in the development of efficient channel coding schemes for satellite communications applications. Through a university grant, as a part of this research, the University of Toledo is investigating the performance of turbocodes, which use parallel concatenation of non-systematic convolutional encoders with an interleaver. The error correcting capacity of these codes is close to the Shannon limit. The research emphasis is on the development of low-complexity, but higher rate (greater than one half), turbocodes and on the iterative decoding of block codes.

Kwatra, Subhash C.↗