Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “task parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Exploratory Climate Data Visualization and Analysis Using DV3D and UVCDAT

Earth system scientists are being inundated by an explosion of data generated by ever-increasing resolution in both global models and remote sensors. Advanced tools for accessing, analyzing, and visualizing very large and complex climate data are required to maintain rapid progress in Earth system research. To meet this need, NASA, in collaboration with the Ultra-scale Visualization Climate Data Analysis Tools (UVCOAT) consortium, is developing exploratory climate data analysis and visualization tools which provide data analysis capabilities for the Earth System Grid (ESG). This paper describes DV3D, a UV-COAT package that enables exploratory analysis of climate simulation and observation datasets. OV3D provides user-friendly interfaces for visualization and analysis of climate data at a level appropriate for scientists. It features workflow inte rfaces, interactive 40 data exploration, hyperwall and stereo visualization, automated provenance generation, and parallel task execution. DV30's integration with CDAT's climate data management system (COMS) and other climate data analysis tools provides a wide range of high performance climate data analysis operations. DV3D expands the scientists' toolbox by incorporating a suite of rich new exploratory visualization and analysis methods for addressing the complexity of climate datasets.

Maxwell, Thomas↗

Productive Programming of Distributed Systems with the SHAD C++ Library

High-performance computing (HPC) is often perceived as a matter of making large-scale systems (e.g., clusters) run as fast as possible, regardless the required programming effort. However, the idea of "bringing HPC to the masses" has recently emerged. Inspired by this vision, we have designed SHAD, the Scalable High-performance Algorithms and Data-structures library. SHAD is open source software, written in C++, for C++ developers. Unlike other HPC libraries for distributed systems, which rely on SPMD models, SHAD adopts a shared-memory programming abstraction, to make C++ programmers feel at home. Underneath, SHAD manages tasking and data-movements, moving the computation where data resides and taking advantage of asynchrony to tolerate network latency. At the bottom of his stack, SHAD can interface with multiple runtime systems: this not only improves developer’s productivity, by hiding the complexity of such software and of the underlying hardware, but also greatly enhance code portability. Thanks to its abstraction layers, SHAD can indeed target different systems, ranging from laptops to HPC clusters, without any need for modifying the user-level code. We have prototyped and open-sourced the implementation of (a subset of) the C++ standard library (STL) targeting multi-node HPC clusters. Our work allows plain STL-based C++ code to scale on HPC systems, with no need for rewriting the code to exploit the complex hardware. SHAD is available under Apache v2 License at https://github.com/pnnl/SHAD. In this paper we overview the design of the SHAD library, depicting its main components: runtime systems abstractions for tasking; parallel and distributed data-structures; STL-compliant interfaces and algorithms.

Castellana, Vito G.↗

The Myth of Fungible FTE: A Quantitative Assessment of Matrixed Resource Allocation

Matrix organizations allow scientific facilities to share specialized personnel across projects, operations, maintenance, and strategic initiatives. Nominal staffing allocations, however, may not capture the schedule consequences of fragmented individual commitments, limited access to specialist groups, and intermittent availability of key decision makers. We developed a stochas- tic, daily-time-step simulation of a hypothetical medium-sized accelerator-facility project com- prising sequential phases and parallel tasks. Each task requires role-specific work measured in FTE-days. Ordinary personnel may be unavailable because they contribute concurrently to other institutional activities, while designated key roles have independently specified daily un- availability probabilities. An organization-wide priority factor scales the number of people from each functional group who can effectively contribute to the project. It is interpreted as a composite proxy for project access and workforce fragmentation across competing commit- ments. We examined project completion time as a function of this factor and Project Lead unavailability using 100 Monte Carlo runs per condition. Increasing priority factor from 0.1 to 1.0 reduced median completion time from 1708.5 days (interquartile range 1681.5–1735.25) to 390 days (interquartile range 379–399). At priority factor = 0.1, increasing Project Lead unavailability from 0.5 to 0.9 increased median completion time from 1713.5 days (interquartile range 1691–1733.25) to 4,417 days (interquartile range 4271.75–4550.5). The model quantifies the commonly expected sensitivity of project schedules to fragmented resource commitments and limited coordination availability. Within this model, the results also indicate a possible threshold regime in which small increases in workforce availability yield only modest sched- ule improvements until sufficient capacity becomes accessible, after which project performance improves sharply. With further validation and calibration, this quantitative framework could support resource-allocation decisions during initial project planning and subsequent schedule rebaselining.

Bai, Mei [SLAC National Accelerator Laboratory (SL↗

Queueing Network Models for Parallel Processing of Task Systems: an Operational Approach

Computer performance modeling of possibly complex computations running on highly concurrent systems is considered. Earlier works in this area either dealt with a very simple program structure or resulted in methods with exponential complexity. An efficient procedure is developed to compute the performance measures for series-parallel-reducible task systems using queueing network models. The procedure is based on the concept of hierarchical decomposition and a new operational approach. Numerical results for three test cases are presented and compared to those of simulations.

Mak, Victor W. K.↗

Two criteria for the selection of assembly plans - Maximizing the flexibility of sequencing the assembly tasks and minimizing the assembly time through parallel execution of assembly tasks

The authors introduce two criteria for the evaluation and selection of assembly plans. The first criterion is to maximize the number of different sequences in which the assembly tasks can be executed. The second criterion is to minimize the total assembly time through simultaneous execution of assembly tasks. An algorithm that performs a heuristic search for the best assembly plan over the AND/OR graph representation of assembly plans is discussed. Admissible heuristics for each of the two criteria introduced are presented. Some implementation issues that affect the computational efficiency are addressed.

Homem De Mello, Luiz S.↗

Simultaneous dual-task performance reveals parallel response selection after practice

E. H. Schumacher, T. L. Seymour, J. M. Glass, D. E. Kieras, and D. E. Meyer (2001) reported that dual-task costs are minimal when participants are practiced and give the 2 tasks equal emphasis. The present research examined whether such findings are compatible with the operation of an efficient response selection bottleneck. Participants trained until they were able to perform both tasks simultaneously without interference. Novel stimulus pairs produced no reaction time costs, arguing against the development of compound stimulus-response associations (Experiment 1). Manipulating the relative onsets (Experiments 2 and 4) and durations (Experiments 3 and 4) of response selection processes did not lead to dual-task costs. The results indicate that the 2 tasks did not share a bottleneck after practice.

Practice (Psychology)↗

Task Assignment Heuristics for Parallel and Distributed CFD Applications

This paper proposes a task graph (TG) model to represent a single discrete step of multi-block overset grid computational fluid dynamics (CFD) applications. The TG model is then used to not only balance the computational workload across the overset grids but also to reduce inter-grid communication costs. We have developed a set of task assignment heuristics based on the constraints inherent in this class of CFD problems. Two basic assignments, the smallest task first (STF) and the largest task first (LTF), are first presented. They are then systematically costs. To predict the performance of the proposed task assignment heuristics, extensive performance evaluations are conducted on a synthetic TG with tasks defined in terms of the number of grid points in predetermined overlapping grids. A TG derived from a realistic problem with eight million grid points is also used as a test case.

Lopez-Benitez, Noe↗

MPI, HPF or OpenMP: A Study with the NAS Benchmarks

Porting applications to new high performance parallel and distributed platforms is a challenging task. Writing parallel code by hand is time consuming and costly, but this task can be simplified by high level languages and would even better be automated by parallelizing tools and compilers. The definition of HPF (High Performance Fortran, based on data parallel model) and OpenMP (based on shared memory parallel model) standards has offered great opportunity in this respect. Both provide simple and clear interfaces to language like FORTRAN and simplify many tedious tasks encountered in writing message passing programs. In our study, we implemented the parallel versions of the NAS Benchmarks with HPF and OpenMP directives. Comparison of their performance with the MPI implementation and pros and cons of different approaches will be discussed along with experience of using computer-aided tools to help parallelize these benchmarks. Based on the study, potentials of applying some of the techniques to realistic aerospace applications will be presented.

Jin, H.↗

MPI, HPF or OpenMP: A Study with the NAS Benchmarks

Porting applications to new high performance parallel and distributed platforms is a challenging task. Writing parallel code by hand is time consuming and costly, but the task can be simplified by high level languages and would even better be automated by parallelizing tools and compilers. The definition of HPF (High Performance Fortran, based on data parallel model) and OpenMP (based on shared memory parallel model) standards has offered great opportunity in this respect. Both provide simple and clear interfaces to language like FORTRAN and simplify many tedious tasks encountered in writing message passing programs. In our study we implemented the parallel versions of the NAS Benchmarks with HPF and OpenMP directives. Comparison of their performance with the MPI implementation and pros and cons of different approaches will be discussed along with experience of using computer-aided tools to help parallelize these benchmarks. Based on the study,potentials of applying some of the techniques to realistic aerospace applications will be presented

Jin, Hao-Qiang↗

Progress of the NASAUSGS Lunar Regolith Simulant Project

Beginning in 2004 personnel at MSFC began serious efforts to develop a new generation of lunar simulants. The first two products were a replication of the previous JSC-1 simulant under a contract to Orbitec and a major workshop in 2005 on future simulant development. It was recognized in early 2006 there were serious limitations with the standard approach of simply taking a single terrestrial rock and grinding it. To a geologist, even a cursory examination of the Lunar Sourcebook shows that matching lunar heterogeneity, crystal size, relative mineral abundances, lack of H2O, plagioclase chemistry and glass abundance simply can not be done with any simple combination of terrestrial rocks. Thus the project refocused its efforts and approached simulant development in a new and more comprehensive manner, examining new approaches in simulant development and ways to more accurately compare simulants to actual lunar materials. This led to a multi-year effort with five major tasks running in parallel. The five tasks are Requirements, Lunar Analysis, Process Development, Feed Stocks, and Standards.

Rickman, Douglas↗

Progress of the NASA/USGS Lunar Regolith Simulant Project

Beginning in 2004 personnel at MSFC began serious efforts to develop a new generation of lunar simulants. The first two products were a replication of the previous JSC-1 simulant under a contract to Orbitec and a major workshop in 2005 on future simulant development. Beginning in 2006 the project refocused its efforts and approached simulant development in a new and more comprehensive manner, examining new approaches in simulant development and ways to more accurately compare simulants to actual lunar materials. This led to a multi-year effort with five major tasks running in parallel. The five tasks are Requirements, Lunar Analysis, Process Development, Feed Stocks, and Standards. Major progress has been made in all five areas. A substantial draft of a formal requirements document now exists and has been largely stable since 2007. It does evolve as specific details of the standards and Lunar Analysis efforts proceed. Lunar Analysis has turned out to be vastly more difficult than anticipated. After great effort to mine existing published and gray literature, the team has realized the necessity of making new measurements of the Apollo samples, an effort that is currently in progress. Process development is substantially ahead of expectations in 2006. It is now practical to synthesize glasses of appropriate composition and purity. It is also possible to make agglutinate particles in significant quantities. A series of minerals commonly found on the Moon has been synthesized. Separation of mineral constituents from starting rock material is also proceeding. Customized grinding and mixing processes have been developed and tested are now being documented. Identification and development of appropriate feedstocks has been both easier and more difficult than anticipated. The Stillwater Mining Company, operating in the Stillwater layered mafic intrusive complex of Montana, has been an amazing resource for the project, but finding adequate sources for some of the components remains a difficult problem. For example the ratio of clino- to ortho-pyroxenes in the Stillwater is not an exact match for lunar materials. One of the sources being examined as an alternative pyroxene source is the Bushveld Complex in South Africa. Standards have been a major success for the project. The Figure of Merit algorithms have been created, tested, and are being considered for an ISO standard. Agreement has been reached in the community about how to make many of the critical measurements. There remains much work to do: (1) driving down the cost of simulants remains a major obstacle; (2) documentation and cost data analysis have not kept up with progress; (3) educating users in the complexity of the lunar regolith and the use of simulants remains a major task. In summary the project has made enormous progress and is successfully placing simulant development and use on a rigorous, scientifically defensible, engineering basis.

Rickman, Doug↗

Load balancing and task decomposition techniques for parallel implementation of integrated vision systems algorithms

Several techniques are presented to perform static and dynamic load balancing schemes for integrated vision systems. These techniques are novel in the sense that they capture the computational requirements of a task by examining the data when they are produced. Furthermore, they can be applied to many integrated vision systems because many algorithms in different systems are either the same or have similar computational characteristics. These techniques are evaluated by applying them to the algorithms in a motion estimation system. It is shown that the performance gains when these techniques are used are significant and the overhead of using these techniques is minimal. The performance is evaluated by implementing the algorithms using the presented techniques on a hypercube multiprocessor system.

Choudhary, Alok N.↗

Benchmarking Ada tasking on tightly coupled multiprocessor architectures

The development of benchmarks and performance measures for parallel Ada tasking is reported with emphasis on the macroscopic behavior of the benchmark across a set of load parameters. The application chosen for the study was the NASREM model for telerobot control, relevant to many NASA missions. The results of the study demonstrate the potential of parallel Ada in accomplishing the task of developing a control system for a system such as the Flight Telerobotic Servicer using the NASREM framework.

Collard, Philippe↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

Contra: A New Language for Task- and Data-Parallellism [Slides]

A new language is beneficial to take advantage of emerging architectures, and existing and new software technologies. Contra is a new language to provide performance portable code and consists of two innovations, which are described. A current status and illustrations of the code are shown.

97 MATHEMATICS AND COMPUTING↗

Research in Parallel Algorithms and Software for Computational Aerosciences

Phase 1 is complete for the development of a computational fluid dynamics CFD) parallel code with automatic grid generation and adaptation for the Euler analysis of flow over complex geometries. SPLITFLOW, an unstructured Cartesian grid code developed at Lockheed Martin Tactical Aircraft Systems, has been modified for a distributed memory/massively parallel computing environment. The parallel code is operational on an SGI network, Cray J90 and C90 vector machines, SGI Power Challenge, and Cray T3D and IBM SP2 massively parallel machines. Parallel Virtual Machine (PVM) is the message passing protocol for portability to various architectures. A domain decomposition technique was developed which enforces dynamic load balancing to improve solution speed and memory requirements. A host/node algorithm distributes the tasks. The solver parallelizes very well, and scales with the number of processors. Partially parallelized and non-parallelized tasks consume most of the wall clock time in a very fine grain environment. Timing comparisons on a Cray C90 demonstrate that Parallel SPLITFLOW runs 2.4 times faster on 8 processors than its non-parallel counterpart autotasked over 8 processors.

Domel, Neal D.↗

Research in Parallel Algorithms and Software for Computational Aerosciences

Phase I is complete for the development of a Computational Fluid Dynamics parallel code with automatic grid generation and adaptation for the Euler analysis of flow over complex geometries. SPLITFLOW, an unstructured Cartesian grid code developed at Lockheed Martin Tactical Aircraft Systems, has been modified for a distributed memory/massively parallel computing environment. The parallel code is operational on an SGI network, Cray J90 and C90 vector machines, SGI Power Challenge, and Cray T3D and IBM SP2 massively parallel machines. Parallel Virtual Machine (PVM) is the message passing protocol for portability to various architectures. A domain decomposition technique was developed which enforces dynamic load balancing to improve solution speed and memory requirements. A host/node algorithm distributes the tasks. The solver parallelizes very well, and scales with the number of processors. Partially parallelized and non-parallelized tasks consume most of the wall clock time in a very fine grain environment. Timing comparisons on a Cray C90 demonstrate that Parallel SPLITFLOW runs 2.4 times faster on 8 processors than its non-parallel counterpart autotasked over 8 processors.

Domel, Neal D.↗