Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

High performance remote sensing data analysis using parallel computation

This paper examines the JPL/Caltech parallel processing system designed for rapid processing and transfer of large quantities of data from remote sensing instruments flown on NASA missions. Two remote sensing analysis applications that use this processing system are described: (1) an analysis system for retrieval of atmospheric parameters (such as species abundance, atmospheric temperature, and water vapor profiles) from data obtained by a Fourier transform IR spectrometer and (2) a prototype airborne SAR processing system. It is shown that a parallel processing system such as the JPL/Caltech system can offer supercomputer computational capability and high-volume data throughput and still be cost-effective.

Patterson, Jean E.↗

Enabling Parallel Execution of System-level Simulations in SAM

This report summarizes the recent code updates related to “element ghosting” in SAM to enable the parallel execution of system-level simulations using multiple processors/cores. Unlike typical MOOSE-based applications, for system-level simulations, SAM mostly deals with a collection of discrete small pieces of meshes, and the connection of physics on these meshes are realized by using “connector” types of components/code structures, such as conjugate heat transfer and flow junctions. The required code implementation is to correctly mark the necessary ghost elements for each type of such components/code structures; thus, the lower-level libraries can correctly perform the necessary data transfer between processors (CPUs) when executed in parallel mode. After the code updates, SAM can now run system-level simulations in the parallel mode. The parallel execution capability was then tested with an ABTR input model with 23k DOFs. Significant speedup was demonstrated when the optimal number of CPUs were used in parallel mode. Future systematic studies on parallelization performance using additional test cases covering different physics/scenarios will be needed to provide additional insights into the scalability of SAM parallelization.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Parallelization of Rocket Engine System Software (Press)

The main goal is to assess parallelization requirements for the Rocket Engine Numeric Simulator (RENS) project which, aside from gathering information on liquid-propelled rocket engines and setting forth requirements, involve a large FORTRAN based package at NASA Lewis Research Center and TDK software developed by SUBR/UWF. The ultimate aim is to develop, test, integrate, and suitably deploy a family of software packages on various aspects and facets of rocket engines using liquid-propellants. At present, all project efforts by the funding agency, NASA Lewis Research Center, and the HBCU participants are disseminated over the internet using world wide web home pages. Considering obviously expensive methods of actual field trails, the benefits of software simulators are potentially enormous. When realized, these benefits will be analogous to those provided by numerous CAD/CAM packages and flight-training simulators. According to the overall task assignments, Hampton University's role is to collect all available software, place them in a common format, assess and evaluate, define interfaces, and provide integration. Most importantly, the HU's mission is to see to it that the real-time performance is assured. This involves source code translations, porting, and distribution. The porting will be done in two phases: First, place all software on Cray XMP platform using FORTRAN. After testing and evaluation on the Cray X-MP, the code will be translated to C + + and ported to the parallel nCUBE platform. At present, we are evaluating another option of distributed processing over local area networks using Sun NFS, Ethernet, TCP/IP. Considering the heterogeneous nature of the present software (e.g., first started as an expert system using LISP machines) which now involve FORTRAN code, the effort is expected to be quite challenging.

Cezzar, Ruknet↗

Molecular solid-state inverter-converter system

A modular approach for aerospace electrical systems has been developed, using lightweight high efficiency pulse width modulation techniques. With the modular approach, a required system is obtained by paralleling modules. The modular system includes the inverters and converters, a paralleling system, and an automatic control and fault-sensing protection system with a visual annunciator. The output is 150 V dc, or a low distortion three phase sine wave at 120 V, 400 Hz. Input power is unregulated 56 V dc. Each module is rated 2.5 kW or 3.6 kVA at 0.7 power factor.

Birchenough, A. G.↗

Multi-Node Program Fuzzing on High Performance Computing Resources

Significant effort is placed on tuning the internal parameters of fuzzers to explore the state space, measured as coverage, of binaries. In this work, we investigate the effects of the external environment on the resulting coverage after fuzzing two binaries with AFL for 24 hours. Parameters such as scaling to multiple nodes, node saturation, and parallel file system type on HPC resources are controlled in order to maximize coverage. It will be shown that employing a parallel file system such as IBM's General Parallel File System offers an advantage for fuzzing operations, since it contains enhancements for performance optimization. When combined with scaling to two and four nodes, while simultaneously restricting the number of coordinated AFL tasks per node on the low end (10-50% of available physical cores), coverage may be enhanced within a shorter period of time. Thus, controlling the external environment is a useful effort.

97 MATHEMATICS AND COMPUTING↗

Optimizing Input/Output Using Adaptive File System Policies

Parallel input/output characterization studies and experiments with flexible resource management algorithms indicate that adaptivity is crucial to file system performance. In this paper we propose an automatic technique for selecting and refining file system policies based on application access patterns and execution environment. An automatic classification framework allows the file system to select appropriate caching and pre-fetching policies, while performance sensors provide feedback used to tune policy parameters for specific system environments. To illustrate the potential performance improvements possible using adaptive file system policies, we present results from experiments involving classification-based and performance-based steering.

Madhyastha, Tara M.↗

Parallel processing spacecraft communication system

An uplink controlling assembly speeds data processing using a special parallel codeblock technique. A correct start sequence initiates processing of a frame. Two possible start sequences can be used; and the one which is used determines whether data polarity is inverted or non-inverted. Processing continues until uncorrectable errors are found. The frame ends by intentionally sending a block with an uncorrectable error. Each of the codeblocks in the frame has a channel ID. Each channel ID can be separately processed in parallel. This obviates the problem of waiting for error correction processing. If that channel number is zero, however, it indicates that the frame of data represents a critical command only. That data is handled in a special way, independent of the software. Otherwise, the processed data further handled using special double buffering techniques to avoid problems from overrun. When overrun does occur, the system takes action to lose only the oldest data.

Bolotin, Gary S.↗

The massively parallel processor

Future sensor systems will utilize massively parallel computing systems for rapid analysis of two-dimensional data. The Goddard Space Flight Center has an ongoing program to develop these systems. A single-instruction multiple data computer known as the Massively Parallel Processor (MPP) is being fabricated for NASA by the Goodyear Aerospace Corporation. This processor contains 16,384 processing elements arranged in a 128 x 128 array. The MPP will be capable of adding more than 6 billion 8-bit numbers per second. Multiplication of eight-bit numbers can occur at a rate of 2 billion per second. Delivery of the MPP to Goddard Space Flight Center is scheduled for 1983.

Schaefer, D. H.↗

Seamless Transition of Critical Infrastructures using Droop Controlled Grid-forming Inverters

Seamless recovery of power to critical infrastructures, after grid failure, is a crucial need arising in scenarios that are increasingly becoming more frequent. Here, this article proposes a seamless transition strategy using a single and unified mode-dependent droop controlled grid-forming inverters. The control strategy achieves the following objectives: 1) regulates the output active and reactive power by the droop-controlled inverters to a desired value while operating in on-grid mode; 2) seamless transition and recovery of power injections into the load after grid failure by inverters that operates in grid-forming mode all the time; 3) requires only a single bit of information on the grid/network status for the mode transition. A framework for assessing the stability of the system and to guide the choice of parameters for controllers is developed using control-oriented modeling. A controller hardware-in-the-loop-based real-time simulation study on a test system based on the realistic electrical network of a commercial-scale medical center is conducted for initial prototyping of the control strategy. A hardware experiment is conducted with two 3 - $\phi$, 480 -V, 125 -kVA grid-forming inverters, a 3 - $\phi$, 480 -V, 270 -kVA grid simulator, a physical grid switch, and a physical load bank. The experimental data establishes the effectiveness of the always grid-forming operation and control of inverters in meeting power delivery objectives when on-grid and off-grid under various kinds of loads and scenarios while minimizing transients during transitions. Furthermore, performance comparison with existing strategies showcases the advantage of the proposed strategy.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Integrating PCLIPS into ULowell's Lincoln Logs: Factory of the future

We are attempting to show how independent but cooperating expert systems, executing within a parallel production system (PCLIPS), can operate and control a completely automated, fault tolerant prototype of a factory of the future (The Lincoln Logs Factory of the Future). The factory consists of a CAD system for designing the Lincoln Log Houses, two workcells, and a materials handling system. A workcell consists of two robots, part feeders, and a frame mounted vision system.

Mcgee, Brenda J.↗

Computational strategies for three-dimensional flow simulations on distributed computer systems

An increasing amount of research activity in computational fluid dynamics has been devoted to the development of efficient algorithms for parallel computing systems. The increasing performance to price ratio of engineering workstations has led to research to development procedures for implementing a parallel computing system composed of distributed workstations. This thesis proposal outlines an ongoing research program to develop efficient strategies for performing three-dimensional flow analysis on distributed computing systems. The PVM parallel programming interface was used to modify an existing three-dimensional flow solver, the TEAM code developed by Lockheed for the Air Force, to function as a parallel flow solver on clusters of workstations. Steady flow solutions were generated for three different wing and body geometries to validate the code and evaluate code performance. The proposed research will extend the parallel code development to determine the most efficient strategies for unsteady flow simulations.

Weed, Richard Allen↗

Configuration space representation in parallel coordinates

By means of a system of parallel coordinates, a nonprojective mapping from R exp N to R squared is obtained for any positive integer N. In this way multivariate data and relations can be represented in the Euclidean plane (embedded in the projective plane). Basically, R squared with Cartesian coordinates is augmented by N parallel axes, one for each variable. The N joint variables of a robotic device can be represented graphically by using parallel coordinates. It is pointed out that some properties of the relation are better perceived visually from the parallel coordinate representation, and that new algorithms and data structures can be obtained from this representation. The main features of parallel coordinates are described, and an example is presented of their use for configuration space representation of a mechanical arm (where Cartesian coordinates cannot be used).

Fiorini, Paolo↗

Queueing Network Models for Parallel Processing of Task Systems: an Operational Approach

Computer performance modeling of possibly complex computations running on highly concurrent systems is considered. Earlier works in this area either dealt with a very simple program structure or resulted in methods with exponential complexity. An efficient procedure is developed to compute the performance measures for series-parallel-reducible task systems using queueing network models. The procedure is based on the concept of hierarchical decomposition and a new operational approach. Numerical results for three test cases are presented and compared to those of simulations.

Mak, Victor W. K.↗

Characterizing parallel file-access patterns on a large-scale multiprocessor

Rapid increases in the computational speeds of multiprocessors have not been matched by corresponding performance enhancements in the I/O subsystem. To satisfy the large and growing I/O requirements of some parallel scientific applications, we need parallel file systems that can provide high-bandwidth and high-volume data transfer between the I/O subsystem and thousands of processors. Design of such high-performance parallel file systems depends on a thorough grasp of the expected workload. So far there have been no comprehensive usage studies of multiprocessor file systems. Our CHARISMA project intends to fill this void. The first results from our study involve an iPSC/860 at NASA Ames. This paper presents results from a different platform, the CM-5 at the National Center for Supercomputing Applications. The CHARISMA studies are unique because we collect information about every individual read and write request and about the entire mix of applications running on the machines. The results of our trace analysis lead to recommendations for parallel file system design. First the file system should support efficient concurrent access to many files, and I/O requests from many jobs under varying load conditions. Second, it must efficiently manage large files kept open for long periods. Third, it should expect to see small requests predominantly sequential access patterns, application-wide synchronous access, no concurrent file-sharing between jobs appreciable byte and block sharing between processes within jobs, and strong interprocess locality. Finally, the trace data suggest that node-level write caches and collective I/O request interfaces may be useful in certain environments.

Purakayastha, Apratim↗

Multiterminal High-Voltage dc Systems with Series-Parallel Valve Group-Based High-Voltage dc Substations

To transfer large amount of power over long distances, multiterminal direct current (MTdc) system based on bipole high-voltage direct current (HVdc) technology is a viable option. However, such system results in large dc transmission loss. The same can be reduced by increasing the dc voltage level. This paper introduces a new MTdc system architecture comprising of series (for increasing dc voltage level) and parallel (for increasing dc current capability) connected HVdc converters. The new architecture is compared with the bipole MTdc architecture in terms of equipment needed and dc transmission loss. The control modifications needed for the MTdc system are identified and the performance of the developed control is verified through electromagnetic transient (EMT) simulations.

Jaldanki, Sreenivasa↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗