Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Parallel solution of closely coupled systems

An odd-even permutation and a nested dissection technique were used to circumvent the strong seriality of a system of closely coupled equations. The effect of transforming the n x n Hermitian definite positive matrix coefficient on the topology of Cholesky factors is discussed. A series of directed graphs is constructed in order to show the computational steps required for the odd-even permutation. Numerical expressions for the speed-up and efficiency of parallel N-processing techniques and sequential processing by a single computer are derived. Similar expressions are derived for the case of insufficient processing capacity. The application of the odd-even permutation to the ensemble class of computer architectures is demonstrated.

Utku, S.↗

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

Accelerating Application Bulk Synchronous Writes in HPC Environments

High-bandwidth storage tiers are becoming more common for their capability to absorb high-rate, bursty I/Os. Notably, the designs of these fast storage tiers differ from system to system. The variation of these layers and non-uniform methods of access can pose chal- lenges for applications seeking to run at multiple HPC facilities. Therefore, in this work, we present Spectral, a rapid-output ab- straction library to accelerate application, bulk-synchronous writes on HPC systems. We design Spectral to enable applications to use high-bandwidth storage, such as node-local storage and dis- tributed, write-caches (e.g., burst buffers) transparently without requiring modifications to the application or file system source code. The key idea is to allow applications to spend most of the time performing productive work and to not require any source code changes for maximum portability on different HPC archi- tectures. Spectral internally re-routes write-only files through available, high-performance I/O resources before ultimately mi- grating them to the shared global parallel file system. For instance, on Summit, Spectral transparently places application outputs on node-local storage and then utilizes asynchronous migration to the center-wide GPFS file system. We evaluate Spectral on the Summit HPC system (1024 nodes) using the IOR benchmark and real scientific applications. Spectral shows linear performance scaling, improving application write performance by over an order of magnitude when compared to GPFS.

Khan, Awais↗

Frontier (HPE Cray EX) Exascale Supercomputer at the Oak Ridge Leadership Computing Facility

Frontier is the HPE Cray EX exascale supercomputer deployed and operated by the Oak Ridge Leadership Computing Facility (OLCF) at Oak Ridge National Laboratory (ORNL). Frontier is designed for large-scale modeling, simulation, and AI workloads and is built from HPE Cray EX system architecture with AMD CPUs and AMD Instinct GPU accelerators connected by the HPE Slingshot interconnect. System composition (representative production configuration): Frontier is composed of approximately 74 cabinets with 128 compute nodes per cabinet (~9,400 compute nodes total). Each compute node contains one 64-core AMD EPYC CPU and four AMD Instinct MI250X GPUs. Nodes are connected using HPE Slingshot (Slingshot-200 class) networking with multiple NIC ports per node providing high injection bandwidth. Frontier is connected to the Orion parallel file system (multi-tier Lustre) providing a large, center-wide high-performance storage namespace. Operational context: Frontier entered public prominence as the first system to reach No. 1 on the TOP500 list in May 2022 (HPL benchmark), establishing the first widely recognized exascale-era performance milestone. The system supports DOE Office of Science mission workloads and enables leadership-class computational science and AI for open science users.

AMD EPYC↗

Improved CDMA Performance Using Parallel Interference Cancellation

This paper considers a general parallel interference cancellation scheme that significantly reduces the degradation effect of user interference but with a lesser implementation complexity than the maximum-likelihood technique. The scheme operates on the fact that parallel processing simultaneously removes from each user the total interference produced by the remaining most reliably received users accessing the channel. The parallel processing can be done in multiple stages. The proposed scheme uses tentative decision devices with different optimum thresholds at the multiple stages to produce the most reliably received data for generation and cancellation of user interference.

CDMA↗

Initial Kernel Timing Using a Simple PIM Performance Model

This presentation will describe some initial results of paper-and-pencil studies of 4 or 5 application kernels applied to a processor-in-memory (PIM) system roughly similar to the Cascade Lightweight Processor (LWP). The application kernels are: * Linked list traversal * Sun of leaf nodes on a tree * Bitonic sort * Vector sum * Gaussian elimination The intent of this work is to guide and validate work on the Cascade project in the areas of compilers, simulators, and languages. We will first discuss the generic PIM structure. Then, we will explain the concepts needed to program a parallel PIM system (locality, threads, parcels). Next, we will present a simple PIM performance model that will be used in the remainder of the presentation. For each kernel, we will then present a set of codes, including codes for a single PIM node, and codes for multiple PIM nodes that move data to threads and move threads to data. These codes are written at a fairly low level, between assembly and C, but much closer to C than to assembly. For each code, we will present some hand-drafted timing forecasts, based on the simple PIM performance model. Finally, we will conclude by discussing what we have learned from this work, including what programming styles seem to work best, from the point-of-view of both expressiveness and performance.

BRIEFING CHARTS↗

System efficiency of a microwave power tube with a multistage depressed collector

The efficiencies of a microwave power tube with a multistage depressed collector and of the power supply driving the tube are computed. An analytical expression for the collector efficiency, which includes the effect of secondary emission and the radial component of velocity, is derived for a hypothetical current probability distribution function. In addition, collector efficiency is calculated with the aid of a digital computer for a specific current distribution. The efficiency of the power supply required to operate the tube in a space environment is estimated by using a simple parallel inverter system.

Dayton, J. A., Jr.↗

Massively parallel support for a case-based planning system

Case-based planning (CBP), a kind of case-based reasoning, is a technique in which previously generated plans (cases) are stored in memory and can be reused to solve similar planning problems in the future. CBP can save considerable time over generative planning, in which a new plan is produced from scratch. CBP thus offers a potential (heuristic) mechanism for handling intractable problems. One drawback of CBP systems has been the need for a highly structured memory to reduce retrieval times. This approach requires significant domain engineering and complex memory indexing schemes to make these planners efficient. In contrast, our CBP system, CaPER, uses a massively parallel frame-based AI language (PARKA) and can do extremely fast retrieval of complex cases from a large, unindexed memory. The ability to do fast, frequent retrievals has many advantages: indexing is unnecessary; very large case bases can be used; memory can be probed in numerous alternate ways; and queries can be made at several levels, allowing more specific retrieval of stored plans that better fit the target problem with less adaptation. In this paper we describe CaPER's case retrieval techniques and some experimental results showing its good performance, even on large case bases.

Kettler, Brian P.↗

Evaluation of Cache-based Superscalar and Cacheless Vector Architectures for Scientific Computations

The growing gap between sustained and peak performance for scientific applications has become a well-known problem in high performance computing. The recent development of parallel vector systems offers the potential to bridge this gap for a significant number of computational science codes and deliver a substantial increase in computing capabilities. This paper examines the intranode performance of the NEC SX6 vector processor and the cache-based IBM Power3/4 superscalar architectures across a number of key scientific computing areas. First, we present the performance of a microbenchmark suite that examines a full spectrum of low-level machine characteristics. Next, we study the behavior of the NAS Parallel Benchmarks using some simple optimizations. Finally, we evaluate the perfor- mance of several numerical codes from key scientific computing domains. Overall results demonstrate that the SX6 achieves high performance on a large fraction of our application suite and in many cases significantly outperforms the RISC-based architectures. However, certain classes of applications are not easily amenable to vectorization and would likely require extensive reengineering of both algorithm and implementation to utilize the SX6 effectively.

Oliker, Leonid↗

PRO-X Fuel Cycle Transportation and Crosscutting Progress Report

The PRO-X program is actively supporting the design of nuclear systems by developing a framework to both optimize the fuel cycle infrastructure for advanced reactors (ARs) and minimize the potential for production of weapons-usable nuclear material. Three study topics are currently being investigated by Sandia National Laboratories (SNL) with support from Argonne National Laboratories (ANL). This multi-lab collaboration is focused on three study topics which may offer proliferation resistance opportunities or advantages in the nuclear fuel cycle. These topics are: 1) Transportation Global Landscape, 2) Transportation Avoidability, and 3) Parallel Modular Systems vs Single Large System (Crosscutting Activity).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Parallel solution of closely coupled systems

The odd-even permutation and associated unitary transformations for reordering the matrix coefficient A are employed as means of breaking the strong seriality which is characteristic of closely coupled systems. The nested dissection technique is also reviewed, and the equivalence between reordering A and dissecting its network is established. The effect of transforming A with odd-even permutation on its topology and the topology of its Cholesky factors is discussed. This leads to the construction of directed graphs showing the computational steps required for factoring A, their precedence relationships and their sequential and concurrent assignment to the available processors. Expressions for the speed-up and efficiency of using N processors in parallel relative to the sequential use of a single processor are derived from the directed graph. Similar expressions are also derived when the number of available processors is fewer than required.

Utku, S.↗

kiloAMPS (Final Report)

This project aims to reduce the testing time of various electrochemical techniques for neural interface electrodes by developing an automated PCB design which would be capable of testing multiple interface electrodes in parallel. The system should be able to perform the full battery of electrochemical tests, though electrochemical impedance spectroscopy (EIS) will be of primary interest. Previous versions of this project have already improved upon traditional interface testing by automating many parts of the process, reducing the test time to around a day and requiring the technician to check up on the process once every six hours. A key bottleneck is that the electrode testing is still done sequentially. Testing time could be heavily reduced if electrodes could be tested in parallel.

42 ENGINEERING↗

ISS EVA80 EMU Water in the Helmet Failure and Associated Analytical Response

On March 23rd, 2022, during Extra-Vehicular Activity (EVA) 80 aboard the International Space Station (ISS) crew identified an 8-10-inch diameter thin film of water in the pressure bubble of the suit during repress. This launched an investigation to determine the cause of the water in the helmet as well as short- and long-term mitigation efforts to enable return to EVA as quickly as possible. The investigation analysis efforts in combination with hardware inspection and testing determined the most likely cause of the water in the helmet to be sublimator carryover as a result of high latent loading put on the system. In parallel with this investigation, efforts were being made to mitigate any future water in the helmet events. The short-term mitigation strategy developed is to install absorbent material in the pressure bubble to capture water as it enters the helmet before it can impact the astronauts. This hardware is referred to as the Helmet Absorption Band (HAB) and Helmet Absorption Pad – Extender (HAP-E). Longer-term mitigation strategies include a device to capture small water events before entering the helmet by installing a water capture system in the vent loop of the Extravehicular Mobility Unit (EMU). This effort is coined the T2 Water Capture System. Additionally, developing an on-orbit sublimator challenge test which will be able to verify sublimator performance before and after EVAs. The sublimator is the hardware which condenses water vapor and removes it from the vent loop, and this hardware is referred to as the EMU Moisture Injection Test System (EMITS). Based on the investigation and short-term water capture solutions the ISS Program decided to return to nominal EVAs on October 7, 2022. With the addition of the long-term mitigation strategies, the team hopes to be able to prevent any future water in the helmet events.

Veronica Lee Pizor↗

Optical to optical interface device

The development, fabrication, and testing of a preliminary model of an optical-to-optical (noncoherent-to-coherent) interface device for use in coherent optical parallel processing systems are described. The developed device demonstrates a capability for accepting as an input a scene illuminated by a noncoherent radiation source and providing as an output a coherent light beam spatially modulated to represent the original noncoherent scene. The converter device developed under this contract employs a Pockels readout optical modulator (PROM). This is a photosensitive electro-optic element which can sense and electrostatically store optical images. The stored images can be simultaneously or subsequently readout optically by utilizing the electrostatic storage pattern to control an electro-optic light modulating property of the PROM. The readout process is parallel as no scanning mechanism is required. The PROM provides the functions of optical image sensing, modulation, and storage in a single active material.

Oliver, D. S.↗

Characterizing Machine Learning I/O Workloads on Leadership Scale HPC Systems

High performance computing (HPC) is no longer solely limited to traditional workloads such as simulation and modeling. With the increase in the popularity of machine learning (ML) and deep learning (DL) technologies, we are observing that an increasing number of HPC users are incorporating ML methods into their workflow and scientific discovery processes, across a wide spectrum of science domains such as biology, earth science, and physics. This gives rise to a diverse set of I/O patterns than the traditional checkpoint/restart-based HPC I/O behavior. The details of the I/O characteristics of such ML I/O workloads have not been studied extensively for large-scale leadership HPC systems. This paper aims to fill that gap by providing an in-depth analysis to gain an understanding of the I/O behavior of ML I/O workloads using darshan - an I/O characterization tool designed for lightweight tracing and profiling. We study the darshan logs of more than 23, 000 HPC ML I/O jobs over a time period of one year running on Summit - the second-fastest supercomputer in the world. This paper provides a systematic I/O characterization of ML I/O jobs running on a leadership scale supercomputer to understand how the I/O behavior differs across science domains and the scale of workloads, and analyze the usage of parallel file system and burst buffer by ML I/O workloads.

Paul, Arnab↗

Velocity and flow angle measurements in the Langley 0.3-meter transonic cryogenic tunnel using a laser transit anemometer

The Laser Transit Anemometer (LTA) system is described. In the LTA system two parallel laser beams of known separation and cross sectional area are focussed at the same location or plane. When a particle in a flow field passes through both beams and the time is recorded for its transit (time of flight), its velocity can be calculated knowing the distance between the beams. By rotating the two beams (spots) around a common center and recording the number of valid events (a particle which passes through both spots in the proper sequence) at each angle the flow angle can be determined by curve fitting a predetermined number of angles or points and calculating the peak of what should be a Gaussian curve. The best angle or flow angle is defined as the angle at which the maximum number of valid events occurs. The LTA system functioned properly although conditions were less than desirable.

Honaker, W. C.↗

Parallel Battery: The Framework and Process for an Intelligent and Ecological Battery System and Related Services

The concept, framework, process methodology and applications of parallel battery were proposed from both virtual and real aspects.The parallel battery was an application of ACP-based parallel intelligence in battery and related energy system areas.The real battery system was running with its equivalent, and the artificial battery system was in a virtual space, in a parallel and interactive manner.The artificial battery system contained the descriptive, predictive, and prescriptive functions on the real battery and related energy systems.There was a closed-loop workflow between the real battery system and the artificial battery system, which iteratively optimizes the battery and related energy systems, leading to a new paradigm of intelligent and ecological parallel battery system management.

ACP approach↗

Grid Cyber-Security Strategy in an Attacker-Defender Model

The progression of cyber-attacks on the cyber-physical system is analyzed by the Probabilistic, Learning Attacker, and Dynamic Defender (PLADD) model. Although our research does apply to all cyber-physical systems, we focus on power grid infrastructure. The PLADD model evaluates the effectiveness of moving target defense (MTD) techniques. We consider the power grid attack scenarios in the AND configurations and OR configurations. In addition, we consider, for the first time ever, power grid attack scenarios involving both AND configurations and OR configurations simultaneously. Cyber-security managers can use the strategy introduced in this manuscript to optimize their defense strategies. Specifically, our research provides insight into when to reset access controls (such as passwords, internet protocol addresses, and session keys), to minimize the probability of a successful attack. Our mathematical proof for the OR configuration of multiple PLADD games shows that it is best if all access controls are reset simultaneously. For the AND configuration, our mathematical proof shows that it is best (in terms of minimizing the attacker's average probability of success) that the resets are equally spaced apart. We introduce a novel concept called hierarchical parallel PLADD system to cover additional attack scenarios that require combinations of AND and OR configurations.

97 MATHEMATICS AND COMPUTING↗