Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Characterizing Machine Learning I/O Workloads on Leadership Scale HPC Systems

High performance computing (HPC) is no longer solely limited to traditional workloads such as simulation and modeling. With the increase in the popularity of machine learning (ML) and deep learning (DL) technologies, we are observing that an increasing number of HPC users are incorporating ML methods into their workflow and scientific discovery processes, across a wide spectrum of science domains such as biology, earth science, and physics. This gives rise to a diverse set of I/O patterns than the traditional checkpoint/restart-based HPC I/O behavior. The details of the I/O characteristics of such ML I/O workloads have not been studied extensively for large-scale leadership HPC systems. This paper aims to fill that gap by providing an in-depth analysis to gain an understanding of the I/O behavior of ML I/O workloads using darshan - an I/O characterization tool designed for lightweight tracing and profiling. We study the darshan logs of more than 23, 000 HPC ML I/O jobs over a time period of one year running on Summit - the second-fastest supercomputer in the world. This paper provides a systematic I/O characterization of ML I/O jobs running on a leadership scale supercomputer to understand how the I/O behavior differs across science domains and the scale of workloads, and analyze the usage of parallel file system and burst buffer by ML I/O workloads.

Paul, Arnab↗

Parallel Battery: The Framework and Process for an Intelligent and Ecological Battery System and Related Services

The concept, framework, process methodology and applications of parallel battery were proposed from both virtual and real aspects.The parallel battery was an application of ACP-based parallel intelligence in battery and related energy system areas.The real battery system was running with its equivalent, and the artificial battery system was in a virtual space, in a parallel and interactive manner.The artificial battery system contained the descriptive, predictive, and prescriptive functions on the real battery and related energy systems.There was a closed-loop workflow between the real battery system and the artificial battery system, which iteratively optimizes the battery and related energy systems, leading to a new paradigm of intelligent and ecological parallel battery system management.

ACP approach↗

Grid Cyber-Security Strategy in an Attacker-Defender Model

The progression of cyber-attacks on the cyber-physical system is analyzed by the Probabilistic, Learning Attacker, and Dynamic Defender (PLADD) model. Although our research does apply to all cyber-physical systems, we focus on power grid infrastructure. The PLADD model evaluates the effectiveness of moving target defense (MTD) techniques. We consider the power grid attack scenarios in the AND configurations and OR configurations. In addition, we consider, for the first time ever, power grid attack scenarios involving both AND configurations and OR configurations simultaneously. Cyber-security managers can use the strategy introduced in this manuscript to optimize their defense strategies. Specifically, our research provides insight into when to reset access controls (such as passwords, internet protocol addresses, and session keys), to minimize the probability of a successful attack. Our mathematical proof for the OR configuration of multiple PLADD games shows that it is best if all access controls are reset simultaneously. For the AND configuration, our mathematical proof shows that it is best (in terms of minimizing the attacker's average probability of success) that the resets are equally spaced apart. We introduce a novel concept called hierarchical parallel PLADD system to cover additional attack scenarios that require combinations of AND and OR configurations.

97 MATHEMATICS AND COMPUTING↗

Evaluating adaptive and predictive power management strategies for optimizing visualization performance on supercomputers

Power is becoming an increasingly scarce resource on the next generation of supercomputers, and should be used wisely to improve overall performance. One strategy for improving power usage is hardware overprovisioning, i.e., systems with more nodes than can be run at full power simultaneously without exceeding the system-wide power limit. With this study, we compare two strategies for allocating power throughout an overprovisioned system – adaptation and prediction – in the context of visualization workloads. While adaptation has been suitable for workloads with more regular execution behaviors, it may not be as suitable on visualization workloads, since they can have variable execution behaviors. This study considers a total of 104 experiments, which vary the rendering workload, power budget, allocation strategy, and node concurrency, including tests processing data sets up to 1 billion cells and using up to 18,432 cores across 512 nodes. Overall, we find that prediction is a superior strategy for this use case, improving performance up to 27% compared to an adaptive strategy.

97 MATHEMATICS AND COMPUTING↗

Kink mechanism in Cu/Nb nanolaminates explored by $\mathcal{in}$ $\mathcal{situ}$ pillar compression

We report Nano metallic laminates (NMLs) exhibit different failure modes depending on the loading conditions due to their mechanical anisotropies. Kinking is a typical failure mode in many NMLs compressed along a layer-parallel direction. However, a detailed description of the microstructure evolution during kink band (KB) formation and an in-depth understanding of the formation mechanisms are lacking. In this work, the KB process is investigated in Cu/Nb NMLs by in situ micro pillar compression in the scanning electron microscope (SEM) along a layer-parallel direction. Post-mortem S/TEM and transmission Kikuchi diffraction (TKD) analyses show that kink banding leads to significant microstructure changes characterized by an accumulation of geometrically necessary dislocations (GNDs) and of tilt geometrically necessary boundaries (GNBs) near KB boundaries (KBBs). The distinct microstructure evolution implies that KB formation is facilitated by the inhomogeneous microstructures resulting in constrained deformation modes. Specifically, dislocations active on slip planes nearly parallel to the interfaces make a major contribution to kink evolution after the onset of kinking. Once layer-parallel slip systems are activated, preexisting lattice dislocations and dislocations nucleating from interfaces will accumulate as GNDs near KBBs via the stochastic storage of lattice dislocations that have certain Burgers vectors. GNDs can further transform into GNBs via cross-slip and climb driven processes near the KBB. Furthermore, GNBs near KBBs can grow by incorporating more GNDs or by coalescence to accommodate the KB evolution. We further hypothesize that microstructural perturbations and their ensuing stresses can initiate KB formation in Cu/Nb NMLs.

36 MATERIALS SCIENCE↗

An innovative approach for atrazine electrochemical oxidation modelling: Process parameter effect, intermediate formation and kinetic constant assessment

Water reuse for irrigation activities is becoming a crucial worldwide challenge due to the depletion of water sources. Anyway, agricultural drainage can potentially contain dangerous contaminants such as metals, pesticides, and herbicides, including atrazine. To address the need for agriculture wastewater purification, we investigated atrazine removal from simulated wastewater by electro-oxidation using platinum-coated titanium electrodes on a lab-scale experimental apparatus. The effects of electrolyte composition and concentration, i.e. ionic strength and applied current density on atrazine removal, were investigated. The results demonstrated that the electrochemical oxidation of the herbicide occurred through two routes, depending on the presence or absence of oxidizing chlorine species. The generation of intermediates during the treatment was monitored and quantified by evaluating the effect of an inert electrolyte (NaClO 4 ) versus an oxidizable chlorine species (NaCl). In both experimental conditions, five intermediates were identified, including desethyl-atrazine (DEA), hydroxyatrazine (ATZ-OH), desisopropyl-atrazine (DIA) and desethyl-desisopropyl-atrazine (DEDIA). A degradation mechanism and a model for describing hydroxyl radicals and active chlorine species contributions at ATZ oxidation were also proposed. Intermediate evolution profiles suggest that ATZ degradation can be considered as a series–parallel reaction system. Finally, the energy requirement assessment for ATZ removal was carried out. The highest ATZ removal (≅98%) was achieved with NaCl = 0.08 M, J = 60 A/m −2 , and E C = 5.83 kWh m −3 . Results highlight that atrazine removal was improved when an active chlorine species (NaCl) was present in the water solution. Moreover, the addition of chlorine species during electro-oxidation is an energy-saving strategy. Collectively, electro-oxidation technique can be efficiently applied to treat polluted water in order to meet the needs of recycling water quality and reduce resource consumption.

Electro-chemical oxidation↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Large language models for batteries

Large Language Models (LLMs) are advanced artificial intelligence systems capable of solving diverse tasks using language, reasoning, and external tools. Despite their growing deployment in academia and industry, their potential remains underexplored in battery research. This review presents a comprehensive overview of existing and emerging applications of LLMs in batterie field, addressing two critical questions: What can LLMs offer to support battery-related tasks, and how to develop more effective models for this purpose. We begin by outlining the principles of LLMs and criteria for selecting appropriate models and tools for battery research and development. We then explore their roles in text-mining, data interpretation, and the development of intelligent battery systems. In parallel, we discuss technical challenges, such as data standardizing and sharing, model evaluation, and tool integration. Lastly, we propose future research directions with short-, medium-, and long-term goals and highlight more broad perspectives for connecting experts and cross-disciplinary collaborations.

SoC↗

Hvac: Removing I/O Bottleneck for Large-Scale Deep Learning Applications

Scientific communities are increasingly adopting deep learning (DL) models in their applications to accelerate scientific discovery processes. However, with rapid growth in the computing capabilities of HPC supercomputers, large-scale DL applications have to spend a significant portion of training time performing I/O to a parallel storage system. Previous research works have investigated optimization techniques such as prefetching and caching. Unfortunately, there exist non-trivial challenges to adopting the existing solutions on HPC supercomputers for large-scale DL training applications, which include non-performance and/or failures at extreme scale, lack of portability and generality in design, complex deployment methodology, and being limited to a specific application or dataset. To address these challenges, we propose High-Velocity AI Cache (HVAC), a distributed read-cache layer that targets and fully exploits the node-local storage or near node-local storage technology. HVAC seamlessly accelerates read I/O by aggregating node-local or near node-local storage, avoiding metadata lookups and file locking while preserving portability in the application code. We deploy and evaluate HVAC on 1,024 nodes (with over 6000 NVIDIA V100 GPUS) of the Summit supercomputer. In particular, we evaluate the scalability, efficiency, accuracy, and load distribution of HVAC compared to GPFS and XFS-on-NVMe. With four different DL applications, we observe an average 25 % performance improvement atop GPFS and 9% drop against XFS-on-NVMe, which scale linearly and are considered the performance upper bound. We envision HVAC as an important caching library for upcoming HPC supercomputers such as Frontier.

Khan, Awais↗

Communication-Avoiding and Memory-Constrained Sparse Matrix-Matrix Multiplication at Extreme Scale

Sparse matrix-matrix multiplication (SpGEMM) is a widely used kernel in various graph, scientific computing and machine learning algorithms. In this paper, we consider SpGEMMs performed on hundreds of thousands of processors generating trillions of nonzeros in the output matrix. Distributed SpGEMM at this extreme scale faces two key challenges: (1) high communication cost and (2) inadequate memory to generate the output. Furthermore, we address these challenges with an integrated communication-avoiding and memory-constrained SpGEMM algorithm that scales to 262,144 cores (more than 1 million hardware threads) and can multiply sparse matrices of any size as long as inputs and a fraction of output fit in the aggregated memory. As we go from 16,384 cores to 262,144 cores on a Cray XC40 supercomputer, the new SpGEMM algorithm runs 10x faster when multiplying large-scale protein-similarity matrices.

97 MATHEMATICS AND COMPUTING↗

TunIO: An AI-powered Framework for Optimizing HPC I/O

I/O operations are a known performance bottleneck of HPC applications. To achieve good performance, users often employ an iterative multistage tuning process to find an optimal I/O stack configuration. However, an I/O stack contains multiple layers, such as high-level I/O libraries, I/O middleware, and parallel file systems, and each layer has many parameters. These parameters and layers are entangled and influenced by each other. The tuning process is time-consuming and complex. In this work, we present TunIO, an AI-powered I/O tuning framework that implements several techniques to balance the tuning cost and performance gain, including tuning the high-impact parameters first. Furthermore, TunIO analyzes the application source code to extract its I/O kernel while retaining all statements necessary to perform I/O. It utilizes a smart selection of high-impact configuration parameters of the given tuning objective. Finally, it uses a novel Reinforcement Learning (RL)-driven early stopping mechanism to balance the cost and performance gain. Experimental results show that TunIO leads to a reduction of up to ≈73% in tuning time while achieving the same performance gain when compared to H5Tuner. It achieves a significant performance gain/cost of 208.4 MBps/min (I/O bandwidth for each minute spent in tuning) over existing approaches under our testing.

Rajesh, Neeraj↗

It’s Time to Talk About HPC Storage: Perspectives on the Past and Future

High-performance computing (HPC) storage systems are a key component of the success of HPC to date. Recently, we have seen major developments in storage-related technologies, as well as changes to how HPC platforms are used, especially in relation to artificial intelligence and experimental data analysis workloads. Additionally, these developments merit a revisit of HPC storage system architectural designs. In this article, we discuss the drivers, identify key challenges to status quo posed by these developments, and discuss directions future research might take to unlock the potential of new technologies for the breadth of HPC applications.

97 MATHEMATICS AND COMPUTING↗

Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints

Vast volumes of data are produced by today's scientific simulations and advanced instruments. These data cannot be stored and transferred efficiently because of limited I/O bandwidth, network speed, and storage capacity. Error-bounded lossy compression can be an effective method for addressing these issues: not only can it significantly reduce data size, but it can also control the data distortion based on user-defined error bounds. In practice, many scientific applications have specific requirements or constraints for lossy compression, in order to guarantee that the reconstructed data are valid for post hoc analysis. For example, some datasets contain irrelevant data that should be isolated in particular and users often have intuition regarding value ranges, geospatial regions, and other data subsets that are crucial for subsequent analysis. Existing state-of-the-art error-bounded lossy compressors, however, do not consider these constraints during compression, resulting in inferior compression ratios with respect to user's post hoc analysis, due to the fact that the data itself provides little or no value for post hoc analysis. In this work we address this issue by proposing an optimized framework that can preserve diverse constraints during the error-bounded lossy compression, e.g., cleaning the irrelevant data, efficiently preserving different precision for multiple value intervals, and allowing users to set diverse precision over both regular and irregular regions. We perform our evaluation on a supercomputer with up to 2,100 cores. Experiments with six real-world applications show that our proposed diverse constraints based error-bounded lossy compressor can obtain a higher visual quality or data fidelity on reconstructed data with the same or even higher compression ratios compared with the traditional state-of-the-art compressor SZ. Furthermore, our experiments also demonstrate very good scalability in compression performance compared with the I/O throughput of the parallel file system.

97 MATHEMATICS AND COMPUTING↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

A massively parallel and scalable multi-CPU material point method

Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100x per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.

Wang, Xinlei↗

PoliMOR

PoliMOR is a scalable, automated, and customizable policy engine framework for multi-tiered parallel file systems. It is composed of single-purpose agents that handle tasks such as gathering file metadata, making policy decisions, and then executing actions based on those policies. These agents are designed to communicate using distributed message queues, allowing the number of individual agents to be scaled up as needed. PoliMOR automates the data management tasks by precluding the need for admin intervention. The agents in PoliMOR can be customized to integrate any utilities/tools that perform tasks like metadata scanning and data placement management.

Brumgard, Christopher↗

Nodeman: A Node Management Tool For Hpc Clusters

NodeMan is a command line tool to manage nodes in an HPC cluster. At it's core, it is an extensible framework composed of bash scripting and GNU parallel. HPC System Administrator will find it useful in that it encapsulates desired functions and allows them to be assembled in a way familiar to administrators - through pipes. In fact, NodeMan functions can work with common command line tools as long as they use stdin/stdout. System Administrators can construct moderately complex logic and filtering on a compact command line that would normally require a substantial shell script. In the spirit of clush and pdsh, it is able to run commands remotely on nodes. Additionally, NodeMan is more flexible. For example, it can interact with IPMI and naturally processes node lists for orchestrating different tools. The library of useful pre-built functions is growing. System administrators can easily create new functions and make it their own.

Serr, ScottM↗