Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Field validation of dynamic mechanical torque measurements using fiber-optic strain sensors for geared wind turbines

Abstract Accurate knowledge of the mechanical loads of wind turbine gearboxes has become essential in modern, highly loaded gearbox designs, as maintaining or even improving gearbox reliability with increasing torque density demands is proving to be challenging. Unfortunately, the traditional method of measuring dynamic mechanical torque using strain gauges placed on the outer surface of a rotating shaft and transmitting the resulting signal is unsuitable for serial deployment due to technical and economic constraints. An alternative method based on fiber-optic strain sensors placed on the stationary outer surface of the gearbox ring gear has been proposed. Like shaft torsion, the radial deformation of the ring gear is proportionate to the rotor torque. Placing the sensors on a stationary component is a cost-effective alternative for serial implementation because the need for complex and expensive data transfer via wireless transmission or a slip ring is eliminated. In this paper, we present the results of an extensive field experiment conducted to evaluate the torque measurement accuracy of this novel sensing solution installed on the gearbox of a Gamesa G97 2-MW wind turbine at the National Renewable Energy Laboratory’s Flatirons Campus. Torque measurements derived from fiber-optic strain sensors placed on the ring gear of the planetary stage are compared to conventional torque measurements from strain gauges placed on the main shaft. Two different torque estimation data processing methods were evaluated, with the method based on operational deflection shapes providing the most accurate results with an average normalized root mean square error below 0.7% for a load revolution distribution analysis. The effect of operating conditions on the torque estimate was also investigated, and the third planet-passing operational deflection shape was found to be the least sensitive to nontorque load-related effects. The fiber-optic strain sensors’ successful operation during the complete test campaign has demonstrated a robust and accurate solution for fleet-wide enhanced gearbox remaining useful life estimation.

17 WIND ENERGY↗

Field Validation of Dynamic Mechanical Torque Measurements for Geared Wind Turbines

Accurate knowledge of the mechanical loads of wind turbine gearboxes has become essential in modern, highly loaded gearbox designs, as maintaining or even improving gearbox reliability with increasing torque density demands is proving to be challenging. Unfortunately, the traditional method of measuring dynamic mechanical torque using strain gauges placed on the outer surface of a rotating shaft and transmitting the resulting signal is unsuitable for serial deployment due to technical and economic constraints. An alternative method based on fiber-optic strain sensors placed on the stationary outer surface of the gearbox ring gear has been proposed. Like shaft torsion, the radial deformation of the ring gear is proportionate to the rotor torque. Placing the sensors on a stationary component is a cost-effective alternative for serial implementation because the need for complex and expensive data transfer via wireless transmission or a slip ring is eliminated. In this paper, we present the results of an extensive field experiment conducted to evaluate the torque measurement accuracy of this novel sensing solution installed on the gearbox of a Gamesa G97 2-MW wind turbine at the National Renewable Energy Laboratory's Flatirons Campus. Torque measurements derived from fiber-optic strain sensors placed on the ring gear of the planetary stage are compared to conventional torque measurements from strain gauges placed on the main shaft. Two different torque estimation data processing methods were evaluated, with the method based on operational deflection shapes providing the most accurate results with an average normalized root mean square error below 0.7% for a load revolution distribution analysis. The effect of operating conditions on the torque estimate was also investigated, and the third planet-passing operational deflection shape was found to be the least sensitive to nontorque load-related effects. The fiber-optic strain sensors' successful operation during the complete test campaign has demonstrated a robust and accurate solution for fleet-wide enhanced gearbox remaining useful life estimation.

17 WIND ENERGY↗

Viability of S3 Object Storage for the ASC Program at Sandia

Recent efforts at Sandia such as DataSEA are creating search engines that enable analysts to query the institution’s massive archive of simulation and experiment data. The benefit of this work is that analysts will be able to retrieve all historical information about a system component that the institution has amassed over the years and make better-informed decisions in current work. As DataSEA gains momentum, it faces multiple technical challenges relating to capacity storage. From a raw capacity perspective, data producers will rapidly overwhelm the system with massive amounts of data. From an accessibility perspective, analysts will expect to be able to retrieve any portion of the bulk data, from any system on the enterprise network. Sandia’s Institutional Computing is mitigating storage problems at the enterprise level by procuring new capacity storage systems that can be accessed from anywhere on the enterprise network. These systems use the simple storage service, or S3, API for data transfers. While S3 uses objects instead of files, users can access it from their desktops or Sandia’s high-performance computing (HPC) platforms. S3 is particularly well suited for bulk storage in DataSEA, as datasets can be decomposed into object that can be referenced and retrieved individually, as needed by an analyst. In this report we describe our experiences working with S3 storage and provide information about how developers can leverage Sandia’s current systems. We present performance results from two sets of experiments. First, we measure S3 throughput when exchanging data between four different HPC platforms and two different enterprise S3 storage systems on the Sandia Restricted Network (SRN). Second, we measure the performance of S3 when communicating with a custom-built Ceph storage system that was constructed from HPC components. Overall, while S3 storage is significantly slower than traditional HPC storage, it provides significant accessibility benefits that will be valuable for archiving and exploiting historical data. There are multiple opportunities that arise from this work, including enhancing DataSEA to leverage S3 for bulk storage and adding native S3 support to Sandia’s IOSS library.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Robust Online Sequential RVFLNs for Data Modeling of Dynamic Time-Varying Systems with Application of an Ironmaking Blast Furnace

In a world where the increasing complexity of modern industrial processes brings difficulties for accurate mathematical modeling, taking advantage of data has become an efficient solution to complex dynamic process modeling issue. In this paper, we develop a novel robust online sequential version of random vector functional-link networks (RVFLNs) for data-driven modeling of dynamic time-varying system and applied it in a blast furnace (BF) ironmaking process. First, to overcome the time-varying dynamics of process and to enable the RVFLNs to learn online with avoiding data saturation, an improved online sequential version of RVFLNs (OS-RFVLNs) is first presented by online sequential learning with forgetting factor. This improved OS-RVFLNs algorithm is not only suitable for the real-time and large data transfer situation, but also can adjust the sensitivity of the algorithm to different samples with the help of the introduced forgetting factor. Second, since the output weights of the improved OS-RVFLNs as well as other RVFLNs algorithms are obtained by the least squares approach, a robustness problem may occur when the training dataset is contaminated with various outliers. To solve this problem, a Cauchy distribution weighted M-estimator is introduced to improve the robustness of the improved OS- RVFLNs. For this proposed robust OS-RVFLNs (R-OS- RVFLNs), since the weights of different outlier data are properly determined by the Cauchy distribution function, their corresponding contribution on modeling can be properly distinguished. Thus robust and better modeling results can be achieved. Experiments using actual industrial data of BF ironmaking process and comparative studies have demonstrated that the proposed method produces a better estimation accuracy and stronger robustness than other methods.

Blast furnace (BF), Modelling, Dynamic systems↗

Low power on-chip data transmission for wafer-scale monolithic active pixel sensors

Here, this paper details the implementation of the digital pulse shaping subsystem within the Backbone Transmission Line Encoding (BTLE) driver, a low-power, long-distance on-chip data transmission solution designed in a 65 nm CMOS process. Digital pulse shaping is critical for minimizing inter-symbol interference (ISI) caused by bandwidth limitations of on-chip interconnects, especially in wafer-scale monolithic active pixel sensors (MAPS). A duobinary encoder coupled with a parallelized polyphase finite impulse response (FIR) filter is used for efficient shaping of the transmitted signal spectrum. This reconfigurable architecture achieves reliable 160 Mb/s data transfer over a 10 cm on-chip link, as validated by simulations demonstrating low power consumption (FoM 37.3 fJ/bit/mm of transmission line length) and effective ISI mitigation.

47 OTHER INSTRUMENTATION↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Modeling and Simulation of Fuel Dispersal During the Loss-of-Coolant Accident

This document is the compilation of the milestone portion to a larger end of project NEUP report. The executive summary of the modeling portion is provided below: In the event of cladding rupture during a postulated LOCA in a pressurized water reactor, fuel particles, along with fission gases, can be expelled into the reactor core from the fractured fuel rod, a phenomenon referred to as fuel dispersal. The initial stage of fuel dispersal is strongly influenced by the high-pressure ejection of fuel fragments, the size and geometry of the ruptured cladding, and the depressurization history of the fuel rod during the postulated LOCA transient. Depending on the location of the burst orifice relative to the quench front, the dispersal event represents an intricate three-phase flow and heat transfer phenomenon, where high-temperature fuel particles carried by the fission gases interact with the coolant within the narrow subchannels of the fuel assemblies, inducing localized phase change. Given the unique multiphysics nature of this phenomena, the current study develops a dedicated computational framework to predict the mass distribution and cooling of dispersing fuel particles, facilitating post-accident assessment and management of the fuel assemblies. Considering the scale of nuclear reactor applications, a continuum three-fluid model is proposed for simulating the transport of solids within the reactor core. With high-temperature fuel fragments within the liquid media, nucleation sites inducing phase changes are dispersed within the flow domain. Coupled with the fact that the transient dispersal event occurs on different time scales than other three-phase flow applications, this study derives a time-averaged three-fluid flow model without losing generality. The assumptions regarding the continuum treatment of the solid phase and the modeling of fuel dispersal behavior are incorporated to simplify the governing equations and derive applicable closure relations. The computational validation of the model was conducted using adiabatic experimental results obtained from ongoing research at Oregon State University, focusing on characterizing fuel dispersal behavior during simulated LOCA conditions. Settlement characteristics of the solids, quantified by the probability distribution of equivalent particles, closely matched the probability density functions reported in experimental studies. The transport of fuel particles within a scaled 5 × 5 lattice of a pressurized-water reactor rod bundle geometry was modeled through a two-fluid Eulerian framework. The required boundary conditions were evaluated from the fuel performance code BISON in a postulated large-break LOCA scenario. The modeling framework considered solid fuel particles as granular matter, interacting with the gaseous dry steam phase and fission gases through the governing interfacial momentum exchange between the participating fluids. The simulation results provided the volume fraction of the solids obtained at the bottom surface of the enclosing tank geometry. Postulated LOCA leading to fuel dispersal phenomena involves the strong coupling between fuel thermomechanics, cladding deformation, thermal-hydraulics, and fuel particle transport. Incorporation of such a strong coupling in numerical simulation is performed by coupling the multiphysics solvers. In the case of fuel dispersal, a strong coupled simulation can be performed by coupling the BISON code for fuel performance, the TRACE code for system-level thermal hydraulics, and fuel particle transport in Multiphysics Object-Oriented Simulation Environment (MOOSE). For such intricate infrastructure, the MOOSE Framework eases the data transfer between codes. The recent version of MOOSE has incorporated the Navier-Stokes module for the fluid flow. An exploratory exercise was done to gain familiarity with finite volume capabilities in the MOOSE framework to incorporate the Spalart-Allmaras (SA) turbulence model. New finite-volume and auxiliary kernels were introduced to assemble the SA transport equation, compute turbulent viscosity, and evaluate wall distance and diagnostic turbulence terms, fully integrated with existing Navier-Stokes modules. A turbulent lid-driven cavity at a Reynolds number of approximately 10,000 is used for verification. MOOSE shows the robust solver convergence and produces the turbulent features. But it underpredicts the velocity profile and turbulent quantities, emphasizing the need to develop improved SA near-wall treatments (e.g., low-Re corrections or wall functions) as a key direction for future work.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A prototype scintillator real‐time beam monitor for ultra‐high dose rate radiotherapy

Background: FLASH Radiotherapy (RT) is an emergent cancer RT modality where an entire therapeutic dose is delivered at more than 1000 times higher dose rate than conventional RT. For clinical trials to be conducted safely, a precise and fast beam monitor that can generate out-of-tolerance beam interrupts is required. This paper describes the overall concept and provides results from a prototype ultra-fast, scintillator-based beam monitor for both proton and electron beam FLASH applications. Purpose: A FLASH Beam Scintillator Monitor (FBSM) is being developed that employs a novel proprietary scintillator material. The FBSM has capabilities that conventional RT detector technologies are unable to simultaneously provide: (1) large area coverage; (2) a low mass profile; (3) a linear response over a broad dynamic range; (4) radiation hardness; (5) real-time analysis to provide an IEC-compliant fast beam-interrupt signal based on true two-dimensional beam imaging, radiation dosimetry and excellent spatial resolution. Methods: The FBSM uses a proprietary low mass, less than 0.5 mm water equivalent, non-hygroscopic, radiation tolerant scintillator material (designated HM: hybrid material) that is viewed by high frame rate CMOS cameras. Folded optics using mirrors enable a thin monitor profile of ∼10 cm. A field programmable gate array (FPGA) data acquisition system generates real-time analysis on a time scale appropriate to the FLASH RT beam modality: 100–1000 Hz for pulsed electrons and 10–20 kHz for quasi-continuous scanning proton pencil beams. An ion beam monitor served as the initial development platform for this work and was tested in low energy heavy-ion beams ( 86 Kr +26 and protons). A prototype FBSM was fabricated and then tested in various radiation beams that included FLASH level dose per pulse electron beams, and a hospital RT clinic with electron beams. Results: Results presented in this report include image quality, response linearity, radiation hardness, spatial resolution, and real-time data processing. Furthermore, the HM scintillator was found to be highly radiation damage resistant. It exhibited a small 0.025%/kGy signal decrease from a 216 kGy cumulative dose resulting from continuous exposure for 15 min at a FLASH compatible dose rate of 237 Gy/s. Measurements of the signal amplitude versus beam fluence demonstrate linear response of the FBSM at FLASH compatible dose rates of >40 Gy/s. Comparison with commercial Gafchromic film indicates that the FBSM produces a high resolution 2D beam image and can reproduce a nearly identical beam profile, including primary beam tails. The spatial resolution was measured at 35–40 µm. Tests of the firmware beta version show successful operation at 20 000 Hz frame rate or 50 µs/frame, where the real-time analysis of the beam parameters is achieved in less than 1 µs. Conclusions: The FBSM is designed to provide real-time beam profile monitoring over a large active area without significantly degrading the beam quality. A prototype device has been staged in particle beams at currents of single particles up to FLASH level dose rates, using both continuous ion beams and pulsed electron beams. Using a novel scintillator, beam profiling has been demonstrated for currents extending from single particles to 10 nA currents. Radiation damage is minimal and even under FLASH conditions would require ≥50 kGy of accumulated exposure in a single spot to result in a 1% decrease in signal output. Beam imaging is comparable to radiochromic films, and provides immediate images without hours of processing. Real-time data processing, taking less than 50 µs (combined data transfer and analysis times), has been implemented in firmware for 20 kHz frame rates for continuous proton beams.

2D beam imaging↗

Cross Inference of Throughput Profiles Using Micro Kernel Network Method

Dedicated network connections are being increasingly deployed in cloud, centralized and edge computing and data infrastructures, whose throughput profiles are critical indicators of the underlying data transfer performance. Due to the cost and disruptions to physical infrastructures, network emulators, such as Mininet, are often used to generate measurements needed to estimate throughput profiles, typically expressed as a function of the connection round trip time. The profiles estimated using measurements from such emulated networks are usually inaccurate for high bandwidth and high latency connections, since they do not accurately reflect the critical network transport dynamics mainly due to computing and memory constraints of the host. We present a machine learning (ML) method to estimate the throughput profiles using emulation measurements to closely match the testbed and production network profiles. In particular, we propose a micro Kernel Network (mKN) that provides baseline throughput measurements on the host running Mininet emulations, which are used to learn a regression map that converts them to the corresponding testbed measurement estimates. Once initially learned, this map is applied to measurements from subsequent network emulations on the same host. We present experimental measurements to illustrate this approach, and derive generalization equations for the proposed mKN-ML method. Using a four-site scenario emulation, we show the effectiveness of this method in providing accurate concave throughput profiles from inaccurate convex or non-smooth ones indicated by Mininet emulation.

Rao, Nageswara↗

IRIS Reimagined: Advancements in Intelligent Runtime System for Task-Based Programming

Task-based programming models are gaining traction in scientific computing. IRIS is a portable runtime system that exploits multiple heterogeneous programming systems and can discover available resources and manage multiple diverse programming systems (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, and OpenMP) simultaneously. It accounts for the constraints of task dependencies and provides customizable scheduling policies to map those tasks to heterogeneous devices. In this paper, we present new capabilities added to IRIS to improve its portability for heterogeneous programming, build-friendliness, and performance efficiency. The new additions include vendor-specific kernel support, a runtime system with a foreign function interface to eliminate writing wrapper or boilerplate code for heterogeneous kernels, an easy-to-use and configurable CMake-based build environment, automatic and efficient data transfers and orchestration, and the Hunter and DAGGER toolchains to evaluate IRIS’s task scheduling algorithms.

Miniskar, Narasinga Rao↗

MatRIS: Addressing the Challenges for Portability and Heterogeneity Using Tasking for Matrix Decomposition (Cholesky)

The ubiquitous in-node heterogeneity of HPC and cloud computing platforms makes software portability and performance optimization extremely challenging. Described here, the MatRIS multilevel math library abstraction framework employs tasking to alleviate these difficulties. MatRIS includes the IRIS task-based runtime on the bottom level and exposes different layers of abstraction to render algorithms architecturally agnostic. MatRIS ensures the decomposition and creation of tasks that represent the necessary encapsulation of the optimized kernels from both vendor and open-source math libraries. Once built, MatRIS can select different combinations of accelerators at runtime, making it portable even on diverse heterogeneous architectures. By leveraging the IRIS runtime’s features for managing heterogeneity, MatRIS deploys algorithms that remove the need to specify orchestration and data transfer. This study describes how the serial task abstraction of a tiled Cholesky factorization is made portable and scalable in the case of multi-device and multi-vendor heterogeneity on a node with NVIDIA and AMD GPUs by using MatRIS. First, we demonstrate that Cholesky in MatRIS provides multi-GPU scalability that offers competitive performance versus cuSolverMG. Then, we present the challenges and opportunities for heterogeneous execution.

Monil, M. A. H.↗

Distributed out-of-memory NMF on CPU/GPU architectures

We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density 10 -6 .

97 MATHEMATICS AND COMPUTING↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

Accelerating Machine Learning Inference with GPUs in ProtoDUNE Data Processing

Abstract We study the performance of a cloud-based GPU-accelerated inference server to speed up event reconstruction in neutrino data batch jobs. Using detector data from the ProtoDUNE experiment and employing the standard DUNE grid job submission tools, we attempt to reprocess the data by running several thousand concurrent grid jobs, a rate we expect to be typical of current and future neutrino physics experiments. We process most of the dataset with the GPU version of our processing algorithm and the remainder with the CPU version for timing comparisons. We find that a 100-GPU cloud-based server is able to easily meet the processing demand, and that using the GPU version of the event processing algorithm is two times faster than processing these data with the CPU version when comparing to the newest CPUs in our sample. The amount of data transferred to the inference server during the GPU runs can overwhelm even the highest-bandwidth network switches, however, unless care is taken to observe network facility limits or otherwise distribute the jobs to multiple sites. We discuss the lessons learned from this processing campaign and several avenues for future improvements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

How fast can one resize a distributed file system?

Efficient resource utilization becomes a major concern as large-scale distributed computing infrastructures keep growing in size. Malleability, the possibility for resource managers to dynamically increase or decrease the amount of resources allocated to a job, is a promising way to save energy and costs. However, state-of-the-art parallel and distributed storage systems have not been designed with malleability in mind. The reason is mainly the supposedly high cost of data transfers required by resizing operations. Nevertheless, as network and storage technologies evolve, old assumptions about potential bottlenecks can be revisited. In this study, we evaluate the viability of malleability as a design principle for a distributed storage system. We specifically model the minimal duration of the commission and decommission operations. To show how our models can be used in practice, we evaluate the performance of these operations in HDFS, a relevant state-of-the-art distributed file system. We show that the existing decommission mechanism of HDFS is good when the network is the bottleneck, but can be accelerated by up to a factor 3 when storage is the limiting factor. We also show that the commission in HDFS can be substantially accelerated. With the highlights provided by our model, we suggest improvements to speed both operations in HDFS. We discuss how the proposed models can be generalized for distributed file systems with different assumptions and what perspectives are open for the design of efficient malleable distributed file systems.

97 MATHEMATICS AND COMPUTING↗

Fast shared-memory streaming multilevel graph partitioning

In this report we show that a fast parallel graph partitioner can benefit many applications by reducing data transfers. The online methods for partitioning graphs have to be fast and they often rely on simple one-pass streaming algorithms, while the offline methods for partitioning graphs contain more involved algorithms and the most successful methods in this category belong to the multilevel approaches. In this work, we assess the feasibility of using streaming graph partitioning algorithms within the multilevel framework. Our end goal is to come up with a fast parallel offline multilevel partitioner that can produce competitive cutsize quality. We rely on a simple but fast and flexible streaming algorithm throughout the entire multilevel framework. This streaming algorithm serves multiple purposes in the partitioning process: a clustering algorithm in the coarsening, an effective algorithm for the initial partitioning, and a fast refinement algorithm in the uncoarsening. Its simple nature also lends itself easily for parallelization. The experiments on various graphs show that our approach is on the average up to 5.1x faster than the multi-threaded MeTiS, which comes at the expense of only 2x worse cutsize.

97 MATHEMATICS AND COMPUTING↗