Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Erasure Codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

MLEC-Sim: A Simulator for Evaluating Multi-Level Erasure Coding

We present MLEC-Sim, a sophisticated simulator for Multi-Level Erasure Coding (MLEC), developed in approximately 13 KLOC. The simulator is engineered to analyze the impact of various system configurations and erasure coding policies on system durability and network overhead. It supports a comprehensive range of parameters including disk capacity, disk I/O bandwidth, failure rates, network bandwidth, and system scale, accommodating various erasure coding approaches such as Single-Level Erasure Coding (SLEC), Multi-Level Erasure Coding (MLEC), and Local Reconstruction Codes (LRC). MLEC-Sim provides support for multiple chunk placement policies, including clustered parity and declustered parity, and encompasses a variety of repair methods like Repair-ALL, Repair-FCO, Repair-HYB, and Repair-MIN. It is capable of simulating disk failures through a variety of means, including distribution-based or trace-based mechanisms, and can handle complex multi-level (de)clustered placements and repair processes. A key feature of MLEC-Sim is its adoption of the splitting simulation method for evaluating system durabilities at extremely high levels, which are challenging to assess with traditional simulation approaches. This feature allows for a detailed evaluation of system resilience under a range of conditions, aiding in the selection of appropriate erasure coding solutions for enhancing system durability. MLEC-Sim contributes to the field of data storage and reliability by providing a tool for the detailed evaluation of the durability and efficiency of erasure coding configurations, intended for use by researchers and practitioners in the design and optimization of storage systems.

Wang, Meng↗

Current possibilities and future opportunities for erasure coded computations

The key capability established through the research funded by this award are erasure coded computations for linear systems, in serial and in parallel. This capability enables powerful efficient and scalable alternatives to existing linear system solvers in fault-prone computational systems.

97 MATHEMATICS AND COMPUTING↗

Design Considerations and Analysis of Multi-Level Erasure Coding in Large-Scale Data Centers

Multi-level erasure coding (MLEC) has seen large deployments in the field, but there is no in-depth study of design considerations for MLEC at scale. In this paper, we provide comprehensive design considerations and analysis of MLEC at scale. We introduce the design space of MLEC in multiple dimensions, including various code parameter selections, chunk placement schemes, and various repair methods. We quantify their performance and durability, and show which MLEC schemes and repair methods can provide the best tolerance against independent/correlated failures and reduce repair network traffic by orders of magnitude. To achieve this, we use various evaluation strategies including simulation, splitting, dynamic programming, and mathematical modeling. We also compare the performance and durability of MLEC with other EC schemes such as SLEC and LRC and show that MLEC can provide high durability with higher encoding throughput and less repair network traffic over both SLEC and LRC.

Wang, Meng↗

JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows

In modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. In this paper, we propose Janus, a resilient and adaptable data transmission approach designed for cross-facility scientific workflows. Unlike traditional TCP-based methods, Janus leverages UDP, integrates erasure coding for fault tolerance, and combines it with error-bounded lossy compression to reduce overhead. This novel design allows users to balance data transmission time and accuracy, optimizing transfer performance based on specific scientific requirements. Additionally, Janus dynamically adjusts erasure coding parameters in response to real-time network conditions, ensuring efficient data transfers even in fluctuating environments. We develop optimization models for determining ideal configurations and implement adaptive data transfer protocols to enhance reliability. Through extensive simulations and real-network experiments, we demonstrate that Janus significantly improves transfer efficiency while maintaining data fidelity.

Esaulov, Vladislav [Georgia State University, Atla↗

RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific Data

In modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead.In this paper, we propose RAPIDS, a hybrid approach that combines the multigrid-based error-bounded lossy compression with erasure coding, to significantly reduce the storage and network overhead required for maintaining high data availability. Our experiments show that RAPIDS reduces the storage overhead by up to 7.5x and network overhead by up to 3x to achieve the same level of availability compared to the regular EC method. We improve RAPIDS by building two models to optimize the fault tolerance configurations and data gathering strategy. We demonstrate that RAPIDS significantly improves performance when running on many CPU cores in parallel or on GPUs.

Wan, Lipeng↗

Gradient Coding With Iterative Block Leverage Score Sampling

Gradient coding is a method for mitigating straggling servers in a centralized computing network that uses erasure-coding techniques to distributively carry out first-order optimization methods. Randomized numerical linear algebra uses randomization to develop improved algorithms for large-scale linear algebra computations. In this study, we propose a method for distributed optimization that combines gradient coding and randomized numerical linear algebra. The proposed method uses a randomized ℓ 2 -subspace embedding and a gradient coding technique to distribute blocks of data to the computational nodes of a centralized network, and at each iteration the central server only requires a small number of computations to obtain the steepest descent update. The novelty of our approach is that the data is replicated according to importance scores, called block leverage scores, in contrast to most gradient coding approaches that uniformly replicate the data blocks. Furthermore, we do not require a decoding step at each iteration, avoiding a bottleneck in previous gradient coding schemes. We show that our approach results in a valid ℓ 2 -subspace embedding, and that our resulting approximation converges to the optimal solution.

97 MATHEMATICS AND COMPUTING↗

Securely Aggregated Coded Matrix Inversion

Coded computing is a method for mitigating straggling workers in a centralized computing network, by using erasure-coding techniques. Federated learning is a decentralized model for training data distributed across client devices. In this work we propose approximating the inverse of an aggregated data matrix, where the data is generated by clients; similar to the federated learning paradigm, while also being resilient to stragglers. To do so, we propose a coded computing method based on gradient coding. We modify this method so that the coordinator does not access the local data at any point; while the clients access the aggregated matrix in order to complete their tasks. Here, the network we consider is not centrally administrated, and the communications which take place are secure against potential eavesdroppers.

97 MATHEMATICS AND COMPUTING↗

Federated Access from DOE Labs to Distributed Storage in the EIC Era of Computing

The Electron Ion Collider (EIC) collaboration and future experiment is a unique scientific ecosystem within Nuclear Physics as the experiment starts right off as a crosscollaboration between Brookhaven National Lab (BNL) & Jefferson Lab (JLab). As a result, this muti-lab computing model tries at best to provide services accessible from anywhere by anyone who is part of the collaboration. While the computing model for the EIC is not finalized, it is anticipated that the computational and storage resources will be made accessible to a wide range of collaborators across the world. The use of federated ID seems to be a critical element to the strategy of providing such services, allowing seamless access to each lab site computing resources. However, providing Federated access to a Federated storage is not a trivial matter and has its share of technical challenges. In this contribution, we focus on the steps we took towards the deployment of a distributed object storage system that integrates with Amazon S3 and Federated ID. We will first cover for and explain the first stage storage solutions provided to the EIC during the detector design phase. Our initial test deployment consisted of Lustre storage using MinIO, hence providing an S3 interface. High Availability load balancers were added later to provide the initial scalability it lacked. Performance of that system will be shown. While this embryonic solution worked well, it had many limitations. Looking ahead, the Ceph object storage is considered a top-of-the-line solution in the storage community - since the Ceph Object Gateway is compatible with the Amazon S3 API out of the box, our next phase will use a native S3 storage. Our Ceph deployment will consist of erasure coded storage nodes to maximize storage potential along with multiple Ceph Object Gateways for redundant access. We will compare performance of our next stage implementations. Finally, we will present how to leverage OpenID Connect with the Ceph Object Gateway’s to enable Federated ID access. We hope this contribution will serve the community needs as we move forward with cross-lab collaborations and the need for Federated ID access to distributed compute facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

ROOT RNTuple and EOS: The Next Generation of Event Data I/O

For several years, the ROOT team is developing the new RNTuple I/O subsystem in preparation of the next generation of collider experiments. Both HL-LHC and DUNE are expected to start data taking by the end of this decade. They pose unprecedented challenges to event data I/O in terms of data rates, event sizes, and event complexity. At the same time, the I/O landscape is becoming more diverse. HPC cluster file systems and object stores, NVMe disk cache layers in analysis facilities, and S3 storage on cloud resources are mixing with traditional XRootD-managed spinning disk pools.The ROOT team will finalize a first production version of the RNTuple binary format by the end of 2024. After this point, ROOT will provide backward compatibility for RNTuple data. This contribution provides an overview of the RNTuple feature set, the related R&D activities and the long-term vision for RNTuple. We report on performance, interface design, tooling, robustness, integration with experiment frameworks, and validation results, as well as recent R&D on parallel reading and writing and exploitation of modern hardware and storage systems. We will give an outlook on possible future features after a first production release.Collaboratively, the IT and EP departments at CERN have launched a formal project within the Research and Computing sector to evaluate the novel data format for physics analysis data utilized in LHC experiments and other fields. This part of the project focuses on validating the scalability of the EOS storage backend during the transition from the over 25 years old TTree production format to the newly developed RNTuple format, using both replicated and erasure-coded storage profiles.

Blomer, Jakob [CERN]↗

Data Processing Unit Services Module

The Data Processing Services Module (DPUSM) provides the ability to perform pluggable compression, erasure coding, checksuming and other important file system operations within the Linux kernel. The pluggable provider interface allows for the use of hardware acceleration of those services. In-kernel file systems are then able to use these functions to use these accelerators to perform operations that are normally run on the processor, resulting in improved file system performance. Third parties will register "providers" with the DPUSM to communicate with their respective accelerators. Providers will implement functions with DPUSM API signatures so that the DPUSM can translate the data inputted by users of the DPUSM into data that providers recognize.

Lee, Jason↗

Relational bulk reconstruction from modular flow

Abstract The entanglement wedge reconstruction paradigm in AdS/CFT states that for a bulk qudit within the entanglement wedge of a boundary subregion$$ \overline{A} $$ A ¯ , operators acting on the bulk qudit can be reconstructed as CFT operators on$$ \overline{A} $$ A ¯ . This naturally fits within the framework of quantum error correction, with the CFT states containing the bulk qudit forming a code protected against the erasure of the boundary subregionA. In this paper, we set up and study a framework for relational bulk reconstruction in holography: given two code subspaces both protected against erasure of the boundary regionA, the goal is to relate the operator reconstructions between the two spaces. To accomplish this, we assume that the two code subspaces are smoothly connected by a one-parameter family of codes all protected against the erasure ofA, and that the maximally-entangled states on these codes are all full-rank. We argue that such code subspaces can naturally be constructed in holography in a “measurement-based” setting. In this setting, we derive a flow equation for the operator reconstruction of a fixed code subspace operator using modular theory which can, in principle, be integrated to relate the reconstructed operators all along the flow. We observe a striking resemblance between our formulas for relational bulk reconstruction and the infinite-time limit of Connes cocycle flow, and take some steps towards making this connection more rigorous. We also provide alternative derivations of our reconstruction formulas in terms of a canonical reconstruction map we call the modular reflection operator.

Physics↗

Distributed Quantum Error Correction for Chip-Level Catastrophic Errors

Quantum error correction holds the key to scaling up quantum computers. Cosmic ray events severely impact the operation of a quantum computer by causing chip-level catastrophic errors, essentially erasing the information encoded in a chip. Here, in this work, we present a distributed error correction scheme to combat the devastating effect of such events by introducing an additional layer of quantum erasure error correcting code across separate chips. We show that our scheme is fault tolerant against chip-level catastrophic errors and discuss its experimental implementation using superconducting qubits with microwave links. Our analysis shows that in state-of-the-art experiments, it is possible to suppress the rate of these errors from 1 per 10 s to less than 1 per month.

97 MATHEMATICS AND COMPUTING↗

Covariant Quantum Error-Correcting Codes with Metrological Entanglement Advantage

Here, we show that a subset of the basis for the irreducible representations of a tensor-product SU(2) rotation forms a covariant approximate quantum error-correcting code with transversal U(1) logical gates. Generalizing previous work on “thermodynamic codes” to general local spin and different irreducible representations using only properties of the angular momentum algebra, we obtain bounds on the code inaccuracy under generic noise on any known 𝑑 sites, under independent and identically distributed noise, and under heralded 𝑑-local erasures. We demonstrate that this family of codes protects a probe state with quantum Fisher information surpassing the standard quantum limit when the sensing parameter couples to the generator of the U(1) logical gate.

quantum error correction↗

Leveraging Qubit Loss Detection in Fault-Tolerant Quantum Algorithms

Qubit loss errors constitute a dominant source of noise in many quantum hardware systems, particularly in neutral-atom quantum computers. We develop a theoretical framework to effectively detect and correct loss errors in logical algorithms and leverage such loss information in decoding. Considering general quantum error correction codes and logical circuits, we introduce a delayed-erasure decoder for experimentally motivated error models which leverages information from delayed loss detection to accurately correct loss errors, even when the precise moment of the error is unknown. Using this decoder, we identify strategies for detecting and correcting loss errors based on the logical circuit structure. For deep circuits prior to logical measurement, we explore methods to integrate loss detection into syndrome extraction with minimal overhead, identifying optimal strategies depending on the qubit loss fraction in the noise and hardware capabilities. In contrast, we find that many key algorithmic subroutines involve frequent gate teleportation, shortening the circuit depth before logical measurement and naturally replacing qubits with no additional experimental overhead. We simulate this setting using a toy model algorithm for small-angle synthesis and find a significant performance improvement as the loss fraction increases. These results provide a path forward for advancing large-scale fault-tolerant quantum computation in systems with loss error detection.

atoms↗

Quantum error correction in the black hole interior

We study the quantum error correction properties of the black hole interior in a toy model for an evaporating black hole: Jackiw-Teitelboim gravity entangled with a non-gravitational bath. After the Page time, the black hole interior degrees of freedom in this system are encoded in the bath Hilbert space. We use the gravitational path integral to show that the interior density matrix is correctable against the action of quantum operations on the bath which (i) do not have prior access to details of the black hole microstates, and (ii) do not have a large, negative coherent information with respect to the maximally mixed state on the bath, with the lower bound controlled by the black hole entropy and code subspace dimension. Thus, the encoding of the black hole interior in the radiation is robust against generic, low-rank quantum operations. For erasure errors, gravity comes within an O (1) distance of saturating the Singleton bound on the tolerance of error correcting codes. For typical errors in the bath to corrupt the interior, they must have a rank that is a large multiple of the bath Hilbert space dimension, with the precise coefficient set by the black hole entropy and code subspace dimension.

2D gravity↗

Information transmission with continuous variable quantum erasure channels

Quantum capacity, as the key figure of merit for a given quantum channel, upper bounds the channel's ability in transmitting quantum information. Identifying different types of channels, evaluating the corresponding quantum capacity, and finding the capacity-approaching coding scheme are the major tasks in quantum communication theory. Quantum channel in discrete variables has been discussed enormously based on various error models, while error model in the continuous variable channel has been less studied due to the infinite dimensional problem. In this paper, we investigate a general continuous variable quantum erasure channel. By defining an effective subspace of the continuous variable system, we find a continuous variable random coding model. We then derive the quantum capacity of the continuous variable erasure channel in the framework of decoupling theory. The discussion in this paper fills the gap of a quantum erasure channel in continuous variable setting and sheds light on the understanding of other types of continuous variable quantum channels.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Dual-rail encoding with superconducting cavities

The design of quantum hardware that reduces and mitigates errors is essential for practical quantum error correction (QEC) and useful quantum computation. To this end, we introduce the circuit-Quantum Electrodynamics (QED) dual-rail qubit in which our physical qubit is encoded in the single-photon subspace, { | 01 〉 , | 10 〉 } , of two superconducting microwave cavities. The dominant photon loss errors can be detected and converted into erasure errors, which are in general much easier to correct. In contrast to linear optics, a circuit-QED implementation of the dual-rail code offers unique capabilities. Using just one additional transmon ancilla per dual-rail qubit, we describe how to perform a gate-based set of universal operations that includes state preparation, logical readout, and parametrizable single and two-qubit gates. Moreover, first-order hardware errors in the cavities and the transmon can be detected and converted to erasure errors in all operations, leaving background Pauli errors that are orders of magnitude smaller. Hence, the dual-rail cavity qubit exhibits a favorable hierarchy of error rates and is expected to perform well below the relevant QEC thresholds with today’s coherence times.

97 MATHEMATICS AND COMPUTING↗