Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “checkpointing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗

A survey on checkpointing strategies: Should we always checkpoint à la Young/Daly?

The Young/Daly formula provides an approximation of the optimal checkpointing period for a parallel application executing on a supercomputing platform. It was originally designed to handle fail-stop errors for preemptible tightly-coupled applications, but has been extended to other application and resilience frameworks. Here, we provide some background and survey various scenarios to assess the usefulness and limitations of the formula, both for preemptible applications and workflow applications represented as a graph of tasks. We also discuss scenarios with uncertainties, and extend the study to silent errors. We exhibit cases where the optimal period is of a different order than that dictated by the Young/Daly formula, and finally we explain how checkpointing can be further combined with replication.

97 MATHEMATICS AND COMPUTING↗

Structural basis for the activity and specificity of the immune checkpoint inhibitor lirilumab

The clinical success of immune checkpoint inhibitors has underscored the key role of the immune system in controlling cancer. Current FDA-approved immune checkpoint inhibitors target the regulatory receptor pathways of cytotoxic T-cells to enhance their anticancer responses. Despite an abundance of evidence that natural killer (NK) cells can also mediate potent anticancer activities, there are no FDA-approved inhibitors targeting NK cell specific checkpoint pathways. Lirilumab, the most clinically advanced NK cell checkpoint inhibitor, targets inhibitory killer immunoglobulin-like receptors (KIRs), however it has yet to conclusively demonstrate clinical efficacy. Here we describe the crystal structure of lirilumab in complex with the inhibitory KIR2DL3, revealing the precise epitope of lirilumab and the molecular mechanisms underlying KIR checkpoint blockade. Notably, the epitope includes several key amino acids that vary across the human population, and binding studies demonstrate the importance of these amino acids for lirilumab binding. These studies reveal how KIR variations in patients could influence the clinical efficacy of lirilumab and reveal general concepts for the development of immune checkpoint inhibitors targeting NK cells.

60 APPLIED LIFE SCIENCES↗

Benchmarking Variables for Checkpointing in HPC Applications

Checkpoint/Restart (C/R) is a widely used fault tolerance mechanism in converged systems of cloud, edge, and HPC. However, users often rely on their experience to determine which variables to checkpoint, as there is currently no benchmark that can provide a reference. This can result in checkpointing redundant or even incorrect variables. To address this issue, we propose a benchmark suite that includes critical variables for checkpointing, which have been manually identified, and a method for identifying those critical variables, with 20 representative HPC applications. Our method involves analyzing data dependency between variables to identify critical variables analytically. We verify the identified variables' correctness with a widely used C/R library FTI by an ablation study. With our benchmark suite and data dependency analysis, HPC practitioners now have a reference for identifying checkpointing variables and better knowledge of what kind of variables to checkpoint.

Fu, Xiang↗

An Efficient Checkpointing System for Large Machine Learning Model Training

As machine learning models increase in size and complexity rapidly, the cost of checkpointing in ML training became a bottleneck in storage and performance (time). For example, the latest GPT-4 model has massive parameters at the scale of 1.76 trillion. It is highly time and storage consuming to frequently writes the model to checkpoints with more than 1 trillion floating point values to storage. This work aims to understand and attempt to mitigate this problem. First, we characterize the checkpointing interface in a collection of representative large machine learning/language models with respect to storage consumption and performance overhead. Second, we propose the two optimizations: i) A periodic cleaning strategy that periodically cleans up outdated checkpoints to reduce the storage burden; ii) A data staging optimization that coordinates checkpoints between local and shared file systems for performance improvement.

machine learning, artificial intelligence↗

Physics-aware adaptive checkpointing with shadow systems for nonlinear PDE simulations

Large-scale simulations of nonlinear partial differential equations (PDEs) that exhibit strongly transient behavior and pattern-forming dynamics produce enormous amounts of data, which, even with modern storage systems, cannot be stored for later curation. Current I/O strategies either write dense time series of snapshots, which is often prohibitive in I/O and storage, or store a few checkpoints that enable restart but incur expensive recomputation cost and provide no control over post-restart error growth, especially when lossy compression is used. Moreover, most, if not all, existing strategies take no account of the actual physical state of the system. Here, we present a simple physics-aware I/O framework in which a low-cost shadow system adaptively triggers lossy checkpoints when the shadow system deviates from the fine-scale simulation. The shadow system can be a coarsened replica of the fine-scale simulation that evolves concurrently. This means that checkpoints are taken based on the physical state of the system: fewer checkpoints are triggered when the system is quiescent while more are taken when the system undergoes a rapid change. This type of behavior is observed in many systems such as Brusselator and FitzHugh–Nagumo. We illustrate that our framework maintains stable restarts, keeps fine-scale restart errors bounded by shadow errors, and reconstructs the time history with significantly lower error and storage than interpolating fixed-interval snapshots, with low-cost shadow replay and modest online synchronization overhead.

Gong, Qian [ORNL] (ORCID:0000000235704142)↗

AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis

Checkpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 13 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint.

HPC↗

Scrutinizing Variables for Checkpoint Using Automatic Differentiation

Checkpoint/Restart (C/R) saves the running state of the programs periodically, which consumes considerable time and system resources. We observe that not every piece of data is involved in the computation in typical HPC applications; such unused data should be excluded from checkpointing for better storage and compute efficiency. We propose a systematic approach that leverages automatic differentiation (AD) to scrutinize every element within variables (e.g., arrays) necessary for checkpointing. This allows us to identify critical and uncritical elements and eliminate uncritical elements from checkpointing. Specifically, we inspect every single element within a variable necessary for checkpointing with an AD tool to determine whether the element has an impact on the application output or not. We validate our approach with all benchmarks from the NPB suite. We visualize the distribution of critical and uncritical elements within a variable with respect to its binary impact (yes or no) on the application output.

Huang, Xin [Kobe University]↗

Lossy checkpoint compression in full waveform inversion: a case study with ZFPv0.5.5 and the overthrust model

This paper proposes a new method that combines checkpointing methods with error-controlled lossy compression for large-scale high-performance full-waveform inversion (FWI), an inverse problem commonly used in geophysical exploration. This combination can significantly reduce data movement, allowing a reduction in run time as well as peak memory. In the exascale computing era, frequent data transfer (e.g., memory bandwidth, PCIe bandwidth for GPUs, or network) is the performance bottleneck rather than the peak FLOPS of the processing unit. Like many other adjoint-based optimization problems, FWI is costly in terms of the number of floating-point operations, large memory footprint during backpropagation, and data transfer overheads. Past work for adjoint methods has developed checkpointing methods that reduce the peak memory requirements during backpropagation at the cost of additional floating-point computations. Combining this traditional checkpointing with error-controlled lossy compression, we explore the three-way tradeoff between memory, precision, and time to solution. We investigate how approximation errors introduced by lossy compression of the forward solution impact the objective function gradient and final inverted solution. Empirical results from these numerical experiments indicate that high lossy-compression rates (compression factors ranging up to 100) have a relatively minor impact on convergence rates and the quality of the final solution.

58 GEOSCIENCES↗

Optimal checkpointing for adjoint multistage time-stepping schemes

Here, we consider checkpointing strategies that minimize the number of recomputations needed when performing discrete adjoint computations using multistage time-stepping schemes that require computing several substeps within one complete time step. Specifically, we propose two algorithms that can generate optimal checkpoint-ing schedules under weak assumptions. The first is an extension of the seminal Revolve algorithm adapted to multistage schemes. The second algorithm, named CAMS, is developed based on dynamic programming, and it requires the least number of recomputations when compared with other algorithms. The CAMS algorithm is made publicly available in a library with bindings to C and Python. Numerical results show that the proposed algorithms can deliver up to two times the speedup compared with that of classical Revolve. Moreover, we discuss the utilization of the CAMS library in mature scientific computing libraries and demonstrate the ease of using it in an adjoint workflow. The proposed algorithms have been adopted by the PETSc TSAdjoint library. Their performance has been demonstrated with a large-scale PDE-constrained optimization problem on a leadership-class supercomputer. This work is a significant extension of the authors' conference paper.

97 MATHEMATICS AND COMPUTING↗

Melanoma Cell Intrinsic GABAA Receptor Enhancement Potentiates Radiation and Immune Checkpoint Inhibitor Response by Promoting Direct and T Cell-Mediated Antitumor Activity

Most patients with metastatic melanoma show variable responses to radiation therapy and do not benefit from immune checkpoint inhibitors. Improved strategies for combination therapy that leverage potential benefits from radiation therapy and immune checkpoint inhibitors are critical.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Checkpoint kinases are required for oocyte meiotic progression by the maintenance of normal spindle structure and chromosome condensation

Highlights: • Inhibition of Chk1/2 has no significant effects on germinal vesicle breakdown. • Chk1/2 inhibition results in the first polar body extrusion defects. • Chk1/2 is critically involved in meiotic spindle organization. • Inhibition of Chk1/2 leads to abnormal chromosome condensation. • Chk1/2 is important for the development of MII stage oocytes. Checkpoint kinases (Chk) 1/2 are known for DNA damage checkpoint and cell cycle control in somatic cells. According to recent findings, the involvement of Chk1 in oocyte meiotic resumption and Chk2 is regarded as an essential regulator for progression at the post metaphase I stage (MI). In this study, AZD7762 (Chk1/2 inhibitor) and SB218078 (Chk1 inhibitor) were used to uncover the joint roles of Chk1/2 and differentiate the importance of Chk1 and Chk2 during oocyte meiotic maturation. Inhibition of Chk1/2 or Chk1 alone had no significant effect on germinal vesicle breakdown (GVBD) but significantly inhibited the first polar body (PB1). Interestingly, inhibition of Chk1 alone could not increase or completely block the extrusion of PB1 like Chk1/2 inhibition. Also, Chk1/2 inhibition resulted in defective meiotic spindle organization and chromosome condensation both in MI and metaphase II (MII) stages of oocytes. The location of γ-tubulin and Securin were abnormal or missing, while P38 MAPK was activated by Chk1/2 inhibition. Meanwhile, Chk1/2 inhibition reduced the percentage of the second polar body extrusion and pronuclear formation. In conclusion, our results further understand the functions and regulatory mechanism of Chk1/2 during oocyte meiotic maturation.

60 APPLIED LIFE SCIENCES↗

Immune checkpoint analysis of T‐cell responses to pp65 and IE‐1 antigens in end‐stage lung diseases

Abstract Lung transplant (LTX) patients are at high risk of cytomegalovirus (CMV) infection, which is often associated with high mortality and morbidity. Reactivation of CMV causes cell injury due to the cytopathic effect of viral replication and triggering of T cell immunity. The aim of this study was to compare expression of immune checkpoints (ICs) (PD‐1, CTLA‐4, LAG‐3 and TIGIT) in CD4, CD8 and CD56 and activation markers CD137, CD154 and CD69 of end‐stage patients awaiting lung transplant. Eighteen pre‐LTX positive for anti‐CMV IgG titres and 18 healthy subjects were enrolled. IC and activation markers have been evaluated through flow cytometric analysis in HC and pre‐LTX patients. Reactive (QF+) and unreactive (QF−) patients were stratified according to QuantiFERON‐CMV assays. ICs' and activation markers' expression were determined before and after in vitro stimulation with pp‐65 and IE‐1 antigens. Lower expression of PD‐1 was observed in CD4 and CD8 cells of pre‐LTX patients than controls, whereas CTLA4 appeared upregulated in CD56 and CD8 cells. TIGIT is increased on the surface of CD4, CD8 and NK cells after peptide stimulation in QF‐negative patients and PD‐1 is only downregulated after stimulation in the QF‐positive patients. This study provides new evidence of immune dysregulation in patients with end‐stage lung disorders, particularly in relation to immune checkpoint cell biology. The change in QF+ mostly happens on cytotoxic cells NK and CD8, while the changes in QF− were observed in adaptive immune cells, including CD4 and CD8.

Bergantini, Laura↗

Selected 3D Flash-X Checkpoints for the Long-Time Evolution of a 9.6 Solar-Mass Core-Collapse Supernova Model

This dataset contains selected 3D Flash-X checkpoint files from the long-time evolution of a low-energy core-collapse supernova explosion of a 9.6 Msun zero-metallicity, low-mass iron-core progenitor. The checkpoints span the shock-breakout phase through the young-remnant phase, ending at approximately 3 yr after core bounce. The dataset is intended for follow-up analysis and post-processing, especially radiation-transport calculations using the hydrodynamic and compositional structure of the ejecta. For a full description of the numerical setup, physical assumptions, limitations, and interpretation of the simulation, users should refer to the associated paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

A Novel DNA Repair Gene Signature for Immune Checkpoint Inhibitor-Based Therapy in Gastric Cancer

Gastric cancer is a heterogeneous group of diseases with only a fraction of patients responding to immunotherapy. The relationships between tumor DNA damage response, patient immune system and immunotherapy have recently attracted attention. Accumulating evidence suggests that DNA repair landscape is a significant factor in driving response to immune checkpoint blockade (ICB) therapy. In this study, to explore new prognostic and predictive biomarkers for gastric cancer patients who are sensitive and responsive to immunotherapies, we developed a novel 15-DNA repair gene signature (DRGS) and its related scoring system and evaluated the efficiency of the DRGS in discriminating different molecular and immune characteristics and therapeutic outcomes of patients with gastric adenocarcinoma, using publicly available datasets. The results demonstrated that DRGS high score patients showed significantly better therapeutic outcomes for ICB compared to DRGS low score patients (p < 0.001). Integrated analysis of multi-omics data demonstrated that the patients with high DRGS score were characteristic of high levels of anti-tumor lymphocyte infiltration, tumor mutation burden (TMB) and PD-L1 expression, and these patients exhibited a longer overall survival, as compared to the low-score patients. Results obtained from HPA and IHC supported significant dysregulation of the genes in DRGS in gastric cancer tissues, and a positive correlation in protein expression between DRGS and PD-L1. Therefore, the DRGS scoring system may have implications in tailoring immunotherapy in gastric cancers. A preprint has previously been published (Yuan et al., 2021).

60 APPLIED LIFE SCIENCES↗

An unsupervised machine-learning checkpoint-restart algorithm using Gaussian mixtures for particle-in-cell simulations

We propose an unsupervised machine-learning checkpoint-restart (CR) algorithm for particle-in-cell (PIC) algorithms using Gaussian mixtures (GM). The algorithm compresses the particle population per spatial cell by constructing a velocity distribution function using GM. Particles are reconstructed at restart time by local resampling of the Gaussians. To guarantee fidelity of the CR process, we ensure the exact preservation of invariants such as charge, momentum, and energy for both compression and reconstruction stages, everywhere on the mesh. We also ensure the preservation of Gauss' law after particle reconstruction by exactly matching the density profile at restart time. As a result, the GM CR algorithm is shown to provide a clean, conservative restart capability while potentially affording orders of magnitude savings in input/output requirements. Here, we demonstrate the algorithm using a recently developed exactly energy- and charge-conserving PIC algorithm using both electrostatic and electromagnetic tests. The tests demonstrate not only a high-fidelity CR capability, but also its potential for enhancing the fidelity of the PIC solution for a given particle resolution.

97 MATHEMATICS AND COMPUTING↗