Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Checkpointing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Structural basis for proficient oxidized ribonucleotide insertion in double strand break repair

Reactive oxygen species (ROS) oxidize cellular nucleotide pools and cause double strand breaks (DSBs). Non-homologous end-joining (NHEJ) attaches broken chromosomal ends together in mammalian cells. Ribonucleotide insertion by DNA polymerase (pol) μ prepares breaks for end-joining and this is required for successful NHEJ in vivo. We previously showed that pol μ lacks discrimination against oxidized dGTP (8-oxo-dGTP), that can lead to mutagenesis, cancer, aging and human disease. Here we reveal the structural basis for proficient oxidized ribonucleotide (8-oxo-rGTP) incorporation during DSB repair by pol μ. Time-lapse crystallography snapshots of structural intermediates during nucleotide insertion along with computational simulations reveal substrate, metal and side chain dynamics, that allow oxidized ribonucleotides to escape polymerase discrimination checkpoints. Abundant nucleotide pools, combined with inefficient sanitization and repair, implicate pol μ mediated oxidized ribonucleotide insertion as an emerging source of widespread persistent mutagenesis and genomic instability.

60 APPLIED LIFE SCIENCES↗

rRNA methylation by Spb1 regulates the GTPase activity of Nog2 during 60S ribosomal subunit assembly

Biogenesis of the large ribosomal (60S) subunit involves the assembly of three rRNAs and 46 proteins, a process requiring approximately 70 ribosome biogenesis factors (RBFs) that bind and release the pre-60S at specific steps along the assembly pathway. The methyltransferase Spb1 and the K-loop GTPase Nog2 are essential RBFs that engage the rRNA A-loop during sequential steps in 60S maturation. Spb1 methylates the A-loop nucleotide G2922 and a catalytically deficient mutant strain ( spb 1 D52A ) has a severe 60S biogenesis defect. However, the assembly function of this modification is currently unknown. Here, we present cryo-EM reconstructions that reveal that unmethylated G2922 leads to the premature activation of Nog2 GTPase activity and capture a Nog2-GDP-AlF 4 - transition state structure that implicates the direct involvement of unmodified G2922 in Nog2 GTPase activation. Genetic suppressors and in vivo imaging indicate that premature GTP hydrolysis prevents the efficient binding of Nog2 to early nucleoplasmic 60S intermediates. We propose that G2922 methylation levels regulate Nog2 recruitment to the pre-60S near the nucleolar/nucleoplasmic phase boundary, forming a kinetic checkpoint to regulate 60S production. Our approach and findings provide a template to study the GTPase cycles and regulatory factor interactions of the other K-loop GTPases involved in ribosome assembly.

59 BASIC BIOLOGICAL SCIENCES↗

Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy

T cell states are prognostic in different cancer types. Recent technologies enable joint profiling of T cell RNA and T cell receptor (TCR) sequences at single-cell resolution. Here we present the TCR-RNA Integrating Model (TRIM), a multi-modal variational autoencoder framework that integrates RNA-TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. We applied TRIM to three independent datasets that included T cells collected before and after checkpoint inhibitor treatment, sourced either from blood and tumor biopsies in patients with head and neck squamous cell carcinoma and colorectal cancer, or from tumor and adjacent tissue in a pan-cancer dataset. In all settings, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment, demonstrating its utility in modeling multimodal T cell data and predicting T cell response to treatment and disease progression.

60 APPLIED LIFE SCIENCES↗

Heterologous synthesis of the complex homometallic cores of nitrogenase P- and M-clusters in Escherichia coli

Nitrogenase is an active target of heterologous expression because of its importance for areas related to agronomy, energy, and environment. One major hurdle for expressing an active Mo-nitrogenase in Escherichia coli is to generate the complex metalloclusters (P- and M-clusters) within this enzyme, which involves some highly unique bioinorganic chemistry/metalloenzyme biochemistry that is not generally dealt with in the heterologous expression of proteins via synthetic biology; in particular, the heterologous synthesis of the homometallic P-cluster ([Fe 8 S 7 ]) and M-cluster core (or L-cluster; [Fe 8 S 9 C]) on their respective protein scaffolds, which represents two crucial checkpoints along the biosynthetic pathway of a complete nitrogenase, has yet to be demonstrated by biochemical and spectroscopic analyses of purified metalloproteins. Here, we report the heterologous formation of a P-cluster-containing NifDK protein upon coexpression of Azotobacter vinelandii nifD, nifK, nifH, nifM, and nifZ genes, and that of an L-cluster-containing NifB protein upon coexpression of Methanosarcina acetivorans nifB, nifS, and nifU genes alongside the A. vinelandii fdxN gene, in E. coli. Our metal content, activity, EPR, and XAS/EXAFS data provide conclusive evidence for the successful synthesis of P- and L-clusters in a nondiazotrophic host, thereby highlighting the effectiveness of our metallocentric, divide-and-conquer approach that individually tackles the key events of nitrogenase biosynthesis prior to piecing them together into a complete pathway for the heterologous expression of nitrogenase. As such, this work paves the way for the transgenic expression of an active nitrogenase while providing an effective tool for further tackling the biosynthetic mechanism of this important metalloenzyme.

59 BASIC BIOLOGICAL SCIENCES↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

The genome of the polyextremophilic yeast, Naganishia friedmannii, reveals adaptations involved in stress response pathways, carbohydrate metabolism expansion, and a limited DNA repair repertoire

Here we report the draft genome sequence of Naganishia friedmannii (formerly Cryptococcus friedmannii) isolate, a Basidiomycota yeast commonly found in some of the most extreme environments of the Earth's cryosphere. We isolated N. friedmannii strain Llullensis from soils at 6000 m above sea level on Volcán Llullaillaco, Argentina. The genome was 22.2 Mb with 6251 identified protein coding genes. Proteins known to be associated with thermal, osmotic, and radiation stress were identified in the genome. Comparative analysis with seven other Naganishia genomes revealed unique features underlying its polyextremophilic lifestyle. Naganishia friedmannii showed an expansion of genes involved in breaking down plant-derived carbohydrates, supporting the hypothesis that it survives at high elevations by metabolizing wind-deposited organic matter. Surprisingly, many genes involved in cell-cycle checkpoints and DNA repair were missing, as in several other Naganishia species. This extensive loss may be adaptive in extreme environments prone to abiotic stress, where a high mutation rate could generate advantageous traits, and reduced cell-cycle control may allow for faster reproduction that would be advantageous for rapid growth during brief periods of soil wetting following rare snow events.

Vimercati, Lara↗

Pot1 promotes telomere DNA replication via the Stn1-Ten1 complex in fission yeast

Abstract Telomeres are nucleoprotein complexes that protect the chromosome-ends from eliciting DNA repair while ensuring their complete duplication. Pot1 is a subunit of telomere capping complex that binds to the G-rich overhang and inhibits the activation of DNA damage checkpoints. In this study, we explore new functions of fission yeast Pot1 by using a pot1-1 temperature sensitive mutant. We show that pot1 inactivation impairs telomere DNA replication resulting in the accumulation of ssDNA leading to the complete loss of telomeric DNA. Recruitment of Stn1 to telomeres, an auxiliary factor of DNA lagging strand synthesis, is reduced in pot1-1 mutants and overexpression of Stn1 rescues loss of telomeres and cell viability at restrictive temperature. We propose that Pot1 plays a crucial function in telomere DNA replication by recruiting Stn1-Ten1 and Polα-primase complex to telomeres via Tpz1, thus promoting lagging-strand DNA synthesis at stalled replication forks.

Carvalho Borges, Pâmela C. (ORCID:0000000244919874↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

UnifyFS: A User-level Shared File System for Unified Access to Distributed Local Storage

We introduce UnifyFS, a user-level file system that aggregates node-local storage tiers available on high performance computing (HPC) systems and makes them available to HPC applications under a unified namespace. UnifyFS employs transparent I/O interception, so it does not require changes to application code and is compatible with commonly used HPC I/O libraries. The design of UnifyFS supports the predominant HPC I/O workloads and is optimized for bulk-synchronous I/O patterns. Furthermore, UnifyFS provides customizable file system semantics to flexibly adapt its behavior for diverse I/O workloads and storage devices. In this paper, we discuss the unique design goals and architecture of UnifyFS and evaluate its performance on a leadership-class HPC system. In our experimental results, we demonstrate that UnifyFS exhibits excellent scaling performance for write operations and can improve the performance of application checkpoint operations by as much as 3× versus a tuned configuration.

Brim, Michael↗

Rotational Millimeter-Wave Shoe Scanner Using the Discrete Fourier Transform for Backprojection-Based Image Reconstruction

An active 3D microwave / millimeter-wave shoe scanner was previously developed at the Pacific Northwest National Laboratory (PNNL) using two linear arrays scanned over a rectilinear aperture. The radar system chirps a frequency sweep from 10-40 GHz. These frequencies allow imaging through optically opaque material such as leather, rubber, plastics, and other dielectrics. The system was designed to detect concealed items in the soles of shoes while allowing people to leave their shoes on through a security checkpoint. To shrink the footprint of the system, a new iteration of the design has been developed that scans the two linear arrays over a circular aperture. This new footprint opens the possibility of it being installed in the floor of a cylindrical millimeter-wave body scanner. The backprojection-based multilayer dielectric image reconstruction developed at PNNL can easily handle arbitrary spatial sampling, accommodating the new rotational shoe scanner design. Commonly, the fast Fourier transform (FFT) is used to efficiently compute the range response from the data collected by the system as a preprocessing step to the backprojection algorithm. It was found that converting to range using the discrete Fourier transform (DFT) directly has some advantages over the FFT. For example, nonlinear and non-uniform frequency sweeps can easily be compensated for during the computation of the DFT and only the range bins of interest need to be computed and their spacing can be chosen arbitrarily. Because the range conversion step of the image reconstruction is the fastest part of the process there is very little speed penalty for using the DFT over the FFT and it can even increase the speed of image reconstruction when the ranges of interest are fewer than the total span that is calculated in the FFT.

Millimeter-wave imaging, microwave imaging, shoe s↗

Phase IB study of ziv-aflibercept plus pembrolizumab in patients with advanced solid tumors

Background The combination of antiangiogenic agents with immune checkpoint inhibitors could potentially overcome immune suppression driven by tumor angiogenesis. We report results from a phase IB study of ziv-aflibercept plus pembrolizumab in patients with advanced solid tumors. Methods This is a multicenter phase IB dose-escalation study of the combination of ziv-aflibercept (at 2–4 mg/kg) plus pembrolizumab (at 2 mg/kg) administered intravenously every 2 weeks with expansion cohorts in programmed cell death protein 1 (PD-1)/programmed death-ligand 1(PD-L1)-naïve melanoma, renal cell carcinoma (RCC), microsatellite stable colorectal cancer (CRC), and ovarian cancer. The primary objective was to determine maximum tolerated dose (MTD) and recommended dose of the combination. Secondary endpoints included overall response rate (ORR) and overall survival (OS). Exploratory objectives included correlation of clinical efficacy with tumor and peripheral immune population densities. Results Overall, 33 patients were enrolled during dose escalation (n=3) and dose expansion (n=30). No dose-limiting toxicities were reported in the initial dose level. Ziv-aflibercept 4 mg/kg plus pembrolizumab 2 mg/kg every 2 weeks was established as the MTD. Grade ≥3 adverse events occurred in 19/33 patients (58%), the most common being hypertension (36%) and proteinuria (18%). ORR in the dose-expansion cohort was 16.7% (5/30, 90% CI 7% to 32%). Complete responses occurred in melanoma (n=2); partial responses occurred in RCC (n=1), mesothelioma (n=1), and melanoma (n=1). Median OS was as follows: melanoma, not reached (NR); RCC, 15.7 months (90% CI 2.5 to 15.7); CRC, 3.3 months (90% CI 0.6 to 3.4); ovarian, 12.5 months (90% CI 3.8 to 13.6); other solid tumors, NR. Activated tumor-infiltrating CD8 T cells at baseline (CD8+PD1+), high CD40L expression, and increased peripheral memory CD8 T cells correlated with clinical response. Conclusion The combination of ziv-aflibercept and pembrolizumab demonstrated an acceptable safety profile with antitumor activity in solid tumors. The combination is currently being studied in sarcoma and anti-PD-1-resistant melanoma. Trial registration number NCT02298959 .

Rahma, Osama E.↗

PETSc TSAdjoint: A Discrete Adjoint ODE Solver for First-Order and Second-Order Sensitivity Analysis

Here, we present a new software system PETSc TSAdjoint for first-order and second order adjoint sensitivity analysis of time-dependent nonlinear differential equations. The derivative calculation in PETSc TSAdjoint is essentially a high-level algorithmic differentiation process. The adjoint models are derived by differentiating the timestepping algorithms and implementing them based on the parallel infrastructure in PETSc. Full differentiation of the library code, including MPI routines, is avoided, and users do not need to derive their own adjoint models for their specific applications. PETSc TSAdjoint can compute the first-order derivative, that is, the gradient of a scalar functional, and the Hessian-vector product, which carries second-order derivative information, while requiring minimal input (a few callbacks) from the users. The adjoint model employs optimal checkpointing schemes in a manner that is transparent to users. Finally, usability, efficiency, and scalability are demonstrated through examples from a variety of applications.

79 ASTRONOMY AND ASTROPHYSICS↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

Enabling Command-and-Control in Advanced In Situ Workflows

Scientific discovery is progressing towards autonomous science with the combination of scientific instruments, high-performance computing, and artificial intelligence in complex workflows. This evolution introduces new requirements for managing scientific workflows, including feedback loops, near real-time constraints, and the ability to dynamically control workflow execution. In situ workflows that analyze and visualize data as it is generated are well-suited to satisfy stringent time constraints and their iterative nature offers greater opportunities for command-and-control. However, only a few of the many workflow management systems available have been specifically designed to manage in situ workflows and often lack support for automated feedback loops that allow analysis and visualization components to interact with the main scientific data producer. To address this need, we present in this paper how to add command-and-control capabilities to a workflow management system. We identify the functional design requirements of such a command-and-control system, detail its architecture, interface, and core mechanisms, and illustrate how advanced in situ workflows can leverage command-and-control in three use cases: graceful termination with checkpoint, dynamic and adaptive data reduction, and event-triggered analysis.

Mehta, Kshitij [ORNL] (ORCID:0000000297149981)↗

AEflow (Autoencoder fluid flow compression network) [SWR-22-29]

As the size of turbulent flow simulations continues to grow, in situ data compression is becoming increasingly important for visualization, analysis, and restart checkpointing. For these applications, single-pass compression techniques with low computational and communication overhead are crucial. In this paper we present a deep-learning approach to in situ compression using an autoencoder architecture that is customized for three-dimensional turbulent flows and is well suited for contemporary heterogeneous computing resources. The autoencoder is compared against a recently introduced randomized single-pass singular value decomposition (SVD) for three different canonical turbulent flows: decaying homogeneous isotropic turbulence, a Taylor-Green vortex, and turbulent channel flow. Our proposed fully convolutional autoencoder architecture compresses turbulent flow snapshots by a factor of 64 with a single pass, allows for arbitrarily sized input fields, is cheaper to compute than the randomized single-pass SVD for typical simulation sizes, performs well on unseen flow configurations, and has been made publicly available. The results reported here show that the autoencoder dramatically outperforms a randomized single-pass SVD with similar compression ratio and yields comparable performance to a higher-rank decomposition with an order of magnitude less compression in regard to preserving a number of important statistical quantities such as turbulent kinetic energy, enstrophy, and Reynolds stresses.

King, Ryan↗

ezAlign

The ezAlign is aimed at clustering coarse grain simulation to find common or uncommon occurrences and convert coarse (CG) grained coordinate and topology files to atomistic formats. We use a PointNet based approach to map individual frames of simulation to points in a latent space. These points are then clustered using a variety of clustering methods. Clusters are analyzed to associate them with states in the simulation. Frames can be chosen from these clusters based on proximity to cluster centers. ezAlign takes CG coordinate and topology files and converts and outputs their corresponding atomistic formats using an alignment and relaxation procedure. ezAlign is designed to convert complex, solvated biological systems including lipid membranes with drug-like molecules using GROMACS. A GROMACS checkpoint (.cpt) file is also outputted to enable continuation simulations that retain the equilibrated atomic velocities. Independent atomistic coordinates and topologies for every molecule must already be included in ezAlign/files. A number of commonly simulated biological molecules are currently provided.

Bennett, WilliamF.↗

SuperNu Version 4.x

We seek to release SuperNu, Version 4.x, as a continuation of development for the open source SuperNu software. The SuperNu, Version 3.x Monte Carlo radiative transfer code for astrophysical transients was released with GPLv3 copyright, asserted by LANL in 2015. For the next release we have features planned for development , including: opacity implementation (including non-local thermodynamic equilibrium effects), generalized source implementation (e.g. for emulating shock heating in Type II supernovae), 3T (electron, ion, radiation) internal energy update, special relativity corrections through O(v^2/c^2), light polarization (e.g. for comparison to spectroplarimetry observations of supernovae and kilonovae), and infrastructure features (checkpoint and restart of simulations, and tools including setup and Slurm batch scripts and simulation post-processing/analysis scripts). These features are intended to improve the fidelity and/or better understand uncertainty in supernova and kilonova light curve calculations.

Wollaeger, Ryan↗