Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Checkpointing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

Enabling Command-and-Control in Advanced In Situ Workflows

Scientific discovery is progressing towards autonomous science with the combination of scientific instruments, high-performance computing, and artificial intelligence in complex workflows. This evolution introduces new requirements for managing scientific workflows, including feedback loops, near real-time constraints, and the ability to dynamically control workflow execution. In situ workflows that analyze and visualize data as it is generated are well-suited to satisfy stringent time constraints and their iterative nature offers greater opportunities for command-and-control. However, only a few of the many workflow management systems available have been specifically designed to manage in situ workflows and often lack support for automated feedback loops that allow analysis and visualization components to interact with the main scientific data producer. To address this need, we present in this paper how to add command-and-control capabilities to a workflow management system. We identify the functional design requirements of such a command-and-control system, detail its architecture, interface, and core mechanisms, and illustrate how advanced in situ workflows can leverage command-and-control in three use cases: graceful termination with checkpoint, dynamic and adaptive data reduction, and event-triggered analysis.

Mehta, Kshitij [ORNL] (ORCID:0000000297149981)↗

AEflow (Autoencoder fluid flow compression network) [SWR-22-29]

As the size of turbulent flow simulations continues to grow, in situ data compression is becoming increasingly important for visualization, analysis, and restart checkpointing. For these applications, single-pass compression techniques with low computational and communication overhead are crucial. In this paper we present a deep-learning approach to in situ compression using an autoencoder architecture that is customized for three-dimensional turbulent flows and is well suited for contemporary heterogeneous computing resources. The autoencoder is compared against a recently introduced randomized single-pass singular value decomposition (SVD) for three different canonical turbulent flows: decaying homogeneous isotropic turbulence, a Taylor-Green vortex, and turbulent channel flow. Our proposed fully convolutional autoencoder architecture compresses turbulent flow snapshots by a factor of 64 with a single pass, allows for arbitrarily sized input fields, is cheaper to compute than the randomized single-pass SVD for typical simulation sizes, performs well on unseen flow configurations, and has been made publicly available. The results reported here show that the autoencoder dramatically outperforms a randomized single-pass SVD with similar compression ratio and yields comparable performance to a higher-rank decomposition with an order of magnitude less compression in regard to preserving a number of important statistical quantities such as turbulent kinetic energy, enstrophy, and Reynolds stresses.

King, Ryan↗

ezAlign

The ezAlign is aimed at clustering coarse grain simulation to find common or uncommon occurrences and convert coarse (CG) grained coordinate and topology files to atomistic formats. We use a PointNet based approach to map individual frames of simulation to points in a latent space. These points are then clustered using a variety of clustering methods. Clusters are analyzed to associate them with states in the simulation. Frames can be chosen from these clusters based on proximity to cluster centers. ezAlign takes CG coordinate and topology files and converts and outputs their corresponding atomistic formats using an alignment and relaxation procedure. ezAlign is designed to convert complex, solvated biological systems including lipid membranes with drug-like molecules using GROMACS. A GROMACS checkpoint (.cpt) file is also outputted to enable continuation simulations that retain the equilibrated atomic velocities. Independent atomistic coordinates and topologies for every molecule must already be included in ezAlign/files. A number of commonly simulated biological molecules are currently provided.

Bennett, WilliamF.↗

SuperNu Version 4.x

We seek to release SuperNu, Version 4.x, as a continuation of development for the open source SuperNu software. The SuperNu, Version 3.x Monte Carlo radiative transfer code for astrophysical transients was released with GPLv3 copyright, asserted by LANL in 2015. For the next release we have features planned for development , including: opacity implementation (including non-local thermodynamic equilibrium effects), generalized source implementation (e.g. for emulating shock heating in Type II supernovae), 3T (electron, ion, radiation) internal energy update, special relativity corrections through O(v^2/c^2), light polarization (e.g. for comparison to spectroplarimetry observations of supernovae and kilonovae), and infrastructure features (checkpoint and restart of simulations, and tools including setup and Slurm batch scripts and simulation post-processing/analysis scripts). These features are intended to improve the fidelity and/or better understand uncertainty in supernova and kilonova light curve calculations.

Wollaeger, Ryan↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

PRIME: Protein Representation Inference for Mutation Evaluation

Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.

Gibson, Kaetlyn [Los Alamos National Lab]↗

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

CodeScribe Agent

SF-26-086 CodeScribe introduces a structured, multi-stage pipeline that combines deterministic program analysis with LLM-powered translation to enable incremental, testable Fortran-to-C++ migration. First, `code-scribe index` traverses the project directory tree and produces `scribe.yaml` metadata files recording all modules, subroutines, and functions at each level, giving the LLM accurate structural context instead of a hallucinated codebase model. Second, `code-scribe draft` performs the deterministic portion of translation — converting Fortran types to C++ equivalents, replacing `use` statements with `#include` and `using namespace` directives, and detecting constructs requiring special handling — while embedding`scribe-prompt` annotations that guide the LLM through non-trivial cases such as statement-function-to-lambda conversions and `extern "C"` wrapper generation. Third, `code-scribe translate` applies project-specific TOML-based few-shot prompt templates and submits the composed prompt to a pluggable LLM backend (OpenAI, Anthropic, Argonne ARGO, any OpenAI-compatible endpoint, or local Hugging Face checkpoints), producing a C++ source file, a header, and a Fortran-C++ interface file for each translated routine so the codebase compiles and runs correctly throughout the migration. Beyond translation, CodeScribe includes a tool-using coding agent (`code-scribe agent`) with read, bash, edit, and write capabilities, and a bounded loop mode (`code-scribe loop`) that runs repeated stateless agent sessions over a task file with restricted tool access — enabling sustained, auditable software development workflows for broader scientific computing tasks.

Dhruv, Akash [Argonne National Laboratory (ANL), A↗

Colorectal Cancer Metastases in the Liver Establish Immunosuppressive Spatial Networking between Tumor-Associated SPP1 + Macrophages and Fibroblasts

Abstract Purpose: The liver is the most frequent metastatic site for colorectal cancer. Its microenvironment is modified to provide a niche that is conducive for colorectal cancer cell growth. This study focused on characterizing the cellular changes in the metastatic colorectal cancer (mCRC) liver tumor microenvironment (TME). Experimental Design: We analyzed a series of microsatellite stable (MSS) mCRCs to the liver, paired normal liver tissue, and peripheral blood mononuclear cells using single-cell RNA sequencing (scRNA-seq). We validated our findings using multiplexed spatial imaging and bulk gene expression with cell deconvolution. Results: We identified TME-specific SPP1-expressing macrophages with altered metabolism features, foam cell characteristics, and increased activity in extracellular matrix (ECM) organization. SPP1+ macrophages and fibroblasts expressed complementary ligand–receptor pairs with the potential to mutually influence their gene-expression programs. TME lacked dysfunctional CD8 T cells and contained regulatory T cells, indicative of immunosuppression. Spatial imaging validated these cell states in the TME. Moreover, TME macrophages and fibroblasts had close spatial proximity, which is a requirement for intercellular communication and networking. In an independent cohort of mCRCs in the liver, we confirmed the presence of SPP1+ macrophages and fibroblasts using gene-expression data. An increased proportion of TME fibroblasts was associated with the worst prognosis in these patients. Conclusions: We demonstrated that mCRC in the liver is characterized by transcriptional alterations of macrophages in the TME. Intercellular networking between macrophages and fibroblasts supports colorectal cancer growth in the immunosuppressed metastatic niche in the liver. These features can be used to target immune-checkpoint–resistant MSS tumors.

60 APPLIED LIFE SCIENCES↗

Novel Anti-LY6G6D/CD3 T-Cell–Dependent Bispecific Antibody for the Treatment of Colorectal Cancer

Abstract New therapeutics and combination regimens have led to marked clinical improvements for the treatment of a subset of colorectal cancer. Immune checkpoint inhibitors have shown clinical efficacy in patients with mismatch-repair–deficient or microsatellite instability–high (MSI-H) metastatic colorectal cancer (mCRC). However, patients with microsatellite-stable (MSS) or low levels of microsatellite instable (MSI-L) colorectal cancer have not benefited from these immune modulators, and the survival outcome remains poor for the majority of patients diagnosed with mCRC. In this article, we describe the discovery of a novel T-cell–dependent bispecific antibody (TDB) targeting tumor-associated antigen LY6G6D, LY6G6D-TDB, for the treatment of colorectal cancer. RNAseq analysis showed that LY6G6D was differentially expressed in colorectal cancer with high prevalence in MSS and MSI-L subsets, whereas LY6G6D expression in normal tissues was limited. IHC confirmed the elevated expression of LY6G6D in primary and metastatic colorectal tumors, whereas minimal or no expression was observed in most normal tissue samples. The optimized LY6G6D-TDB, which targets a membrane-proximal epitope of LY6G6D and binds to CD3 with high affinity, exhibits potent antitumor activity both in vitro and in vivo. In vitro functional assays show that LY6G6D-TDB–mediated T-cell activation and cytotoxicity are conditional and target dependent. In mouse xenograft tumor models, LY6G6D-TDB demonstrates antitumor efficacy as a single agent against established colorectal tumors, and enhanced efficacy can be achieved when LY6G6D-TDB is combined with PD-1 blockade. Our studies provide evidence for the therapeutic potential of LY6G6D-TDB as an effective treatment option for patients with colorectal cancer.

60 APPLIED LIFE SCIENCES↗

An African-Specific Variant of TP5 3 Reveals PADI4 as a Regulator of p53-Mediated Tumor Suppression

TP53 is the most frequently mutated gene in cancer, yet key target genes for p53-mediated tumor suppression remain unidentified. Here, we characterize a rare, African-specific germline variant of TP53 in the DNA-binding domain Tyr107His (Y107H). Nuclear magnetic resonance and crystal structures reveal that Y107H is structurally similar to wild-type p53. Consistent with this, we find that Y107H can suppress tumor colony formation and is impaired for the transactivation of only a small subset of p53 target genes; this includes the epigenetic modifier PADI4, which deiminates arginine to the nonnatural amino acid citrulline. Surprisingly, we show that Y107H mice develop spontaneous cancers and metastases and that Y107H shows impaired tumor suppression in two other models. We show that PADI4 is itself tumor suppressive and that it requires an intact immune system for tumor suppression. We identify a p53–PADI4 gene signature that is predictive of survival and the efficacy of immune-checkpoint inhibitors.

60 APPLIED LIFE SCIENCES↗

Inhibition of the eukaryotic initiation factor-2α kinase PERK decreases risk of autoimmune diabetes in mice

Preventing the onset of autoimmune type 1 diabetes (T1D) is feasible through pharmacological interventions that target molecular stress–responsive mechanisms. Cellular stresses, such as nutrient deficiency, viral infection, or unfolded proteins, trigger the integrated stress response (ISR), which curtails protein synthesis by phosphorylating eukaryotic translation initiation factor-2α (eIF2α). In T1D, maladaptive unfolded protein response (UPR) in insulin-producing β cells renders these cells susceptible to autoimmunity. We found that inhibition of the eIF2α kinase PKR-like ER kinase (PERK), a common component of the UPR and ISR, reversed the mRNA translation block in stressed human islets and delayed the onset of diabetes, reduced islet inflammation, and preserved β cell mass in T1D-susceptible mice. Single-cell RNA-Seq of islets from PERK-inhibited mice showed reductions in the UPR and PERK signaling pathways and alterations in antigen-processing and presentation pathways in β cells. Spatial proteomics of islets from these mice showed an increase in the immune checkpoint protein programmed death-ligand 1 (PD-L1) in β cells. Golgi membrane protein 1, whose levels increased following PERK inhibition in human islets and EndoC-βH1 human β cells, interacted with and stabilized PD-L1. Collectively, our studies show that PERK activity enhances β cell immunogenicity and that inhibition of PERK may offer a strategy for preventing or delaying the development of T1D.

Research & Experimental Medicine↗

The Helix-Loop-Helix motif of human EIF3A regulates translation of proliferative cellular mRNAs

Improper regulation of translation initiation, a vital checkpoint of protein synthesis in the cell, has been linked to a number of cancers. Overexpression of protein subunits of eukaryotic translation initiation factor 3 (eIF3) is associated with increased translation of mRNAs involved in cell proliferation. In addition to playing a major role in general translation initiation by serving as a scaffold for the assembly of translation initiation complexes, eIF3 regulates translation of specific cellular mRNAs and viral RNAs. Mutations in the N-terminal Helix-Loop-Helix (HLH) RNA-binding motif of the EIF3A subunit interfere with Hepatitis C Virus Internal Ribosome Entry Site (IRES) mediated translation initiation in vitro . Here we show that the EIF3A HLH motif controls translation of a small set of cellular transcripts enriched in oncogenic mRNAs, including MYC . We demonstrate that the HLH motif of EIF3A acts specifically on the 5' UTR of MYC mRNA and modulates the function of EIF4A1 on select transcripts during translation initiation. In Ramos lymphoma cell lines, which are dependent on MYC overexpression, mutations in the HLH motif greatly reduce MYC expression, impede proliferation and sensitize cells to anti-cancer compounds. These results reveal the potential of the EIF3A HLH motif in eIF3 as a promising chemotherapeutic target.

59 BASIC BIOLOGICAL SCIENCES↗

Student Programs FY21 Conversion Report

Sandia National Labs has created a noteworthy and effective internship program whose focus is creating a talent pipeline for the laboratory. Our program utilizes industry standard conversion calculations to examine the effectiveness of the program and to compare to our competitors. Sandia defines students eligible for conversion as graduating in the given fiscal year and in their final degree program. Students indicate to SIP upon hire and at certain checkpoints throughout their internship if they are in their final degree program or not. This means that they will not continue to a higher degree program after they graduate. For instance, someone who is graduating with a master’s degree in the current FY and does not plan to pursue a PhD would be considered eligible, while an undergrad student who is graduating in the same year, but plans to pursue a graduate degree, would not be considered eligible for conversion. Conversion data pulled for this report includes all eligible interns for fiscal year 2021. We use a rolling population, which includes anyone who was an intern at some point during FY21. To calculate conversion, we narrow our population down to the students who graduated between October 2020 through September 2021, who have indicated that they are in their final degree program. The conversion data was pulled on 10/29/2021, so any conversions completed after this date will not be included in the calculation. Our conversion data includes students who separated from Sandia and returned as a staff member. Conversions also include FTE, LTE, and postdoc positions. We do not include conversion to contractor positions in our calculations.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

STNS01-44 BEE – FY22-1: Enhanced BEE Client [Slide]

BEE provides a portable, modular, HPC-focused workflow engine capable of managing containerized applications at scale. In FY22 BEE is completing enhancements and refinements that will complete the major development work of the workflow system. The first major milestone is the development of a graphical client. The second milestone will be the ability for BEE to automatically restart checkpointed tasks. The final milestone for FY22 will be the ability to launch and manage multiple simultaneous workflows.

97 MATHEMATICS AND COMPUTING↗