Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Checkpointing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Pot1 promotes telomere DNA replication via the Stn1-Ten1 complex in fission yeast

Abstract Telomeres are nucleoprotein complexes that protect the chromosome-ends from eliciting DNA repair while ensuring their complete duplication. Pot1 is a subunit of telomere capping complex that binds to the G-rich overhang and inhibits the activation of DNA damage checkpoints. In this study, we explore new functions of fission yeast Pot1 by using a pot1-1 temperature sensitive mutant. We show that pot1 inactivation impairs telomere DNA replication resulting in the accumulation of ssDNA leading to the complete loss of telomeric DNA. Recruitment of Stn1 to telomeres, an auxiliary factor of DNA lagging strand synthesis, is reduced in pot1-1 mutants and overexpression of Stn1 rescues loss of telomeres and cell viability at restrictive temperature. We propose that Pot1 plays a crucial function in telomere DNA replication by recruiting Stn1-Ten1 and Polα-primase complex to telomeres via Tpz1, thus promoting lagging-strand DNA synthesis at stalled replication forks.

Carvalho Borges, Pâmela C. (ORCID:0000000244919874↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

UnifyFS: A User-level Shared File System for Unified Access to Distributed Local Storage

We introduce UnifyFS, a user-level file system that aggregates node-local storage tiers available on high performance computing (HPC) systems and makes them available to HPC applications under a unified namespace. UnifyFS employs transparent I/O interception, so it does not require changes to application code and is compatible with commonly used HPC I/O libraries. The design of UnifyFS supports the predominant HPC I/O workloads and is optimized for bulk-synchronous I/O patterns. Furthermore, UnifyFS provides customizable file system semantics to flexibly adapt its behavior for diverse I/O workloads and storage devices. In this paper, we discuss the unique design goals and architecture of UnifyFS and evaluate its performance on a leadership-class HPC system. In our experimental results, we demonstrate that UnifyFS exhibits excellent scaling performance for write operations and can improve the performance of application checkpoint operations by as much as 3× versus a tuned configuration.

Brim, Michael↗

Rotational Millimeter-Wave Shoe Scanner Using the Discrete Fourier Transform for Backprojection-Based Image Reconstruction

An active 3D microwave / millimeter-wave shoe scanner was previously developed at the Pacific Northwest National Laboratory (PNNL) using two linear arrays scanned over a rectilinear aperture. The radar system chirps a frequency sweep from 10-40 GHz. These frequencies allow imaging through optically opaque material such as leather, rubber, plastics, and other dielectrics. The system was designed to detect concealed items in the soles of shoes while allowing people to leave their shoes on through a security checkpoint. To shrink the footprint of the system, a new iteration of the design has been developed that scans the two linear arrays over a circular aperture. This new footprint opens the possibility of it being installed in the floor of a cylindrical millimeter-wave body scanner. The backprojection-based multilayer dielectric image reconstruction developed at PNNL can easily handle arbitrary spatial sampling, accommodating the new rotational shoe scanner design. Commonly, the fast Fourier transform (FFT) is used to efficiently compute the range response from the data collected by the system as a preprocessing step to the backprojection algorithm. It was found that converting to range using the discrete Fourier transform (DFT) directly has some advantages over the FFT. For example, nonlinear and non-uniform frequency sweeps can easily be compensated for during the computation of the DFT and only the range bins of interest need to be computed and their spacing can be chosen arbitrarily. Because the range conversion step of the image reconstruction is the fastest part of the process there is very little speed penalty for using the DFT over the FFT and it can even increase the speed of image reconstruction when the ranges of interest are fewer than the total span that is calculated in the FFT.

Millimeter-wave imaging, microwave imaging, shoe s↗

Phase IB study of ziv-aflibercept plus pembrolizumab in patients with advanced solid tumors

Background The combination of antiangiogenic agents with immune checkpoint inhibitors could potentially overcome immune suppression driven by tumor angiogenesis. We report results from a phase IB study of ziv-aflibercept plus pembrolizumab in patients with advanced solid tumors. Methods This is a multicenter phase IB dose-escalation study of the combination of ziv-aflibercept (at 2–4 mg/kg) plus pembrolizumab (at 2 mg/kg) administered intravenously every 2 weeks with expansion cohorts in programmed cell death protein 1 (PD-1)/programmed death-ligand 1(PD-L1)-naïve melanoma, renal cell carcinoma (RCC), microsatellite stable colorectal cancer (CRC), and ovarian cancer. The primary objective was to determine maximum tolerated dose (MTD) and recommended dose of the combination. Secondary endpoints included overall response rate (ORR) and overall survival (OS). Exploratory objectives included correlation of clinical efficacy with tumor and peripheral immune population densities. Results Overall, 33 patients were enrolled during dose escalation (n=3) and dose expansion (n=30). No dose-limiting toxicities were reported in the initial dose level. Ziv-aflibercept 4 mg/kg plus pembrolizumab 2 mg/kg every 2 weeks was established as the MTD. Grade ≥3 adverse events occurred in 19/33 patients (58%), the most common being hypertension (36%) and proteinuria (18%). ORR in the dose-expansion cohort was 16.7% (5/30, 90% CI 7% to 32%). Complete responses occurred in melanoma (n=2); partial responses occurred in RCC (n=1), mesothelioma (n=1), and melanoma (n=1). Median OS was as follows: melanoma, not reached (NR); RCC, 15.7 months (90% CI 2.5 to 15.7); CRC, 3.3 months (90% CI 0.6 to 3.4); ovarian, 12.5 months (90% CI 3.8 to 13.6); other solid tumors, NR. Activated tumor-infiltrating CD8 T cells at baseline (CD8+PD1+), high CD40L expression, and increased peripheral memory CD8 T cells correlated with clinical response. Conclusion The combination of ziv-aflibercept and pembrolizumab demonstrated an acceptable safety profile with antitumor activity in solid tumors. The combination is currently being studied in sarcoma and anti-PD-1-resistant melanoma. Trial registration number NCT02298959 .

Rahma, Osama E.↗

PETSc TSAdjoint: A Discrete Adjoint ODE Solver for First-Order and Second-Order Sensitivity Analysis

Here, we present a new software system PETSc TSAdjoint for first-order and second order adjoint sensitivity analysis of time-dependent nonlinear differential equations. The derivative calculation in PETSc TSAdjoint is essentially a high-level algorithmic differentiation process. The adjoint models are derived by differentiating the timestepping algorithms and implementing them based on the parallel infrastructure in PETSc. Full differentiation of the library code, including MPI routines, is avoided, and users do not need to derive their own adjoint models for their specific applications. PETSc TSAdjoint can compute the first-order derivative, that is, the gradient of a scalar functional, and the Hessian-vector product, which carries second-order derivative information, while requiring minimal input (a few callbacks) from the users. The adjoint model employs optimal checkpointing schemes in a manner that is transparent to users. Finally, usability, efficiency, and scalability are demonstrated through examples from a variety of applications.

79 ASTRONOMY AND ASTROPHYSICS↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

Enabling Command-and-Control in Advanced In Situ Workflows

Scientific discovery is progressing towards autonomous science with the combination of scientific instruments, high-performance computing, and artificial intelligence in complex workflows. This evolution introduces new requirements for managing scientific workflows, including feedback loops, near real-time constraints, and the ability to dynamically control workflow execution. In situ workflows that analyze and visualize data as it is generated are well-suited to satisfy stringent time constraints and their iterative nature offers greater opportunities for command-and-control. However, only a few of the many workflow management systems available have been specifically designed to manage in situ workflows and often lack support for automated feedback loops that allow analysis and visualization components to interact with the main scientific data producer. To address this need, we present in this paper how to add command-and-control capabilities to a workflow management system. We identify the functional design requirements of such a command-and-control system, detail its architecture, interface, and core mechanisms, and illustrate how advanced in situ workflows can leverage command-and-control in three use cases: graceful termination with checkpoint, dynamic and adaptive data reduction, and event-triggered analysis.

Mehta, Kshitij [ORNL] (ORCID:0000000297149981)↗

AEflow (Autoencoder fluid flow compression network) [SWR-22-29]

As the size of turbulent flow simulations continues to grow, in situ data compression is becoming increasingly important for visualization, analysis, and restart checkpointing. For these applications, single-pass compression techniques with low computational and communication overhead are crucial. In this paper we present a deep-learning approach to in situ compression using an autoencoder architecture that is customized for three-dimensional turbulent flows and is well suited for contemporary heterogeneous computing resources. The autoencoder is compared against a recently introduced randomized single-pass singular value decomposition (SVD) for three different canonical turbulent flows: decaying homogeneous isotropic turbulence, a Taylor-Green vortex, and turbulent channel flow. Our proposed fully convolutional autoencoder architecture compresses turbulent flow snapshots by a factor of 64 with a single pass, allows for arbitrarily sized input fields, is cheaper to compute than the randomized single-pass SVD for typical simulation sizes, performs well on unseen flow configurations, and has been made publicly available. The results reported here show that the autoencoder dramatically outperforms a randomized single-pass SVD with similar compression ratio and yields comparable performance to a higher-rank decomposition with an order of magnitude less compression in regard to preserving a number of important statistical quantities such as turbulent kinetic energy, enstrophy, and Reynolds stresses.

King, Ryan↗

ezAlign

The ezAlign is aimed at clustering coarse grain simulation to find common or uncommon occurrences and convert coarse (CG) grained coordinate and topology files to atomistic formats. We use a PointNet based approach to map individual frames of simulation to points in a latent space. These points are then clustered using a variety of clustering methods. Clusters are analyzed to associate them with states in the simulation. Frames can be chosen from these clusters based on proximity to cluster centers. ezAlign takes CG coordinate and topology files and converts and outputs their corresponding atomistic formats using an alignment and relaxation procedure. ezAlign is designed to convert complex, solvated biological systems including lipid membranes with drug-like molecules using GROMACS. A GROMACS checkpoint (.cpt) file is also outputted to enable continuation simulations that retain the equilibrated atomic velocities. Independent atomistic coordinates and topologies for every molecule must already be included in ezAlign/files. A number of commonly simulated biological molecules are currently provided.

Bennett, WilliamF.↗

SuperNu Version 4.x

We seek to release SuperNu, Version 4.x, as a continuation of development for the open source SuperNu software. The SuperNu, Version 3.x Monte Carlo radiative transfer code for astrophysical transients was released with GPLv3 copyright, asserted by LANL in 2015. For the next release we have features planned for development , including: opacity implementation (including non-local thermodynamic equilibrium effects), generalized source implementation (e.g. for emulating shock heating in Type II supernovae), 3T (electron, ion, radiation) internal energy update, special relativity corrections through O(v^2/c^2), light polarization (e.g. for comparison to spectroplarimetry observations of supernovae and kilonovae), and infrastructure features (checkpoint and restart of simulations, and tools including setup and Slurm batch scripts and simulation post-processing/analysis scripts). These features are intended to improve the fidelity and/or better understand uncertainty in supernova and kilonova light curve calculations.

Wollaeger, Ryan↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

PRIME: Protein Representation Inference for Mutation Evaluation

Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.

Gibson, Kaetlyn [Los Alamos National Lab]↗

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

CodeScribe Agent

SF-26-086 CodeScribe introduces a structured, multi-stage pipeline that combines deterministic program analysis with LLM-powered translation to enable incremental, testable Fortran-to-C++ migration. First, `code-scribe index` traverses the project directory tree and produces `scribe.yaml` metadata files recording all modules, subroutines, and functions at each level, giving the LLM accurate structural context instead of a hallucinated codebase model. Second, `code-scribe draft` performs the deterministic portion of translation — converting Fortran types to C++ equivalents, replacing `use` statements with `#include` and `using namespace` directives, and detecting constructs requiring special handling — while embedding`scribe-prompt` annotations that guide the LLM through non-trivial cases such as statement-function-to-lambda conversions and `extern "C"` wrapper generation. Third, `code-scribe translate` applies project-specific TOML-based few-shot prompt templates and submits the composed prompt to a pluggable LLM backend (OpenAI, Anthropic, Argonne ARGO, any OpenAI-compatible endpoint, or local Hugging Face checkpoints), producing a C++ source file, a header, and a Fortran-C++ interface file for each translated routine so the codebase compiles and runs correctly throughout the migration. Beyond translation, CodeScribe includes a tool-using coding agent (`code-scribe agent`) with read, bash, edit, and write capabilities, and a bounded loop mode (`code-scribe loop`) that runs repeated stateless agent sessions over a task file with restricted tool access — enabling sustained, auditable software development workflows for broader scientific computing tasks.

Dhruv, Akash [Argonne National Laboratory (ANL), A↗

Colorectal Cancer Metastases in the Liver Establish Immunosuppressive Spatial Networking between Tumor-Associated SPP1 + Macrophages and Fibroblasts

Abstract Purpose: The liver is the most frequent metastatic site for colorectal cancer. Its microenvironment is modified to provide a niche that is conducive for colorectal cancer cell growth. This study focused on characterizing the cellular changes in the metastatic colorectal cancer (mCRC) liver tumor microenvironment (TME). Experimental Design: We analyzed a series of microsatellite stable (MSS) mCRCs to the liver, paired normal liver tissue, and peripheral blood mononuclear cells using single-cell RNA sequencing (scRNA-seq). We validated our findings using multiplexed spatial imaging and bulk gene expression with cell deconvolution. Results: We identified TME-specific SPP1-expressing macrophages with altered metabolism features, foam cell characteristics, and increased activity in extracellular matrix (ECM) organization. SPP1+ macrophages and fibroblasts expressed complementary ligand–receptor pairs with the potential to mutually influence their gene-expression programs. TME lacked dysfunctional CD8 T cells and contained regulatory T cells, indicative of immunosuppression. Spatial imaging validated these cell states in the TME. Moreover, TME macrophages and fibroblasts had close spatial proximity, which is a requirement for intercellular communication and networking. In an independent cohort of mCRCs in the liver, we confirmed the presence of SPP1+ macrophages and fibroblasts using gene-expression data. An increased proportion of TME fibroblasts was associated with the worst prognosis in these patients. Conclusions: We demonstrated that mCRC in the liver is characterized by transcriptional alterations of macrophages in the TME. Intercellular networking between macrophages and fibroblasts supports colorectal cancer growth in the immunosuppressed metastatic niche in the liver. These features can be used to target immune-checkpoint–resistant MSS tumors.

60 APPLIED LIFE SCIENCES↗