Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

TChem v2.0 - A Software Toolkit for the Analysis of Complex Kinetic Models

TChem is an open source software library for solving complex computational chemistry problems and analyzing detailed chemical kinetic models. The software provides support for: complex kinetic models for gas-phase and surface chemistry; thermodynamic properties based on NASA polynomials; species production/consumption rates; stable time integrator for solving stiff time ordinary differential equations; and, reactor models such as homogenous gas-phase ignition (with analytical Jacobian matrices), continuously stirred tank reactor, plug-flow reactor. This toolkit builds upon earlier versions that were written in C and featured tools for gas-phase chemistry only. The current version of the software was completely refactored in C++, uses an object-oriented programming model, and adopts Kokkos as its portability layer to make it ready for the next generation computing architectures i.e., multi/many core computing platforms with GPU accelerators. We have expanded the range of kinetic models to include surface chemistry and have added examples pertaining to Continuously Stirred Tank Reactors (CSTR) and Plug Flow Reactor (PFR) models to complement the homogenous ignition examples present in the earlier versions. To exploit the massive parallelism available from modern computing platforms, the current software interface is designed to evaluate samples in parallel, which enables large scale parametric studies, e.g. for sensitivity analysis and model calibration.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Photon Detection System for DUNE Low-Energy Physics Study and the Demonstration of a Timing Resolution of a Few Nanoseconds Using ProtoDUNE-SP PDS

Photon detection systems (PDS) are an integral part of liquid-argon neutrino detectors. Besides providing the timing information for an event, which is necessary for reconstructing the drift coordinates of ionizing particle tracks, photon detectors can be effectively used for other purposes, including triggering events, background rejection, and calorimetric energy estimation. PDS in particular for the DUNE Far Detector Module 2 is designed to achieve a more extended optical coverage (→4 ) with new-generation large-size PD modules based on the ARAPUCA technology. This will provide enhanced opportunities for the study of low-energy neutrino physics using PDS. The ARAPUCA technology was extensively tested within the ProtoDUNE-SP detector operated at the CERN neutrino platform. Here, we present a study of the timing resolution of ARAPUCA detectors using light emitted from a sample of energetic cosmic ray muons traveling parallel to the PDS. An intrinsic timing resolution in the order of 3 ns is observed for the ARAPUCA detectors. The excellent timing resolution ability of PDS can be exploited for further enhancing physics studies using the DUNE far detectors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

First Examination of Irradiated Fuel with Pulsed Neutrons at LANSCE (Preliminary Results)

We present preliminary results on the characterization of an irradiated U-lOZr-lPd fuel sample that was prepared from the irradiated AFC-3A-R5A sample. U-lOZr metallic fuels are researched as host materials for potential transmutation fuels and the addition of palladium strives to bind lanthanides, thus preventing fuel-cladding chemical interactions (FCCI). These interactions limit the lifetime of metallic fuels and are caused by migration of lanthanide fission products to the periphery of the fuel slug, where they start to interact with the D9 or HT9 steel cladding, ultimately leading to failure of the mechanical integrity of the cladding. Neutrons offer bulk characterization of irradiated materials for which X-ray tomography methods are not suitable due to the immense gamma background emitted from the samples. In particular pulsed neutrons provide information from the ability to resolve the neutron energy using their time-of-flight and thus the potential to utilize neutron absorption resonance to characterize the spatial distribution of isotopes. This, in turn, may allow to characterize the distribution of fission and neutron capture products non-destructively and may ultimately be applied to the bulk of an irradiation capsule prior to destructive post-irradiation examination to identify regions of interest. To allow the characterization of entire irradiation capsules, a cask is under development in the advanced post-irradiation work package at LANL and progress on this development was reported elsewhere. In parallel, an irradiated U-lOZr-lPd sample cut from the AFC-3AR5A irradiation was shipped to LANL and will be fully characterized with an NSUF funded rapid turnaround experiment (RTE) in the 2020 LANSCE run cycle. The sample emits at a dose rate of ~3R/hr on contact and is therefore manageable with remote handling, without requiring a cask. The disk-shaped material is larger than samples prepared for analysis using electron or X-ray methods and is therefore an intermediate step towards characterization of bulk samples at LANSCE. However, since it covers the full diameter of the irradiated fuel slug, some insight on redistribution of elements, spatially resolved information on microstructure, e.g. phase composition and texture, will be possible using the pulsed neutron-based methods developed for fuel characterization at LANSCE. This report describes the development of procedures to handle the sample at LANSCE as well as preliminary data and results from tests conducted in December 2019 on the energy-resolved neutron imaging (ERNI) beam line at flight path 5 and the high pressure-preferred orientation diffractometer (HIPPO) at LANSCE. This effort is a collaboration between LANL, INL, and ORNL. To compare our capabilities with prior work, we present an overview of previously reported bulk characterization of irradiated or spent fuels. The overview addresses neutron diffraction and neutron absorption resonance spectroscopy, both of which have only few reported applications on irradiated or spent nuclear fuel, as well as neutron radiography. This literature review was already described in a previous report but is repeated here to put our current efforts in context of previous work.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Frequency-domain computing using nonlinear acoustic-wave device on lithium niobate

Abstract Multiply-accumulation are crucial computing operations in signal processing, numerical simulations, and machine learning. In recent years, optical analog approaches have demonstrated higher computing performance and better power efficiency than their digital counterparts. However, analog computing chips usually need large areas and complex structures for parallel computing, as a single device element only executes one computing operation at a single time. Here, we demonstrate frequency-domain computing using the nonlinear acoustic-wave devices on lithium niobate, featuring a normalized external second-harmonic generation conversion efficiency of ~ 5.7 × 10-4 W-1. The second-order sum-frequency nonlinear process of lithium niobate enables multiplication of inputs encoded in the frequency domain. Compared to the analog schemes, our device features a notably simpler design, and nanofabrication requires only one lift-off. Using a single acoustic-wave device within an area of 0.03 mm2, we can simultaneously conduct over 130,000 multiply-accumulation operations. Our acoustic-wave device shows applications in real and complex vector convolutions and image processing. This demonstration sets the stage for experimental realizations into frequency-domain integrated nonlinear acoustic computing systems, potentially shaping future developments in acoustic neural networks and quantum computing.

chai, mingzhao (ORCID:0009000466226341)↗

Towards performance portability in the Spark astrophysical magnetohydrodynamics solver in the Flash-X simulation framework

Simulations of core-collapse supernovae, and other astrophysical phenomena, are quintessential extreme-scale computing challenges. For core-collapse supernova simulations to be carried out by the ExaStar project under the Exascale Computing Project umbrella, a robust, efficient, and state-of-the-art magnetohydrodynamics solver is a critical requirement. In Flash-X, the primary software instrument for ExaStar, a new magnetohydrodynamics solver has been designed and implemented from the ground up to achieve accuracy and efficiency for simulations of complex astrophysical flows. This new solver, dubbed Spark, uses high-order spatial reconstruction, Runge-Kutta time integration, and an efficient cell-centered approach to satisfying the divergence-free condition for the magnetic fields. Spark was written to be optimized for data locality in cache hierarchy of CPUs. Since data locality optimizations for cache hierarchy are not directly compatible with those of accelerators, we have taken the approach of using program synthesis to avoid massive amounts of code replication that would be necessary if we were to maintain two different versions of the solver. Our program synthesis relies on a simple key-dictionary approach, implemented in python, that enables us to assemble the version of the solver suitable for the target hardware from code fragments identified by specific keys. In this work, we describe the data locality optimizations of the solver for CPUs and accelerators and the program synthesis tools that enable this portability. We also detail the parallel performance of Spark for both CPUs and accelerators.

97 MATHEMATICS AND COMPUTING↗

Core-edge integrated predictive studies of ST40 and NSTX plasmas with the scrape-off layer box model

The ability to model the interplay between the core and edge of tokamak plasmas is crucial to designing both the plasma operating scenario of a fusion pilot plant and the design of the tokamak itself. Scrape-off-layer (SOL) models that are tailored to integrated scenario modeling need to have fast turn-around time and minimal computational burden to enable wide parameter-space coverage for design scoping. The SOL 0-D Box model is a reduced SOL model based on global power and particle balance that captures the essential physics of SOL transport with little computational cost. The usage of the 0-D Box model in core-edge coupled simulations has been demonstrated in both interpretive and predictive modes on a variety of devices. This paper presents a sensitivity study of the 0-D Box model to the input SOL heat-flux width for an ST40 plasma. This study demonstrates that accurate prediction of this width is crucial to predicting global performance parameters of a plasma scenario, such as energy confinement time and flux consumption. We also present an extension of the Box model to 1-D to allow for parallel variation of plasma parameters along the magnetic field lines. The 1-D Box model is then compared with SOLPS-ITER simulations of an NSTX plasma. Advantages and limitations of the Box model are discussed, and future directions are outlined.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Advanced sorbents for modular oxygen production for REMS gasifiers

Under a DOE funded effort, Thermosolv LLC has been developing a sorbent-based oxygen production technology for gasification and oxy-combustion applications. The sorbents utilize oxygen-storage properties of certain perovskites to (1) selectively adsorb oxygen at moderate temperature from compressed air and (2) release the adsorbed oxygen into a vacuum or a sweep gas such as CO2 and/or steam. Through cyclic operations of multiple sorbent beds such a process can be made continuous. Pressure drop across the sorbent bed and other similar design considerations dictate that the sorbent be in the form of a high surface-area-to-volume pellet and yet still possess sufficient crush strength and integrity. Process efficiency is determined by the sorption and desorption kinetics and sorbent capacity. Typical process involves cycle times in the order of 100 seconds, a time too short to fully utilize the full volume of the sorbent and the sorption capacity of the sorbent pellets. As a part of this project, Thermosolv LLC undertook the development of oxygen sorbents as a high-surface area supported sorbent on a light-weight inert support to (1) reduce the cost, (2) increase the productivity, and (3) reduce the overall weight of the reactor. Using a previously developed perovskite (LSCF 1991) with its well-characterized performance, composite sorbent pellets each consisting of an inert core coated by a thin layer of the functional perovskite material were produced. Several low-density, inexpensive inert solids supports in the 1/8”-1/4” size were selected to keep the pressure drop across the sorbent bed in the acceptable range. A number of commercial technologies including spray coating, dip coating, incipient impregnation of porous substrate followed by thermal annealing at various temperatures and duration were employed and tested to produce robust composite supported sorbent pellets. A sense of optimum coating thickness was developed by testing sorbent pellets of various increasing diameter pellets. LSCF-1991extrudates ranging in size from 1/32” to 3/16” were tested in a TGA and in a fixed-bed pressure swing test set-up for sorption/desorption cycles of interest. The data show that for the operational conditions of interest the coating thickness for a supported sorbent pellet needs to be approximately 1/32” (about 0.8 mm). Subsequent work thereby concentrated on developing composite pellets of various substrates, shapes and sizes with coating layer of about this size range. Candidate support materials were chosen based on cost, inertness, mechanical strength, thermal expansion and chemical stability with sorbent material, and in a size range to give an acceptable pressure drop in fixed-bad reactor configurations. Coating application methods used for application of the sorbent onto the support included precipitation, spray coating and dip coating from sorbent slurries and sol gels. For all coated supports where we could successfully apply a uniform coating of desired thickness with an acceptable handling performance in terms of exfoliation, TGA-based cyclic sorption/desorption testing was performed to determine cycling capacity of oxygen. Among nearly twenty different support materials tested, best adhesion performance was obtained from stainless saddle supports. In the bench-scale fixed-bad tests stainless steel saddle support composite pellets with a coating thickness in the 0.6 mm or so range, the composite sorbent pellet performance approached up to 95% of that of the 2 mm parent material pellets. In a parallel approach to reducing the cost of the sorbent, alternate sorbent formulations replacing/reducing the amount of cobalt in the LSCF family were also investigated. Successful formulations that could match the performance of LSCF 1991 were identified based on TGA cyclic tests as LSCF 1919 and LSF 1910. In the range of operational envelop of cycle times and other relevant process operating conditions, the reduced Co formulations showed comparable performance in the bench-scale fixed-bad cyclic operations. As a part of this project, we also attempted substitution of La and Sr with Ca and Ba, and Co with Mn, Ni and Cu with little success, but the overall project goal of reducing sorbent cost in terms of raw material and manufacturing expenses was successful.

01 COAL, LIGNITE, AND PEAT↗

Unified Language Frontend for Physic-Informed AI/ML

Artificial intelligence and machine learning (AI/ML) are becoming important tools for scientific modeling and simulation as in several other fields such as image analysis and natural language processing. ML techniques can leverage the computing power available in modern systems and reduce the human effort needed to configure experiments, interpret and visualize results, draw conclusions from huge quantities of raw data, and build surrogates for physics based models. Domain scientists in fields like fluid dynamics, microelectronics and chemistry can automate many of their most difficult and repetitive tasks or improve the design times by use of the faster ML-surrogates. However, modern ML and traditional scientific highperformance computing (HPC) tend to use completely different software ecosystems. While ML frameworks like PyTorch and TensorFlow provide Python APIs, most HPC applications and libraries are written in C++. Direct interoperability between the two languages is possible but is tedious and error-prone. In this work, we show that a compiler-based approach can bridge the gap between ML frameworks and scientific software with less developer effort and better efficiency. We use the MLIR (multi-level intermediate representation) ecosystem to compile a pre-trained convolutional neural network (CNN) in PyTorch to freestanding C++ source code in the Kokkos programming model. Kokkos is a programming model widely used in HPC to write portable, shared-memory parallel code that can natively target a variety of CPU and GPU architectures. Our compiler-generated source code can be directly integrated into any Kokkosbased application with no dependencies on Python or cross-language interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Encoder–decoder neural network for solving the nonlinear Fokker–Planck–Landau collision operator in XGC

An encoder–decoder neural network has been used to examine the possibility for acceleration of a partial integro-differential equation, the Fokker–Planck–Landau collision operator. This is part of the governing equation in the massively parallel particle-in-cell code XGC, which is used to study turbulence in fusion energy devices. The neural network emphasizes physics-inspired learning, where it is taught to respect physical conservation constraints of the collision operator by including them in the training loss, along with the ℓ 2 loss. In particular, network architectures used for the computer vision task of semantic segmentation have been used for training. A penalization method is used to enforce the ‘soft’ constraints of the system and integrate error in the conservation properties into the loss function. During training, quantities representing the particle density, momentum and energy for all species of the system are calculated at each configuration vertex, mirroring the procedure in XGC. This simple training has produced a median relative loss, across configuration space, of the order of 10 –4 , which is low enough if the error is of random nature, but not if it is of drift nature in time steps. The run time for the current Picard iterative solver of the operator is O(n 2 ), where n is the number of plasma species. As the XGC1 code begins to attack problems including a larger number of species, the collision operator will become expensive computationally, making the neural network solver even more important, especially since its training only scales as O(n). Here, a wide enough range of collisionality has been considered in the training data to ensure the full domain of collision physics is captured. An advanced technique to decrease the losses further will be subject of a subsequent report. Eventual work will include expansion of the network to include multiple plasma species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enabling Large-Scale Condensed-Phase Hybrid Density Functional Theory Based Ab Initio Molecular Dynamics. 1. Theory, Algorithm, and Performance

By including a fraction of exact exchange (EXX), hybrid functionals reduce the self-interaction error in semilocal density functional theory (DFT) and thereby furnish a more accurate and reliable description of the underlying electronic structure in systems throughout biology, chemistry, physics, and materials science. However, the high computational cost associated with the evaluation of all required EXX quantities has limited the applicability of hybrid DFT in the treatment of large molecules and complex condensed-phase materials. To overcome this limitation, we describe a linear-scaling approach that utilizes a local representation of the occupied orbitals (e.g., maximally localized Wannier functions (MLWFs)) to exploit the sparsity in the real-space evaluation of the quantum mechanical exchange interaction in finite-gap systems. In this work, we present a detailed description of the theoretical and algorithmic advances required to perform MLWF-based ab initio molecular dynamics (AIMD) simulations of large-scale condensed-phase systems of interest at the hybrid DFT level. We focus our theoretical discussion on the integration of this approach into the framework of Car–Parrinello AIMD, and highlight the central role played by the MLWF-product potential (i.e., the solution of Poisson’s equation for each corresponding MLWF-product density) in the evaluation of the EXX energy and wave function forces. We then provide a comprehensive description of the exx algorithm implemented in the open-source Quantum ESPRESSO program, which employs a hybrid MPI/OpenMP parallelization scheme to efficiently utilize the high-performance computing (HPC) resources available on current- and next-generation supercomputer architectures. Furthermore, this is followed by a critical assessment of the accuracy and parallel performance (e.g., strong and weak scaling) of this approach when AIMD simulations of liquid water are performed in the canonical (NVT) ensemble. With access to HPC resources, we demonstrate that exx enables hybrid DFT-based AIMD simulations of condensed-phase systems containing 500–1000 atoms (e.g., (H₂O)₂₅₆) with a wall time cost that is comparable to that of semilocal DFT. In doing so, exx takes us one step closer to routinely performing AIMD simulations of complex and large-scale condensed-phase systems for sufficiently long time scales at the hybrid DFT level of theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

Stage-local partitioned two-step runge-kutta methods for large systems of ordinary differential equations

We introduce stage-local partitioned two-step Runge-Kutta methods are an extension of standard two-step Runge-Kutta methods, which are an alternative to the standard additive two-step Runge-Kutta methods currently existing in the literature. Furthermore, these new schemes are designed with an eye towards truly N-partitioned systems and leverage local stage approximations to make several computationally interesting approximations viable. Specifically, the focus on local stage approximations makes possible the construction of truly asynchronous schemes, in the parallel sense, possible. In addition, we show that an implicit-explicit approach to these schemes can lead to methods that require the inversion of only local nonlinear systems.

Applied Dynamical Systems↗

Time-discretization of a plasma-neutral MHD model with a semi-implicit leapfrog algorithm

The semi-implicit leapfrog time-discretization is a workhorse algorithm for initial-value MHD codes to bridge between vastly separated time scales. Inclusion of atomic interactions with neutrals breaks the functional structure of the MHD equations that exploited by the leapfrog. In this work, we address how to best integrate atomic physics into the semi-implicit leapfrog. Following the Crank-Nicolson method, one approach is to time-center the atomic interactions in the linear solver and use a Newton method to include the nonlinear contributions. Alternatively, another family of methods are based on operator-splitting the terms associated with the atomic interactions using a Strang-splitting technique. These methods naturally break equations into constituent ODE and PDE parts and preserve the structure exploited by the semi-implicit leapfrog. We study the accuracy and efficiency of these methods through a battery of 0D and 1D cases and show that a second-order-in-time Douglas-Rachford inspired coupling between the ODE and PDE advances is effective in reducing the time-discretization error to be comparable to that of Crank-Nicolson with Newton iteration of the nonlinear terms. Splitting ODE and PDE parts results in independent matrix solves for each field which reduces the computational cost considerably and provides parallelization over species relative to Crank-Nicolson.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Lab-Scale Cable-Driven Parallel Robot Prototype for Automated Prefabricated Component Manipulation

This paper presents the design and evaluation of a lab-scale cable-driven parallel robot (CDPR) developed as a flexible platform for automated installation of prefabricated components onto exterior building envelopes. Traditional manual installation methods for prefabricated components, which depend on scaffolding, cranes, cherry pickers, and verbal coordination, are not only labor-intensive and error-prone but also face significant limitations in dense urban environments due to site access constraints. To address these challenges, we developed a lab-scale CDPR platform capable of autonomously transporting building envelope components from a designated pickup zone to their target installation location, minimizing the need for human intervention. This study describes the system’s mechanical design, actuation architecture, real-time feedback system, and control strategy of the CDPR, and evaluates its performance in a laboratory environment. The robot’s actuation system uses torque control for end-effector manipulation. The robot’s real-time pose feedback comes from a construction-grade total station and a wireless inertial measurement unit (IMU), which together support precise end-effector control. Experimental results demonstrate the successful integration of the hardware, sensing, state estimation, and control subsystems. Preliminary tests showed that our lab-scale prototype can position the end effector with an error of less than 3 mm, which is a level of precision not previously achieved by existing CDPRs in construction applications. The key findings are twofold: (1) torque-only control is necessary but not sufficient for minimizing final pose error, and (2) incorporating real-time pose feedback can achieve the desired placement accuracy.

Liu, Yifang [Oak Ridge National Laboratory (ORNL),↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

In Situ High-Temperature Ultrafast Electron Diffraction through Integrated Furnace and MEMS Platforms

Temperature fundamentally governs phase stability, defect evolution, and transport behavior in materials. Despite its central role, direct measurements of structural evolution at elevated temperatures on ultrafast timescales have remained limited. Here, we report the design, integration, and validation of 2 complementary in situ heating platforms that substantially extend the thermal operating range of ultrafast electron diffraction (UED). A compact furnace-type heating stage enables stable diffraction measurements from room temperature to 800 K with ±0.1 K stability under ultrahigh vacuum, achieved through multi-sensor feedback control, dual air-cooling channels, and a thermally isolated motion stage. In parallel, a microelectromechanical system (MEMS)-based heating platform provides rapid thermal response and access to extreme temperatures ≥1,373 K with ±0.1 K stability over hundreds-micrometer regions while supporting simultaneous electrical biasing for electrothermal coupling studies. Absolute temperature calibration is established using diffraction-based thermometry via aluminum lattice expansion and independently validated through in situ melting of bismuth thin films. UED measurements further reveal pronounced temperature-dependent nonequilibrium lattice dynamics in bismuth, including modifications to electron–phonon coupling and Debye–Waller behavior, as well as enhanced ultrafast diffuse scattering in aluminum at elevated temperatures. Together, these developments establish a practical framework for quantitative, time-resolved studies of temperature-driven kinetics and nonequilibrium structural dynamics under extreme thermal environments.

Bai, Qianqian [Chinese Academy of Sciences (CAS), ↗

Modeling of ExB effects on tungsten re-deposition and transport in the DIII-D divertor

Mixed-material DIVIMP-WallDYN modelling, now incorporating ExB drifts, is presented that simultaneously reproduces tungsten (W) erosion and deposition patterns observed during the DIII-D Metal Rings Campaign, in which a toroidally symmetric set of W-coated tiles were installed in the carbon (C) DIII-D divertor. Since most reactor plasma facing component (PFC) designs call for mixed-material environments, including ITER’s W/Be enviroment, the divertor targets will quickly evolve into reconstitued surfaces of multiple elements. This work identifies controlling physics that affects material migration patterns in the divertor, which impact PFC lifetimes and impurity leakage from the divertor to the core. These simulations indicate that radial and poloidal ExB transport dominates over parallel force balance for high-Z impurities such as W in the divertor region of DIII-D. It is demonstrated that ExB drifts are required to reproduce the experimental observation of non-local W and C co-accumulation in a band ~7-9 cm outboard of the outer-strike-point (OSP) W source, for attached Lmode conditions in the unfavorable ion grad-B drift direction. In addition, W gross erosion is localized to the region outboard of the OSP, as the formation of C co-deposits suppresses W erosion at the strike point. Time-dependent simulations with scaled ExB impurity drifts (60% of the OEDGE-calculated drift velocity) and W re-erosion quantitatively reproduce these features, including depth-resolved W/C ratios, within a factor of 2 over ~115 seconds of accumulated plasma exposure. The location of co-deposition regions is shown to be well represented by an analytic leakage model, driven largely by poloidal ExB drifts. Qualitative agreement is also found between campaign-integrated W deposition measurements and simulations for the favorable ion grad-B drift direction, the standard mode of operation for most tokamaks. Furthermore, these results imply that a longterm inward radial migration of material from the outer divertor through the private flux region may occur in future devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗