Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel in time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

An Open-Source Parallel EMT Simulation Framework

As the integration level of inverter-based resources (IBRs) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗

Effects of magnetic field assisted heat treatment on the microstructure and mechanical properties of Fe-0.63 %C alloy

This study investigates the influence of an applied magnetic field on the microstructural evolution and mechanical properties of hypoeutectoid steels subjected to heat treatment. Tensile tests and microstructural analysis were performed on samples processed under varying magnetic field strengths (0 T, 5 T, and 9 T) and different austenitization incubation times. The results indicate that the application of a magnetic field alters the fraction of proeutectoid ferrite phase without changing the cooling rates and heat treatment process. Additionally, pearlite microstructural features such as lamellar spacing and misorientation angles exhibit variations under different field strengths. While the pearlite nodule diameter remains largely unaffected, an increase in percentage elongation and strength is observed in the 5 T treated sample, attributed to changes in microstructural features with the magnetic field. Additionally, the percentage elongation is reduced in the samples heat treated with reduced austenitization incubation times. The study further demonstrates that low-angle misorientations increased in the samples taken parallel to the magnetic field direction, influencing the mechanical response. These findings suggest that applying a magnetic field during heat treatment provides an additional driving force for phase transformations, offering a manufacturing process for tailoring microstructures and optimizing mechanical properties. Moreover, integrating magnetic fields in heat treatment processes has potential benefits in energy efficiency.

High magnetic field↗

Frequency-domain computing using nonlinear acoustic-wave device on lithium niobate

Abstract Multiply-accumulation are crucial computing operations in signal processing, numerical simulations, and machine learning. In recent years, optical analog approaches have demonstrated higher computing performance and better power efficiency than their digital counterparts. However, analog computing chips usually need large areas and complex structures for parallel computing, as a single device element only executes one computing operation at a single time. Here, we demonstrate frequency-domain computing using the nonlinear acoustic-wave devices on lithium niobate, featuring a normalized external second-harmonic generation conversion efficiency of ~ 5.7 × 10-4 W-1. The second-order sum-frequency nonlinear process of lithium niobate enables multiplication of inputs encoded in the frequency domain. Compared to the analog schemes, our device features a notably simpler design, and nanofabrication requires only one lift-off. Using a single acoustic-wave device within an area of 0.03 mm2, we can simultaneously conduct over 130,000 multiply-accumulation operations. Our acoustic-wave device shows applications in real and complex vector convolutions and image processing. This demonstration sets the stage for experimental realizations into frequency-domain integrated nonlinear acoustic computing systems, potentially shaping future developments in acoustic neural networks and quantum computing.

chai, mingzhao (ORCID:0009000466226341)↗

Investigation of pedestal parameters and divertor heat fluxes in small ELM regimes in DIII-D

Abstract Divertor heat flux and its correlation with pedestal parameters within various small edge localized mode (ELM) regimes, including high beta poloidal, type-II and ELMs with negative triangularity H-modes were investigated in DIII-D. The parallel energy fluences of type-II and high beta poloidal small ELM regimes fall below the linear scaling with pedestal electron pressure for type-I ELMs put forward in Eich et al 2017 ( Nucl. Mater. Energy 12 84–90). The negative triangularity of H-mode ELMs follow the Eich scaling for type-I ELMs. The parallel heat flux and total heat loads to the divertor were determined using high-time resolution infrared thermography, while pedestal parameters were obtained through self-consistent kinetic equilibrium reconstructions. Linear regressions for the type-II and high beta poloidal regimes demonstrate that an equivalent 7.5 MA small ELM scenario in ITER would fall below the ~5 MJ m − 2 leading edge melting limit for tungsten (Gunn et al 2017 Nucl. Fusion 57 046025). Utilizing fast thermography, the scrape-off layer power fall-off length for both inter-ELM and intra-ELM was determined and compared to the Eich scaling with poloidal magnetic field in Eich et al (ASDEX Upgrade Team and JET EFDA Contributors 2013 Nucl. Fusion 53 093031). Except for the high beta poloidal scenario, all the small ELM regimes during both inter- and intra-ELM periods had power fall-off lengths ( λ q ) larger then would be expected from the B pol , MP − 1 scaling associated with type-I ELMs, signifying their potential in managing heat loads and offering a solution for core–edge integration.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predicting nonequilibrium Green’s function dynamics and photoemission spectra via nonlinear integral operator learning

Understanding the dynamics of nonequilibrium quantum many-body systems is an important research topic in a wide range of fields across condensed matter physics, quantum optics, and high-energy physics. However, numerical studies of large-scale nonequilibrium phenomena in realistic materials face serious challenges due to intrinsic high-dimensionality of quantum many-body problems and the absence of time-invariance. The nonequilibrium properties of many-body systems can be described by the dynamics of the correlator, or the Green's function of the system, whose time evolution is given by a high-dimensional system of integro-differential equations, known as the Kadanoff–Baym equations (KBEs). The time-convolution term in KBEs, which needs to be recalculated at each time step, makes it difficult to perform long-time numerical simulation. In this paper, we develop an operator-learning framework based on recurrent neural networks (RNNs) to address this challenge. We utilize RNNs to learn the nonlinear mapping between Green's functions and convolution integrals in KBEs. By using the learned operators as a surrogate model in the KBE solver, we obtain a general machine-learning scheme for predicting the dynamics of nonequilibrium Green's functions. Besides significant savings per each time step, the new methodology reduces the temporal computational complexity from $O(N_t^3)$ to $O(N_t)$ where N t is the number of steps taken in a simulation, thereby making it possible to study large many-body problems which are currently infeasible with conventional KBE solvers. Through various numerical examples, we demonstrate the effectiveness of the operator-learning based approach in providing accurate predictions of physical observables such as the reduced density matrix and time-resolved photoemission spectra. Moreover, our framework exhibits clear numerical convergence and can be easily parallelized, thereby facilitating many possible further developments and applications.

97 MATHEMATICS AND COMPUTING↗

Fokker-Planck simulations of fast ion ICRF and electron EC heating in a mirror plasma using CQL3D-m

The CQL3D-m continuum bounce-average Fokker-Planck code is adapted for magnetic mirror plasmas [1] and is now routinely used in no-free-parameter classical integrated modeling of mirror devices [2, 3]. In the present effort, we report on two RF methods of plasma heating in mirror machine. The fast ions (FI) are heated by Fast waves at 2nd-4th harmonic, where FIs originate from neutral beam injection at 45 degrees to the magnetic field. The scenario shows an efficient ion heating near the FI bouncing point. The electrons are heated by X-mode launched from the high magnetic field side towards the resonance. Different from the tokamak applications, CQL3D-m provides an evolving self-consistent ambipolar parallel electric field, which determines the shape of the loss cone and hence an accurate confinement time of both ions and electrons. Also, it includes a description of ion and electron sources and sinks (related to charge exchange and impact ionization) which are updated at every time step. CQL3D-m utilizes a fully nonlinear Coulomb collision operator that is important for the significantly non-Maxwellian ion distributions typically established in mirror plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Floquet-driven indirect exchange interaction mediated by topological insulator surface states

Light drives offer a potential tool for the dynamical control of magnetic interactions in matter. We theoretically investigate the indirect exchange coupling between two parallel chains of magnetic impurities on the surface of a topological insulator, driven by a time-periodic circularly polarized light field in the high-frequency, off-resonant regime. We derive a closed-form analytic expression for the spin susceptibility of the photon-dressed topological insulator surface states and obtain the irradiation dependence of the Ising, Heisenberg, and Dzyaloshinskii-Moriya exchange couplings between the impurity chains. Our results show a two-pronged modification of these exchange couplings by periodic drives. First, the Ruderman-Kittel-Kasuya-Yosida (RKKY) oscillation period of the exchange couplings can be extended by enhancing the driving strength. Second, increasing driving strength enhances the envelope of RKKY oscillations of the Ising type while suppressing those of the Heisenberg type and Dzyaloshinskii-Moriya type. Furthermore, our work provides useful insights for realizing Floquet engineering of collinear and noncollinear indirect exchange interactions in topological insulating systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Packaging a 650V/400A GaN Half-bridge Power Module with Ultra-low Parasitics for Electric Vehicle Drive Applications

This paper proposes a compact and efficient half-bridge power module with three 650 V / 150 A GaN dies in parallel. The power module incorporates a main power printed circuit board (PCB), an interface PCB, and a flex PCB to achieve low parasitics in both power loop and gate-side connection, resolving the issue of high parasitics typically encountered with wire bonding in high-current applications. Additionally, the interface PCB decouples the design constraints between the power loop and the gate loops. The proposed design is optimized with a vertical loop configuration to reduce power loop inductance through magnetic flux cancellation. Finite element analysis indicates that the power loop inductance is 0.58 nH at 100 MHz, while the maximum die junction temperature reaches 131 °C under an ambient temperature of 65 °C and a load current of 385 A. The proposed multi-piece PCB structure reduces the inductance of the drive circuit to minimize EMI and to mitigate false triggering. At the same time, it reduces impedance mismatches across different driver circuits, thereby achieving dynamic current sharing in multi-chip parallel configurations. Under simulation conditions of 400 V / 385 A, the current imbalance among chips was limited to 5 A. A 400 V / 385 A double-pulse test was conducted to experimentally validate the performance of the proposed power module.

30 DIRECT ENERGY CONVERSION↗

Large-scale real-time signal processing in physics experiments: the ALICE TPC FPGA pipeline

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB s -1 of raw detector data. This requirement is met by a custom FPGA-based processing pipeline that performs the complete front-end data treatment fully in-stream, including common-mode correction, pedestal subtraction, ion-tail filtering, zero suppression, and dense data packing. A central element of the design is a highly parallel common-mode correction algorithm operating directly on the streaming data. It robustly identifies signal-free readout channels on a time-bin basis and applies pad-dependent scaling to compensate for local variations in capacitive coupling in the GEM readout. In combination with pedestal subtraction and ion-tail filtering, this enables accurate baseline restoration under extreme high-occupancy conditions, preventing signal loss while efficiently suppressing noise prior to zero suppression. The pipeline operates continuously at the full detector bandwidth and reduces the raw input rate of approximately 3 TB s -1 to about 900 GBps for Pb-Pb collisions at the target interaction rate. Overall, it represents a large-scale FPGA-based real-time signal-processing implementation for high-energy physics detector readout.

Digital signal processing (DSP)↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

A microfluidic spore chamber for long-term imaging of single-spore hyphal development.

Understanding the life cycle of fungal spores is essential for elucidating their roles in pathogenesis, dispersal, and survival. However, studying spore development under controlled, spatially defined conditions remains challenging. Here, we present the Spore Chamber, a custom-built microfluidic platform engineered for parallel trapping and long-term imaging of individual spores under defined media conditions, enabling real-time visualization of hyphal development. Using Aspergillus fumigatus as a model organism, we demonstrate that sparse trapping of individual spores within size-matched trap geometries enables long-term time-lapse imaging of key developmental stages, including germination, polarized hyphal elongation, branching, and conidiophore formation. To assess the device's capacity to resolve morphogenetic responses to exogenous signals, we introduced lipochitooligosaccharides (LCOs) and short-chain chitooligosaccharides (COs). Rhizobium-derived, non-sulfated LCO (nsLCO) mixtures induced enhanced secondary branching (hyperbranching), a response not previously reported in A. fumigatus under these signal conditions, to our knowledge, whereas sulfated LCOs and CO4 did not significantly alter branching patterns. In addition, long-term confinement and imaging revealed rare developmental morphologies previously described primarily in mutant strains, including split conidiophore formation, elongated phialides, and stress-associated phenomena such as microcyclic conidiation, and chlamydospore development. Together, these results establish the Spore Chamber as a targeted microfluidic platform for single-spore phenotyping and long-term developmental analysis, with applications in fungal biology, chemical signaling studies, and host–microbe interaction research.

Antifungal screening↗

Dataset for "A Microfluidic Spore Chamber for Long-Term Imaging of Single-Spore Hyphal Development"

Understanding the life cycle of fungal spores is essential for elucidating their roles in pathogenesis, dispersal, and survival. However, studying spore development under controlled, spatially defined conditions remains challenging. Here, we present the Spore Chamber, a custom-built microfluidic platform engineered for parallel trapping and long-term imaging of individual spores under defined media conditions, enabling real-time visualization of hyphal development. Using Aspergillus fumigatus as a model organism, we demonstrate that sparse trapping of individual spores within size-matched trap geometries enables long-term time-lapse imaging of key developmental stages, including germination, polarized hyphal elongation, branching, and conidiophore formation. To assess the device’s capacity to resolve morphogenetic responses to exogenous signals, we introduced lipochitooligosaccharides (LCOs) and short-chain chitooligosaccharides (COs). Rhizobium-derived, non-sulfated LCO (nsLCO) mixtures induced enhanced secondary branching (hyperbranching), a response not previously reported in A. fumigatus under these signal conditions, to our knowledge, whereas sulfated LCOs and CO4 did not significantly alter branching patterns. In addition, long-term confinement and imaging revealed rare developmental morphologies previously described primarily in mutant strains, including split conidiophore formation, elongated phialides, microcyclic conidiation, and chlamydospore development. Together, these results establish the Spore Chamber as a targeted microfluidic platform for single-spore phenotyping and long-term developmental analysis, with applications in fungal biology, chemical signaling studies, and host–microbe interaction research. Videos of the observed phenomena are included in this data set.

59 BASIC BIOLOGICAL SCIENCES↗

Computational Algorithms for Unit Commitment with AC Power Flows (Final Report)

Security-constrained unit commitment (SCUC) is a key component in power system operations. When AC power flow constraints are considered in the SCUC model (AC-SCUC), the problem becomes extremely difficult due to its discrete and non-convex nature, as described in “Grid Optimization Competition Challenge 3 Problem Formulation (GOCC)”. There are four main challenges: (i) Discrete decisions regarding unit online/offline status and start-up/shut-down procedures for every single unit. The number of discrete decision variables increases considerably when a system integrates multiple generators; (ii) Configuration-based combined-cycle formulations, and multi-commodity models that include ramping products, spin/non-spin products, and regulation up/down products. The combined-cycle units introduce additional discrete decision variables and auxiliary service products further complicate the model by connecting multi-commodity products’ continuous and discrete variables; (iii) SCUC models with AC power flow constraints are far more complex due to massive bilinear terms in the large-scale nonlinear power balance equations. The nonlinear power balance equations are further complicated by the discrete step control variables of shunts; (iv) N − 1 contingency analysis. The size of the model increases linearly with the number of contingencies considered, greatly increasing the size of the optimization model. Accordingly, there is an emergent need to develop a robust algorithm capable of deriving a high-quality solution in a short time and passing through contingency tests simultaneously. In this project, we explore innovative techniques to address this challenging problem by integrating advanced polyhedral theory, approximation methods, relaxation strategies, decomposition techniques, and parallel computing. Each technique approaches the problem from a different perspective, leveraging its specific strengths to tackle distinct challenges. Each individual method has demonstrated its effectiveness in the PI’s previous research. Their integration is expected to significantly reduce the computational time required to solve the proposed complex problem. Successful completion of this project has the potential to transform the industry by enhancing optimization solvers capable of handling large-scale day-ahead energy market clearing models within strict time constraints, while incorporating AC power flow constraints. This advancement will lead to reduced overall generation costs and, consequently, increased social welfare.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗