Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Large Dataset Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Finding Nuclear Clusters in the Short-Baseline Near Detector

Production of nuclear clusters such as deuterons, tritons, helions and alpha particles has only recently been simulated for neutrino-nucleus interactions, and has not yet been measured in neutrino experiments. The Short-Baseline Near Detector (SBND) is the first Liquid Argon Time Projection Chamber (LArTPC) with high enough resolution and a large enough neutrino flux to accurately measure the cross-section for the production of heavier-than-proton fragments in neutrino interactions. SBND is a 112-ton LArTPC, and lies 110 m downstream from the Booster Neutrino Beam target, where it has already collected the world's largest dataset of neutrino-argon interactions. In these processes, the intranuclear cascade and nuclear de-excitation stages can both produce nuclear clusters. With LArTPC technology's excellent reconstruction capabilities and SBND's unprecedented neutrino statistics, a measurement of these nuclear clusters could distinguish between nuclear models and improve neutrino energy reconstruction. This poster presents the status of simulation, reconstruction, and prospects for a measurement of nuclear clusters in SBND.

Beever, Anna [Sheffield U.] (ORCID:000900069339924↗

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES↗

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗

The 4D Camera: An 87 kHz Direct Electron Detector for Scanning/Transmission Electron Microscopy

We describe the development, operation, and application of the 4D Camera—a 576 by 576 pixel active pixel sensor for scanning/transmission electron microscopy which operates at 87,000 Hz. The detector generates data at ~480 Gbit/s which is captured by dedicated receiver computers with a parallelized software infrastructure that has been implemented to process the resulting 10–700 Gigabyte-sized raw datasets. The back illuminated detector provides the ability to detect single electron events at accelerating voltages from 30 to 300 kV. Through electron counting, the resulting sparse data sets are reduced in size by 10--300× compared to the raw data, and open-source sparsity-based processing algorithms offer rapid data analysis. The high frame rate allows for large and complex scanning diffraction experiments to be accomplished with typical scanning transmission electron microscopy scanning parameters.

47 OTHER INSTRUMENTATION↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Optimizing Geospatial Assessments for Nuclear Safeguards Applications with Large Language Models

A multidisciplinary team at Argonne National Laboratory evaluated the ability of large language models (LLMs) to identify geographic locations from open-source text and assessed post-processing measures to strengthen the reliability of those extractions in support of international nuclear safeguards. The study focused on addressing challenges such as toponym ambiguity, imprecise descriptions, and misinformation, which often undermine the accuracy of LLM-derived geospatial assessments. By integrating authoritative geospatial datasets, employing rigorous validation techniques, and leveraging human-in-the-loop processes, the project aimed to enhance the precision, transparency, and reproducibility of geospatial localization workflows. The findings demonstrate that while LLMs exhibit significant potential for accelerating geospatial analysis, their outputs require systematic grounding and verification to ensure reliability in high-stakes applications. This work contributes to the broader field of geospatial intelligence and supports strategic objectives of international organizations such as the International Atomic Energy Agency (IAEA) and the U.S. Department of Energy (DOE).

97 MATHEMATICS AND COMPUTING↗

Climate adaptation and sustainability in switchgrass: exploring plant-microbe-soil interactions across continental scale environmental gradients

Less carbon-intensive energy sources are needed to reduce greenhouse gas emissions and their predicted role in climate change. There is growing interest in the potential of biofuels for meeting this need. A critical question is whether large-scale biofuel production can be sustainable over the time scales needed to mitigate our carbon debt from fossil fuel consumption. The carbon balance and ultimately the sustainability of biofuel feedstock production is the result of complex climate-coupled interactions between carbon fixation, sequestration, and release through combustion. Similarly, the long-term productivity of biofuels depends on the environmental factors limiting plant growth. These factors are often related to soil resources which involve complex interactions at the plant-microbe-soil interface impacting their availability and cycling. Our collaborative project addressed sustainable switchgrass (Panicum virgatum) production by exploring Plant Systems, Plant-Microbiome Interactions, and Ecosystem Processes through the integrating lens of Multi-Scale Modeling. Our research was based on detailed characterization of genetically diverse switchgrass genotypes planted in common gardens across a continental latitudinal gradient. The underlying theme of our Plant Systems research was the use of locally adapted plant material to explore plant function, to understand the mechanistic basis of environmental interactions, and to discover the plant genes important for adaptation and sustainability in the face of climate change. Our Plant-Microbiome Interaction project characterized the microbial communities associated with switchgrass using genomic tools. Our Ecosystem Processes research focused on carbon cycle responses at the ecosystem level using stand level plantings. Finally, our Multi-Scale Modeling helped to define conditions of a sustainable biofuel system and identify key tradeoffs between genetic diversity, productivity, and ecosystem services. Genome-wide association analyses were used to identify alleles that contribute to successful establishment and biomass production across North America. Together, our work provided a baseline analyses of the potential of switchgrass as a biofuel feedstock. Our project resulted in a number of successful outcomes. First, we were successful in collecting switchgrass germplasm across the species range, propagating the material, and establishing common garden experiments across the species range. In collaboration with DOE JGI, we successfully assembled the first tetraploid switchgrass genome and published this resource with an analyses of the genetic basis local adaptation from our gardens (Lowry et al. 2019, Lovell et al. 2021). The gardens were used to characterize the genetic architecture for a number of important plant phenotypes. Our project also conducted extensive sampling and sequencing to characterize the bacterial and fungal associates of switchgrass roots and leaves. We showed that host genotype, location, and harvesting practices can play a role in microbiome assembly (Singer et al. 2019 & 2022, Van Wallendael et al. 2020 & 2022, Edwards et al. 2023). Our ecosystem processes work created baseline dataset of carbon and nutrient cycling in realistic stand plantings of switchgrass. Data from this experiment provided new insight into the role of plant traits, phenology, and local environments in ecosystem processes like soil respiration, net-ecosystem exchange, and dynamics of soil and plant nutrients (Ricketts et al. 2023). Finally, our crop modelling experiments help to characterize the sensitivity of common modeling frameworks to parameters, identify key limiters of productivity across large geographic scales, and leverage patterns of local adaptation in prediction. Ultimately, these studies help to identify critical plant-microbe-soil traits that may be manipulated, through breeding or agronomic management, to improve the sustainability of biofuel feedstocks.

09 BIOMASS FUELS↗

Tree-level carbon stock estimations across diverse species using multi-source remote sensing integration

Forests are critical carbon sinks, and remote sensing has been increasingly widely used for forest monitoring and biomass estimations. However, species-specific tree-level studies remain limited. In this study, we demonstrated the feasibility of integrating UAV-based LiDAR with high-resolution optical satellite imagery (0.5 m) to estimate biomass for individual trees across different species. The proposed method accurately estimated biomass for 53 trees (R² = 0.82, rRMSE = 0.44), with species-specific datasets, showing an average 25.2% increase in R² and a 14.8% reduction in rRMSE. A novel vegetation index combining forest structure parameters with vegetation indices (VIs) was developed using high-resolution multispectral satellite data (3 m) to explore its relationship with individual tree biomass. Combining forest structural parameters with VIs further improved estimation accuracy, achieving an R²of 0.89 and an rRMSE of 0.34. Species-specific datasets show an 11.6% increase in R²compared to methods without VIs, and a 22.2% improvement over methods using only VIs. SHapley Additive exPlanations (SHAP) analysis shows that the volume feature played a key role in model performance and remained stable throughout the training process. Altogether, the proposed approach enhances individual tree biomass and carbon sink estimations, showing great potential for large-scale precise forest carbon monitoring using multi-source remote sensing data.

59 BASIC BIOLOGICAL SCIENCES↗

Hierarchical Bayesian modeling for Inverse Uncertainty Quantification of system thermal-hydraulics code using critical flow experimental data

The best estimate plus uncertainty methodology in nuclear system thermal-hydraulic studies necessitates a comprehensive understanding of uncertainties in system code predictions. The forward uncertainty quantification (UQ) process involves the propagation of input uncertainties through the computational models to obtain uncertainties in the outputs. To this end, achieving an accurate estimation of input uncertainties is important, which is the focus of inverse UQ (IUQ). Traditionally, research in Bayesian IUQ within the nuclear engineering domain has largely relied on single-level Bayesian inference. While being effective for relatively small datasets, this approach encounters limitations for cases with large datasets. The use of a single-level model may prove inefficient, as the resultant posterior distributions can significantly differ when distinct subsets of data are employed. To address this issue, we employ an hierarchical Bayesian model for IUQ. Furthermore, this approach involves organizing observations into different groups based on the test conditions, thereby accommodating varying calibration parameters across these distinct groups. In this study, we developed and implemented a hierarchical Bayesian IUQ method to consider the grouping effect of critical flow measurement data from various geometries. Comparing the outcomes of IUQ under different selections of test data using hierarchical Bayesian IUQ against those obtained from single-level Bayesian IUQ, the forward propagation of hierarchical Bayesian IUQ results demonstrates a notably improved agreement with the experimental data.

42 ENGINEERING↗

Calibration of a soft secondary vertex tagger using proton-proton collisions at s = 13 TeV with the ATLAS detector

Several processes studied by the ATLAS experiment at the Large Hadron Collider produce low-momentum b-flavored hadrons in the final state. This paper describes the calibration of a dedicated tagging algorithm that identifies b-flavored hadrons outside of hadronic jets by reconstructing the soft secondary vertices originating from their decays. The calibration is based on a proton-proton collision dataset at a center-of-mass energy of 13 TeV corresponding to an integrated luminosity of 140 fb -1 . Scale factors used to correct the algorithm’s performance in simulated events are extracted for the b-tagging efficiency and the mistag rate of the algorithm using a data sample enriched in $t\overline{t}$ events. Several orthogonal measurement regions are defined, binned as a function of the multiplicities of soft secondary vertices and jets containing a b-flavored hadron in the event. The mistag rate scale factors are estimated separately for events with low and high average numbers of interactions per bunch crossing. The results, which are derived from events with low missing transverse momentum, are successfully validated in a phase space characterized by high missing transverse momentum and therefore are applicable to new physics searches carried out in either phase space regime.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

The first ensemble of kilometer-scale simulations of a hydrological year over the third pole

An accurate understanding of the current and future water cycle over the Third Pole is of great societal importance, given the role this region plays as a water tower for densely populated areas downstream. An emerging and promising approach for skillful climate assessments over regions of complex terrain is kilometer-scale climate modeling. As a foundational step towards such simulations over the Third Pole, we present a multi-model and multi-physics ensemble of kilometer-scale regional simulations for the hydrological year of October 2019 to September 2020. The ensemble consists of 13 simulations performed by an international consortium of 10 research groups, configured with a horizontal grid spacing ranging from 2.2 to 4 km covering all of the Third Pole region. These simulations are driven by ERA5 and are part of a Coordinated Regional Climate Downscaling EXperiment Flagship Pilot Study on Convection-Permitting Third Pole. The simulations are compared against available gridded and in-situ observations and remote-sensing data, to assess the performance and spread of the model ensemble compared to the driving reanalysis during the cold and warm seasons. Although ensemble evaluation is hindered by large differences between the gridded precipitation datasets used as a reference over this region, we show that the ensemble improves on many warm-season precipitation metrics compared with ERA5, including most wet-day and hour statistics, and also adds value in the representation of wet spells in both seasons. As such, the ensemble will provide an invaluable resource for future improvements in the process understanding of the hydroclimate of this remote but important region.

54 ENVIRONMENTAL SCIENCES↗

Location-Specific Microstructure Characterization Within AM Bench 2022 Nickel Alloy 718 3D Builds

Abstract The Additive Manufacturing Benchmark Test Series (AM Bench) is a broad effort to produce rigorous measurement datasets for validating AM computer simulations across the range of processing, structure, and properties, for many additive manufacturing (AM) build methods and material classes. Here, the microstructures of nickel alloy 718 AM Bench 2022 test artifacts produced using laser-based powder bed fusion (PBF-LB), in both as-built and fully heat-treated conditions, are examined. Cross sections are primarily characterized using large area scanning electron microscopy (SEM) electron backscatter diffraction (EBSD) and example analyses of the crystallographic textures are described. These data are part of a large set of in situ and ex situ measurements from both three-dimensional builds and laser tracks on bare plates. All the measurement data are available online with download links at www.nist.gov/ambench .

Levine, L. E. (ORCID:0000000334484229)↗

Vision Foundation Models in Remote Sensing: A survey

Artificial intelligence (AI) technologies have profoundly transformed the field of remote sensing (RS), revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, RS research has been significantly enhanced by the advent of foundation models (FMs)—large-scale pretrained AI models capable of performing a wide array of tasks with unprecedented accuracy and efficiency. This article provides a comprehensive survey of FMs in the RS domain. We categorize these models based on their architectures, pretraining datasets, and methodologies. Through detailed performance comparisons, we highlight emerging trends and the significant advancements achieved by those FMs. Additionally, we discuss technical challenges, practical implications, and future research directions, addressing the need for high-quality data, computational resources, and improved model generalization. Our research also finds that pretraining methods, particularly self-supervised learning (SSL) techniques like contrastive learning (CL) and masked autoencoders (MAEs), remarkably enhance the performance and robustness of FMs. This survey aims to serve as a resource for researchers and practitioners by providing a panorama of advances and promising pathways for the continued development and application of FMs in RS.

data models↗

How deep is your soil? Quantifying and spatially analyzing understudied deep soil in the United States

Deep soil is largely understudied and important in understanding biogeochemical processes in soil. Here, understudied soil is defined as the difference between soil studied to a known depth and the estimated bedrock depth. To understand more about deep soil, the understudied soil in the US was quantified and spatially analyzed using soil survey data and model estimates of bedrock depth. An equation was derived to find understudied soil using the dataset parameters “max lower depth studied”, “depth to bedrock”, and “likelihood of bedrock in the top 200 cm”. The survey data and bedrock model revealed that soil has been studied to an average depth of 1-2 meters, and the average depth to bedrock is 20 meters. Soil data density in the soil surveys was greatest in the West Coast, Midwest, and areas historically managed for agricultural, while the non-contiguous US and interior West were underrepresented. The soil had been studied deeper than the estimated soil depth in 455 out of 56,889 observation points concentrated in Alaska, California, Texas, Florida, Puerto Rico, and the US Virgin Islands. To understand the diversity and any taxonomic bias of the global soil data available, soil order was compared to US-based National Resource Conservation Service percentages and it was found that Oxisols, Alfisols, Ultisols, Andisols, and Histosols were overrepresented while Gelisols, Aridisols, Vertisols, Entisols, and Spodosols are underrepresented. Soil depth is important in exploring the complexity of biogeochemical processes that take place in soil.

Bedrock↗

Non -degenerate marginal-likelihood calibration with application to quantum characterization

Here, we propose a marginal likelihood strategy within the Kennedy-O’Hagan (KOH) Bayesian framework, where a Gaussian process (GP) models the discrepancy between a physical system and its simulator. Our approach introduces a novel marginalized likelihood by integrating out the degenerate eigenspace of the covariance matrix, rather than approximating the original likelihood. Unlike approximation methods that compromise accuracy for computational efficiency, our method defines an exact likelihood—distinct from the original but preserving all relevant information. This formulation achieves computational efficiency and stability, even for large datasets where the covariance matrix nears degeneracy. Applied to the characterization of a superconducting quantum device at Lawrence Livermore National Laboratory, the approach enhances the predictive accuracy of the Lindblad master equations for modeling Ramsey measurement data by effectively quantifying uncertainties consistent with the quantum data.

general physics↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗