Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Surrogate Hessian accelerated structural optimization for stochastic electronic structure theories

In this work, we present an efficient energy-based method for structural optimization with stochastic electronic structure theories, such as diffusion quantum Monte Carlo (DMC). This method is based on robust line-search energy minimization in reduced parameter space, exploiting approximate but accurate Hessian information from a surrogate theory, such as density functional theory. The surrogate theory is also used to characterize the potential energy surface, allowing for simple but reliable ways to maximize statistical efficiency while retaining controllable accuracy. We demonstrate the method by finding the minimum DMC energy structures of the selected flake-like aromatic molecules, such as benzene, coronene, and ovalene, represented by 2, 6, and 19 structural parameters, respectively. In each case, the energy minimum is found within two parallel line-search iterations. The method is near-optimal for a line-search technique and suitable for a broad range of applications. It is easily generalized to any electronic structure method where forces and stresses are still under active development and implementation, such as diffusion Monte Carlo, auxiliary-field Monte Carlo, and stochastic configuration interaction, as well as deterministic approaches such as the random-phase approximation. Accurate and efficient means of geometry optimization could shed light on a broad class of materials and molecules, showing high sensitivity of induced properties to structural variables.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multitask graph neural networks for elastoplastic response prediction in dual-phase polycrystals

Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.

Crystal plasticity↗

Modeling Stochastic Variability in Multiband Time-series Data

In preparation for the era of time-domain astronomy with upcoming large-scale surveys, we propose a state-space representation of a multivariate damped random walk process as a tool to analyze irregularly-spaced multifilter light curves with heteroscedastic measurement errors. We adopt a computationally efficient and scalable Kalman filtering approach to evaluate the likelihood function, leading to maximum O(k 3 n) complexity, where k is the number of available bands and n is the number of unique observation times across the k bands. This is a significant computational advantage over a commonly used univariate Gaussian process that can stack up all multiband light curves in one vector with maximum O(k 3 n 3 ) complexity. Using such efficient likelihood computation, we provide both maximum likelihood estimates and Bayesian posterior samples of the model parameters. Three numerical illustrations are presented: (i) analyzing simulated five-band light curves for a comparison with independent single-band fits; (ii) analyzing five-band light curves of a quasar obtained from the Sloan Digital Sky Survey Stripe 82 to estimate short-term variability and timescale; (iii) analyzing gravitationally lensed g- and r-band light curves of Q0957+561 to infer the time delay. Two R packages, Rdrw and timedelay, are publicly available to fit the proposed models.

79 ASTRONOMY AND ASTROPHYSICS↗

Performance of reverse osmosis membrane with large feed pressure fluctuations from a wave-driven desalination system

Wave-driven desalination systems are proposed water treatment systems that involve reverse osmosis of seawater powered directly by wave motion. Such a configuration would result in drastic feed pressure fluctuations. For a technology conventionally operated with a constant feed condition, the effect of these variable pressures on membrane integrity and performance is unknown. Here, experiments were conducted with spiral wound membranes coupled to a system capable of producing feed pressure fluctuations of more than 400 psi. Feed composition included 5, 20, and 35 g/L NaCl, and a synthetic seawater at normal and 1.5x concentration. The variable feed conditions included sine-like pressure waves swings of 200-500 and 500-900 psi with frequencies of 1.25, 7.5, and 12 waves/min, and a model-generated random waveform. Between each wave experiment we performed membrane integrity tests at 650 psi and 25 g/L NaCl feed, which showed a 7.4% drop in the membrane's water permeability coefficient, an 18.4% flux decline, and more than 99% salt rejection over 1770 h of cumulative experimental time. Analysis of permeate samples showed high salt rejection. In general, variable feed pressure had no significant deleterious effect on membrane integrity or performance.

16 TIDAL AND WAVE POWER↗

Seasonal drivers of dissolved oxygen across a tidal creek–marsh interface revealed by machine learning

Abstract Dissolved oxygen (DO) is a key biogeochemical control in coastal systems, and its concentration and drivers vary markedly through time and space. This makes it difficult to accurately represent coastal DO and associated biogeochemical processes in models, limiting our ability to predict how these systems will respond to global change. We obtained high‐frequency (5‐min) in situ measurements of DO collected at three locations across the interface of a tidal creek and coastal marsh in the Pacific Northwest, USA. Random Forest machine learning models quantified the importance of three categories of environmental drivers (Aquatic, Climatic, and Terrestrial) of DO variability across the creek–marsh interface. We selected two 4‐month datasets representing Summer and Winter seasonal periods to test two hypotheses on the dominant drivers of DO at the coastal interface. We found that the Terrestrial driver—characterized by long periods of anaerobic conditions and episodic pulses in DO after floods—was most important during the Winter, whereas the Aquatic driver—characterized by variability over tidal, diel, and lunar cycles—was most important during the Summer. We explored how future climate change scenarios could alter the drivers of DO variability using a cumulative sums driver–response framework. Our results suggest that under climate change, Aquatic and Climatic drivers may increase in importance during the Summer, potentially linked to changing metabolic regimes and sea level, with Terrestrial driver importance potentially increasing during the Winter. Our approach highlights useful methods for understanding the spatiotemporal complexity of oxygen across coastal interfaces and quantifying the relative importance of distinct environmental drivers.

54 ENVIRONMENTAL SCIENCES↗

Identifying Transient Candidates in the Dark Energy Survey Using Convolutional Neural Networks

The ability to discover new transient candidates via image differencing without direct human intervention is an important task in observational astronomy. For these kind of image classification problems, machine learning techniques such as Convolutional Neural Networks (CNNs) have shown remarkable success. In this work, we present the results of an automated transient candidate identification on images with CNNs for an extant data set from the Dark Energy Survey Supernova program, whose main focus was on using Type Ia supernovae for cosmology. By performing an architecture search of CNNs, we identify networks that efficiently select non-artifacts (e.g., supernovae, variable stars, AGN, etc.) from artifacts (image defects, mis-subtractions, etc.), achieving the efficiency of previous work performed with random Forests, without the need to expend any effort in feature identification. The CNNs also help us identify a subset of mislabeled images. Performing a relabeling of the images in this subset, the resulting classification with CNNs is significantly better than previous results, lowering the false positive rate by 27% at a fixed missed detection rate of 0.05.

79 ASTRONOMY AND ASTROPHYSICS↗

Understanding Growth Dynamics and Yield Prediction of Sorghum Using High Temporal Resolution UAV Imagery Time Series and Machine Learning

Unmanned aerial vehicles (UAV) carrying multispectral cameras are increasingly being used for high-throughput phenotyping (HTP) of above-ground traits of crops to study genetic diversity, resource use efficiency and responses to abiotic or biotic stresses. There is significant unexplored potential for repeated data collection through a field season to reveal information on the rates of growth and provide predictions of the final yield. Generating such information early in the season would create opportunities for more efficient in-depth phenotyping and germplasm selection. This study tested the use of high-resolution time-series imagery (5 or 10 sampling dates) to understand the relationships between growth dynamics, temporal resolution and end-of-season above-ground biomass (AGB) in 869 diverse accessions of highly productive (mean AGB = 23.4 Mg/Ha), photoperiod sensitive sorghum. Canopy surface height (CSM), ground cover (GC), and five common spectral indices were considered as features of the crop phenotype. Spline curve fitting was used to integrate data from single flights into continuous time courses. Random Forest was used to predict end-of-season AGB from aerial imagery, and to identify the most informative variables driving predictions. Improved prediction of end-of-season AGB (RMSE reduction of 0.24 Mg/Ha) was achieved earlier in the growing season (10 to 20 days) by leveraging early- and mid-season measurement of the rate of change of geometric and spectral features. Early in the season, dynamic traits describing the rates of change of CSM and GC predicted end-of-season AGB best. Late in the season, CSM on a given date was the most influential predictor of end-of-season AGB. The power to predict end-of-season AGB was greatest at 50 days after planting, accounting for 63% of variance across this very diverse germplasm collection with modest error (RMSE 1.8 Mg/ha). End-of-season AGB could be predicted equally well when spline fitting was performed on data collected from five flights versus 10 flights over the growing season. This demonstrates a more valuable and efficient approach to using UAVs for HTP, while also proposing strategies to add further value.

54 ENVIRONMENTAL SCIENCES↗

Inter-year Variability in EDS Standards

This report provides a deep dive into how the stability of energy dispersive spectroscopy (EDS) standards yielded insight into the shortfalls of using software to randomly select points for spectral acquisition when using heterogeneous standards. Some procedural recommendations are included which are meant to minimize the impact of anomalous reference standard spectra by detecting them before the standard is implemented into data processing.

36 MATERIALS SCIENCE↗

Reassessing the MCNP Random Number Generator

Random number generators are integral components to Monte Carlo codes. They provide the pseudorandom number sequence used to actually sample the distributions of interest. As a result, they are one of the most important components to the software. The current recommended MCNP random number generator is a 63-bit linear congruential generator (LCG). This generator is quite fast, but it has some drawbacks. First, it only has a period of 2 63 . Due to the necessarily non-optimal usage of random numbers to ensure parallel reproducibility, this amount is too few to guarantee random number sequences are not reused in all configurations the code runs under. As simulation size increases, users will need to be aware of the limitations of the generator and tune configuration variables to best suit their simulations, or they will need to assume that reuse is not negatively affecting their answers. Neither of these are optimal. Second, small LCGs are fairly weak in bit generation quality, and this can have an unknown impact on the quality of the simulation. This paper is an investigation into whether or not more modern random number generators can supersede the current ones. The goal is to find a generator that is similar or superior in speed to the LCGs, has a state space large enough to make strong guarantees about random number reuse, and passes all modern random number test suites. If such a generator is found, it would eliminate the need for the user to even be aware of the limitations of the random number generator and would simplify the use of the code. This paper will be broken into several parts. Sec. 2 will discuss the evolution of the random number generator within the MCNP code. Sec. 3 will go over what a Monte Carlo code needs from a generator to be reproducible and portable and how the current generator behaves in that light. Sec. 4 goes through how each generator was tested. Finally, Sec. 5 will discuss improvements that could be made to the code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Iterative self-organizing SCEne-LEvel sampling (ISOSCELES) for large-scale building extraction

Convolutional neural networks (CNN) provide state-of-the-art performance in many computer vision tasks, including those related to remote-sensing image analysis. Successfully training a CNN to generalize well to unseen data, however, requires training on samples that represent the full distribution of variation of both the target classes and their surrounding contexts. With remote sensing data, acquiring a sufficiently representative training set is a challenge due to both the inherent multi-modal variability of satellite or aerial imagery and the general high cost of labeling data. To address this challenge, we have developed ISOSCELES, an Iterative Self-Organizing SCEne LEvel Sampling method for hierarchical sampling of large image sets. Using affinity propagation, ISOSCELES automates the selection of highly representative training images. Compared to random sampling or using available reference data, the distribution of the training is principally data driven, reducing the chance of oversampling uninformative areas or undersampling informative ones. In comparison to manual sample selection by an analyst, ISOSCELES exploits descriptive features, spectral and/or textural, and eliminates human bias in sample selection. Using a hierarchical sampling approach, ISOSCELES can obtain a training set that reflects both between-scene variability, such as in viewing angle and time of day, and within-scene variability at the level of individual training samples. We verify the method by demonstrating its superiority to stratified random sampling in the challenging task of adapting a pre-trained model to a new image and spatial domain for country-scale building extraction. Using a pair of hand-labeled training sets comprising 1,987 sample image chips, a total of 496,000,000 individually labeled pixels, we show, across three distinct model architectures, an increase in accuracy, as measured by F1-score, of 2.2–4.2%.

42 ENGINEERING↗

The LSST AGN Data Challenge: Selection Methods

Abstract Development of the Rubin Observatory Legacy Survey of Space and Time (LSST) includes a series of Data Challenges (DCs) arranged by various LSST Scientific Collaborations that are taking place during the project's preoperational phase. The AGN Science Collaboration Data Challenge (AGNSC-DC) is a partial prototype of the expected LSST data on active galactic nuclei (AGNs), aimed at validating machine learning approaches for AGN selection and characterization in large surveys like LSST. The AGNSC-DC took place in 2021, focusing on accuracy, robustness, and scalability. The training and the blinded data sets were constructed to mimic the future LSST release catalogs using the data from the Sloan Digital Sky Survey Stripe 82 region and the XMM-Newton Large Scale Structure Survey region. Data features were divided into astrometry, photometry, color, morphology, redshift, and class label with the addition of variability features and images. We present the results of four submitted solutions to DCs using both classical and machine learning methods. We systematically test the performance of supervised models (support vector machine, random forest, extreme gradient boosting, artificial neural network, convolutional neural network) and unsupervised ones (deep embedding clustering) when applied to the problem of classifying/clustering sources as stars, galaxies, or AGNs. We obtained classification accuracy of 97.5% for supervised models and clustering accuracy of 96.0% for unsupervised ones and 95.0% with a classic approach for a blinded data set. We find that variability features significantly improve the accuracy of the trained models, and correlation analysis among different bands enables a fast and inexpensive first-order selection of quasar candidates.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Monitoring covariance in multivariate time series: Comparing machine learning and statistical approaches

Abstract In complex systems with multiple variables monitored at high‐frequency, variables are not only temporally autocorrelated, but they may also be nonlinearly related or exhibit nonstationarity as the inputs or operation changes. One approach to handling such variables is to detrend them prior to monitoring and then apply control charts that assume independence and stationarity to the residuals. Monitoring controlled systems is even more challenging because the control strategy seeks to maintain variables at prespecified mean levels, and to compensate, correlations among variables may change, making monitoring the covariance essential. In this paper, a vector autoregressive model (VAR) is compared with a multivariate random forest (MRF) and a neural network (NN) for detrending multivariate time series prior to monitoring the covariance of the residuals using a multivariate exponentially weighted moving average (MEWMA) control chart. Machine learning models have an advantage when the data's structure is unknown or may change. We design a novel simulation study with nonlinear, nonstationary, and autocorrelated data to compare the different detrending models and subsequent covariance monitoring. The machine learning models have superior performance for nonlinear and strongly autocorrelated data and similar performance for linear data. An illustration with data from a reverse osmosis process is given.

Weix, Derek↗

Strong Correspondence in Evapotranspiration and Carbon Dioxide Fluxes Between Different Eddy Covariance Systems Enables Quantification of Landscape Heterogeneity in Dryland Fluxes

Abstract The eddy covariance method is widely used to investigate fluxes of energy, water, and carbon dioxide at landscape scales, providing important information on how ecological systems function. Flux measurements quantify ecosystem responses to environmental perturbations and management strategies, including nature‐based climate‐change mitigation measures. However, due to the high cost of conventional instrumentation, most eddy covariance studies employ a single system, limiting spatial representation to the flux footprint. Insufficient replication may be limiting our understanding of ecosystem behavior. To address this limitation, we deployed eight lower‐cost eddy covariance systems in two clusters around two conventional eddy covariance systems in the Chihuahuan Desert of North America for a period of 2 years. These dryland settings characterized by large temperature variations and relatively low carbon dioxide fluxes represented a challenging setting for eddy covariance. We found very good closure of energy and water balance across all systems (within ±9% of unity). We found very good correspondence between the lower‐cost and conventional systems' fluxes of sensible heat (with concordance correlation coefficient (CCC) of ≥0.87), latent energy (evapotranspiration; CCC ≥ 0.89), and useful correspondence in the net ecosystem exchange ((NEE); with CCC ≥ 0.4) at the daily temporal resolution. Relative to the conventional systems, the low‐frequency systems were characterized by a higher level of random error, particularly in the NEE fluxes. Lower‐cost systems can enable wider deployment affording better replication and sampling of spatiotemporal variability at the expense of greater measurement noise that might be limiting for certain applications. Replicated eddy covariance observations may be useful when addressing gaps in the existing monitoring of critical and underrepresented ecosystems and for measuring areas larger than a single flux footprint.

54 ENVIRONMENTAL SCIENCES↗

Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems

This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.

Google Engine↗

Synthesis of ARM User Facility Surface Rainfall Datasets to Construct a Best Estimate Value Added Product (PrecipBE)

Surface precipitation measurements are essential for Earth system model (ESM) evaluation and understanding cloud processes. An ever-growing need for robust, temporally evolving, and easy-to-use statistical datasets provides motivation for a baseline ground-based precipitation properties data product. The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility operates an extensive suite of precipitation instruments with various sensitivities and operating mechanisms, which render the decision of which instrument to use based on one or more fixed thresholds challenging and prone to errors and bias. Using a long-term instrument inter-comparison from a unique per-precipitation event perspective, rather than instantaneous sample comparison, we demonstrate that ARM rainfall-measuring instruments are generally consistent with each other at the statistical level. Inter-instrument deviations at the single event level can be large, especially for specific rainfall event properties such as maximum precipitation rates. A machine-learning (ML) analysis using a random forest regressor indicates that in some cases, depending on instrument, local site climatology, and/or specific deployment configuration, certain atmospheric state variables influence the measured quantities in an unpredictable manner. Thus, a-priori weighting of different instruments does not necessarily lead to more accurate and less biased synthesis of instrument data. These results motivate the design of the ARM precipitation best-estimate (PrecipBE) value-added product, which incorporates all valid precipitation data while considering data quality and other instrument limitations. PrecipBE consists of time series and tabular statistics datasets in an easy-to-use and insightful per-precipitation event format. It provides a large set of precipitation event properties supplemented with ancillary data from ARM datasets that correspond to the detected precipitation events. We describe the PrecipBE algorithm and demonstrate its use via the examination of a single-day output as well as a long-term trend analysis of precipitation events at the ARM Southern Great Plains (SGP) site, covering more than 30 years of data. The trend analysis tentatively suggests a long-term temporal tendency for mainly shorter and less intense precipitation events at the SGP site, but a long-term increase in annual rainfall by more than 36 mm (5 %) per decade. This rainfall trend is catalyzed primarily by more extreme event properties of relatively rare, intense precipitation events, with event total and 1 min maximum precipitation rate at a 1 year timeframe increasing up to 5 mm and 9 mm h −1 (several percent) per decade, respectively. While the currently available PrecipBE datasets (at https://adc.arm.gov/discovery/, last access: 8 December 2025) cover rainfall from multiple ARM deployments up to March 2025, PrecipBE is planned to be expanded to include solid-phase precipitation and will soon become an operational product with a several-day lag from real-time. We invite the ARM user community to leverage this new product and welcome user feedback to further enhance the dataset.

Silber, Israel [Pacific Northwest National Laborat↗

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

Combined Effects of Stream Hydrology and Land Use on Basin‐Scale Hyporheic Zone Denitrification in the Columbia River Basin

Abstract Denitrification in the hyporheic zone (HZ) of river corridors is crucial to removing excess nitrogen in rivers from anthropogenic activities. However, previous modeling studies of the effectiveness of river corridors in removing excess nitrogen via denitrification were often limited to the reach‐scale and low‐order stream watersheds. We developed a basin‐scale river corridor model for the Columbia River Basin with random forest models to identify the dominant factors associated with the spatial variation of HZ denitrification. Our modeling results suggest that the combined effects of hydrologic variability in reaches and substrate availability influenced by land use are associated with the spatial variability of modeled HZ denitrification at the basin scale. Hyporheic exchange flux can explain most of spatial variation of denitrification amounts in reaches of different sizes, while among the reaches affected by different land uses, the combination of hyporheic exchange flux and stream dissolved organic carbon (DOC) concentration can explain the denitrification differences. Also, we can generalize that the most influential watershed and channel variables controlling denitrification variation are channel morphology parameters (median grain size (D50), stream slope), climate (annual precipitation and evapotranspiration), and stream DOC‐related parameters (percent of shrub area). The modeling framework in our study can serve as a valuable tool to identify the limiting factors in removing excess nitrogen pollution in large river basins where direct measurement is often infeasible.

54 ENVIRONMENTAL SCIENCES↗

GraphAlign: Graph-Enabled Machine Learning for Seismic Event Filtering

This report summarizes results from a 2 year effort to improve the current automated seismic event processing system by leveraging machine learning models that can operated over the inherent graph data structure of a seismic sensor network. Specifically, the GraphAlign project seeks to utilize prior information on which stations are more likely to detect signals originating from particular geographic regions to inform event filtering. To date, the GraphAlign team has developed a Graphical Neural Network (GNN) model to filter out false events generated by the Global Associator (GA) algorithm. The algorithm operates directly on waveform data that has been associated to an event by building a variable sized graph of station waveforms nodes with edge relations to an event location node. This builds off of previous work where random forest models were used to do the same task using hand crafted features. The GNN model performance was analyzed using an 8 week IMS/IDC dataset, and it was demonstrated that the GNN outperforms the random forest baseline. We provide additional error analysis of which events the GNN model performs well and poorly against concluded by future directions for improvements.

58 GEOSCIENCES↗