Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multi-source data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

51 records · Page 3

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Power generation forecasting for solar plants based on Dynamic Bayesian networks by fusing multi-source information

A Dynamic Bayesian network (DBN) model for solar power generation forecasting in solar plants is proposed in this paper. The key idea is to fuse sensor data, operational indicators, meteorological data, lagged output power information, and model errors for more accurate short-term (e.g., hours) and mid-term (e.g., days to weeks) power generation forecasting. The proposed DBN augments automated data-driven structure learning with expert knowledge encoding using continuous and categorical data given constraints to represent causal relationships within a solar inverter system. Additionally, an error compensation mechanism is proposed to capture temporal fluctuation. The effectiveness of the DBN on solar power generation forecasting was evaluated by rolling window analysis with one-year testing data collected from a local solar plant. The proposed DBN is compared with four state-of-art methods including support-vector regression (SVR), k-nearest neighbors (kNN), artificial neural network (ANN), and long short-term memory (LSTM) models. The result show that the proposed DBN achieves better accuracy in general, and it is not as data-hungry as some neural network-based models. The proposed DBN is also shown to have robust and consistent forecasting power with different forecasting horizons. The accuracy is 92% - 95% from one hour to one week ahead forecasting.

14 SOLAR ENERGY↗

Distributed Thermal Response Multi-Source Modeling to Evaluate Heterogeneous Subsurface Properties

A thorough assessment of thermal properties in heterogeneous subsurface is necessary in design of low-temperature borehole heat exchangers (BHEs). For this study, a distributed thermal response test (DTRT), which combines distributed temperature sensing (DTS) with a conventional thermal response test (TRT), was conducted in a U-bend geothermal loop installed in an open borehole at the University of Illinois at Urbana-Champaign to estimate thermal properties by analyzing the thermal response of different geologic materials while applying a constant heat input rate. Fiber-optic cables in the DTRT were deployed both inside the U-bend geothermal loop and in the center of the borehole to improve the accuracy of calculated heat-loss rates and borehole temperature profile measurements. To assess the subsurface thermal conductivity during the heating phase of the DTRT, a single-source model and a multi-source model, both based on the infinite line source method, were developed using the borehole temperature data and temperatures inside and along the outside of the loop, separately. The two models returned similar thermal conductivity values. The multi-source modeling has the advantage of predicting the thermal conductivity of heterogeneous geologic materials from borehole temperature profiles during the DTRT heating phase. Additionally, based on the distributed thermal conductivity measured in the borehole, estimates were made for both radial thermal impacts and the rate of heat loss in the BHE.

58 GEOSCIENCES↗

Remote Sensing and GIS data at 1km-grid over Chesapeake Bay used in “He et al. 2024, Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem”

The package contains the data layers used in “He et al. 2024, Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem”. The study aims to use multi-source remote sensing and GIS datasets to investigate the spatial heterogeneity and identify spatial zones with similar environmental characteristics and understand the primary driving factors affecting soil respiration within sub-ecosystems of the coastal ecosystem. We employed unsupervised hierarchical clustering analysis to identify spatial regions with distinct environmental characteristics, then determined the main driving factors using Random Forest regression and SHapley Additive exPlanations (SHAP). Spatial data layers include soil respiration, kernel Normalized Difference Vegetation Index (kNDVI) computed from Harmonized Landsat 8 and Sentinel-2 time series, climate variables from the Daymet dataset, land cover, biodiversity, topographical metrics, soil property, and tidal elevation.

54 ENVIRONMENTAL SCIENCES↗

Detection of Synchrophasor False Data Injection Attack using Feature Interactive Network

The synchrophasor data recorded by Phasor Measurement Units (PMUs) plays an increasingly critical role in the regulation and situational awareness of power systems. However, the widely installed PMUs are vulnerable to multiple malicious attacks from cyber hackers during data transmission and storage. To address this problem, a Modified Ensemble Empirical Mode Decomposition (MEEMD) is proposed first to extract the intrinsic mode functions of each Synchrophasor Data Attacks (SDA). The frequency-based adaptive screening criterion embedded in MEEMD is used to eliminate the false intrinsic mode functions. Next, a Multivariate Convolutional Neural Network (MCNN) is proposed to identify multiple SDA by utilizing the extracted intrinsic mode functions and original SDA as input vectors. A fusion block as the main structure of MCNN is also leveraged to increase the diversity of features and compress the model parameters. Integrating MEEMD and MCNN, a framework with automatic feature extraction and multi-source information fusion capability, referred to as Feature Interactive Network (FIN), is proposed to detect multiple SDA. Based on the proposed FIN framework, six types of SDA are explored for the first time using actual synchrophasor data in FNET/Grideye that was collected from different locations in the U.S. Eastern Interconnection. Finally, a large quantity of experiments with different attack strengths are used to evaluate the adaptability and classification performance of the proposed FIN.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Sequential Bayesian inference of transition rates in the hidden Markov model for multi-state system degradation

The more easily available system performance data and advances in data analytics have provided us with opportunities to optimize maintenance programs for engineered systems, for example nuclear power plants. One key task in maintenance optimization is to obtain an accurate model for system degradation. Here, in this research, we propose a Bayesian method to address this problem. Noting that systems usually exhibit multiple states and that the actual state of a system usually is not directly observable, in the method we first model the system degradation process and the observation process based on a hidden Markov model. Then we develop a sequential Bayesian inference algorithm based on importance sampling and the forward algorithm to infer the posterior distributions of the transition rates in the hidden Markov model based on available observations. The proposed Bayesian method allows us to take advantage of evidence from multiple sources, and also allows us to perform Bayesian inference sequentially, without the need to use the entire history of observations every time new observations are collected. We demonstrate the proposed method using both synthetic data for a nuclear power plant feedwater pump and realistic data for a nuclear power plant chemistry analytical device.

42 ENGINEERING↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

A systematic analytical framework for multi-source municipal solid waste characterization for energy recovery

Advancing municipal solid waste (MSW) management from disposal-oriented practices toward circular, value-driven systems requires standardized methodologies capable of identifying material composition and resource recoverable potential at the point of generation. Despite extensive research, MSW characterization remains fragmented due to inconsistences in sampling methodologies, waste sorting categories, and temporal coverage across previous studies which limit cross-site comparability, reproducibility, and constrain the reliable evaluation of potential resource recovery pathways. This lack of consistency has hindered the development of a unified framework for MSW characterization and resource assessment. This study introduces a standardized, field-validated protocol for MSW sampling and composition analysis that ensures consistent, traceable data across diverse waste sources. The protocol integrates randomized spatial sampling, systematic material sorting, and controlled subsampling for multi-site and multi-season field campaigns. Validation included MSW collection from residential, grocery, restaurant, and school MSW streams across five U.S. states, including Maryland, Idaho, Virginia, Ohio, and Mississippi, to demonstrate the protocol’s ability to identify source-based composition patterns relevant to resource recovery applications. Grocery and restaurant streams were dominated by food waste and high-moisture organics, while school waste contained higher paper content and residential waste showed greater heterogeneity. Aggregation into energy-relevant fractions highlighted practical recovery pathways via anaerobic digestion or gasification, supporting data-driven planning, policy, and circular economy strategies for sustainable waste management across waste sources.

09 BIOMASS FUELS↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Multi-Sourced Collaboration for the Production and Refining of Rare Elements and Critical Metals (Final Technical Report)

The project objective was to develop a feasible and cost-effective method for recovering rare earth elements (REEs) and critical materials (CMs) from coal and coal byproducts, resulting in high-purity individually separated REEs and CMs. The targeted REEs included Y, Pr, Nd, Gd, Dy, and Sm, with a purity of over 99.5%, while the CMs included Co, Mn, Ga, Sr, Li, Ni, Zn, and Ge, with a purity of over 90%. The project aimed to design a prototype facility capable of producing 1-3 tonnes/day of high-purity REO mixes. The work was divided into four designated circuits: 1) REE extraction and concentration, 2) REE separation and purification, 3) RE metal production, and 4) CM production. To achieve these goals, the project involved 11 tasks, including technology reviews, research, process flow diagram development, mass balance estimation, and preliminary technical-economic analysis. The project team included researchers from the University of Kentucky, University of Alabama and Virginia Tech as well as process specialists from Argonne National Laboratory. MP Materials provided technical support regarding rare earth markets and processing while Alliance Coal performed resource assessment. The project included a market analysis for Nd/Pr, Tb, Dy, Gd, Y, Co, Mn, Li, Sr, Ga, Ni, Zn, and Ge. These analyses provided insights into the supply and demand trends as well as historic and future projections of market price relative to purity requirements for these elements. Two coal resources were selected for the project: the West Kentucky No. 13 (Baker) Seam and an undisclosed lignite resource in the Illinois coal basin. The estimated quantities of REEs in these resources were calculated based on production samples and drilling data. It was estimated that there is adequate supply for an operation producing one metric ton daily of higher purity mixed rare earth oxides (MREO) for approximately 20 years at a site located in western Kentucky. In Circuit 1, project data was obtained from a pilot heap leach and REE concentration facility. It was concluded that the existing circuit, which generated a MREO concentrate, two types of CM mixed products, and Li- and Sr-containing waters, would be suitable feed for circuits 2-4. Data from the first-of-its-kind coal coarse refuse heap leach pilot pad played a crucial role in estimating reliable elemental concentrations of the pregnant leaching solution (PLS). The average total REE concentration in the PLS was found to be 28.6 ppm. In Circuit 2, several concepts were explored including a novel process referred to as solvent-assisted chromatography (SAC). This concept involved a novel columnar reactor that incorporated multiple mixer/settlers, thereby enabling the operation of counter-flowing aqueous and organic phases. Unfortunately, due to project time constraints, a complete fundamental modeling analysis could not be completed to fully evaluate the technology. Molten salt electrowinning was considered as an alternative for circuit 3 following circuit 2 purification circuit utilizing the novel SAC process. A mass and energy balance of Nd reduction to metal in a fluoride containing molten salt electrolyte was conducted. Comparisons were made with the current state of Asian molten salt electrorefining, and potential improvements in siphoning rare earth metals (REM) from the reactor were presented. A cost estimate was performed for the production of 1 tonne per day, which yielded a total of $2.29 million for the nine electrowinning (EW) cells required. The selected option for circuits 2 and 3 was a plasma distillation process, which initially separates rare earth elements (REEs) from other elements. This is followed by selective electrowinning in various ionic liquids. The selection was made on the basis of thermodynamic modeling and experimental data previously published by a project partner. The combination offers an innovative approach to integrated refining and RE metal production. For Circuit 4, an extensive literature review was conducted for the processing of the CMs. The ultimate decision was to utilize a combined plasma and ionic liquid process as well to produce individual high-purity concentrates of Zn, Ni, Co, Mn, and Mg. A separate flowsheet for Li and Sr was recommended, which would yield carbonates of these elements. Due to the lack of suitable experimental data at this time, a process recommendation could not be provided but several methods have been proposed for consideration. Lastly, a techno-economic analysis (TEA) was conducted to assess the effectiveness of the proposed process for further investigation. The TEA results revealed a capital expense (CapEx) of $737 million and an annual operational expense (OpEx) of $220 million. Due to the selected elements, the hypothetical heap leach pad can produce 1 metric tonne per day of REO equivalent, but a conscious decision was made to only treat targeted REEs, resulting in the production of 0.4 metric tonne of REM. An estimated annual revenue of $90.87 million was projected based on standard market pricing information provided by the funding agency. During the TEA, ten different modules were evaluated for costing purposes. The precipitation circuit was identified as the largest single operational expense, followed by the Mg/Mn process due to the amount of treated metal. In terms of capital expenditures, the heap leach process incurred the highest cost, followed by the Mg/Mn process. The scalability of the plasma process is a crucial consideration since the reactors cannot be scaled beyond the largest demonstrated size due to their reliance on surface area of the slag and vapor phase. The purity estimate for the REEs are generally 98%±2% to produce a metal. The purity level being lower than the project objective was due to the lack of specific experimental data needed to tighten the tolerance of the estimates. Based on literature and previous experience, the CMs are estimated as follows; Ga (95%+, metal), Sr (95%+, carbonate), Li (95%+, carbonate), Ni (98%±2%, metal), Zn (95%+, metal sponge), Ge (95%+, metal), Co (98%±2%, metal), and Mn (98%±2%, metal).

01 COAL, LIGNITE, AND PEAT↗

Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonization

Objective: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. Methods: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. Results: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. Conclusions: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes.

Zhou, Doudou↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Evaluating the feasibility of using downwind methods to quantify point source oil and gas emissions using continuously monitoring fence-line sensors

The dependable reporting of methane (CH 4 ) emissions from point sources, such as fugitive leaks from oil and gas infrastructure, is important for profit maximization (retaining more hydrocarbons), evaluating climate impacts, assessing CH 4 fees for regulatory programs, and validating CH 4 intensity in differentiated gas programs. Currently, there are disagreements between emissions reported by different quantification techniques for the same sources. It has been suggested that downwind CH 4 quantification methods using CH 4 measurements on the fence line of production facilities could be used to generate emission estimates from oil and gas operations at the site level, but it is currently unclear how accurate the quantified emissions are. To investigate the accuracy of downwind methods, this study uses fence-line simulated data collected during controlled-release experiments as input for a non-standard closed-path eddy covariance (EC), the Gaussian plume inverse model (GPIM), and the backward Lagrangian stochastic (bLs) model in a range of atmospheric conditions. This study's EC attempt was unsuccessful due to data collection and instrumentation issues, resulting in invalid results characterized by underestimated emissions, large negative fluxes, and cospectra/ogives that deviated from their ideal shapes. Consequently, the EC results could not be compared with the GPIM and bLS model. The bLs model demonstrated the highest accuracy for single-release single-point emissions, though it exhibited greater uncertainty than GPIM under multi-release conditions. Across the GPIM and bLs model, the most reliable quantification was achieved with 15 min averaging and a narrow 5° wind sector range. Although EC was limited in this context, future studies should consider employing a standard EC system and further optimizing GPIM and bLs approaches – particularly for complex multi-source scenarios – to enhance quantification accuracy and reduce uncertainty.

03 NATURAL GAS↗