Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multi-source data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Tree-level carbon stock estimations across diverse species using multi-source remote sensing integration

Forests are critical carbon sinks, and remote sensing has been increasingly widely used for forest monitoring and biomass estimations. However, species-specific tree-level studies remain limited. In this study, we demonstrated the feasibility of integrating UAV-based LiDAR with high-resolution optical satellite imagery (0.5 m) to estimate biomass for individual trees across different species. The proposed method accurately estimated biomass for 53 trees (R² = 0.82, rRMSE = 0.44), with species-specific datasets, showing an average 25.2% increase in R² and a 14.8% reduction in rRMSE. A novel vegetation index combining forest structure parameters with vegetation indices (VIs) was developed using high-resolution multispectral satellite data (3 m) to explore its relationship with individual tree biomass. Combining forest structural parameters with VIs further improved estimation accuracy, achieving an R²of 0.89 and an rRMSE of 0.34. Species-specific datasets show an 11.6% increase in R²compared to methods without VIs, and a 22.2% improvement over methods using only VIs. SHapley Additive exPlanations (SHAP) analysis shows that the volume feature played a key role in model performance and remained stable throughout the training process. Altogether, the proposed approach enhances individual tree biomass and carbon sink estimations, showing great potential for large-scale precise forest carbon monitoring using multi-source remote sensing data.

59 BASIC BIOLOGICAL SCIENCES↗

Mapping Pine Plantations in the Southeastern U.S. Using Structural, Spectral, and Temporal Remote Sensing Data

The southeastern U.S. produces the most industrial roundwood in the U.S. each year, largely from commercial pine plantations. The extent of plantation forests and management dynamics can be difficult to ascertain from periodic forest inventories, yet short-rotation tree plantations also present challenges for remote sensing. Here, we integrated spectral, temporal, and structural information from airborne and satellite platforms to distinguish pine plantations from natural forests and evaluate the contribution from planted forests to regional forest coverin the southeastern U.S. Within flight lines from NASA Goddard's Lidar, Hyperspectral, and Thermal (G-LiHT) Airborne Imager, lidar metrics of forest structure had the highest overall accuracy for pine plantations among single-source classifications (90%), but the combination of spectral and temporal metrics from Landsat generated comparable accuracy (91%). Combined structural, temporal, and spectral information from G-LiHT and Landsat had the highest accuracy for plantations (92%) and natural forests (88%). At a regional scale, classifications using Landsat spectral and temporal metrics had between 74 and 82% mean class accuracy for plantations.Regionally, plantations accounted for 28% of forest cover in the southeastern U.S., a result similar to plot-based estimates, albeit with greater spatial detail. Regional maps of plantation forests differed from existing map products, including the National Land Cover Database. Combining plantation extent in 2011 with Landsat based forest change data identified strong regional gradients in plantation dynamics since 1985, with distinct spatial patterns of rotation age (east-west) and plantation expansion (interior). Our analysis demonstrates the potential to improve the characterization of dynamic land cover classes, including economically important timber plantations, by integrating diverse remote sensing datasets. Critically, multi-source remote sensing provides an approach to leverage periodic forest inventory data for annual monitoring of managed forest landscapes.

Fagan, M. E.↗

Retrieval of biophysical parameters with AVIRIS and ISM: The Landes Forest, south west France

The first steps of an experiment for investigating the capability of airborne spectrometer data for retrieval of biophysical parameters of vegetation, especially water conditions are presented. Airborne Visible/Infrared Imaging Spectrometer (AVIRIS) and ISM data were acquired in the frame of the 1991 NASA/JPL and CNES campaigns on the Landes, South west France, a large and flat forest area with mainly maritime pines. In-situ measurements were completed at that time; i.e. reflectance spectra, atmospheric profiles, sampling for further laboratory analyses of elements concentrations (lignin, water, cellulose, nitrogen,...). All information was integrated in an already existing data base (age, LAI, DBH, understory cover,...). A methodology was designed for (1) obtaining geometrically and atmospherically corrected reflectance data, (2) registering all available information, and (3) analyzing these multi-source informations. Our objective is to conduct comparative studies with simulation reflectance models, and to improve these models, especially in the MIR.

Zagolski, F.↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Space Weathering Investigations Enabled by NASA's Virtual Heliophysical Observatories

Structural and chemical impact of the heliospheric space environment on exposed planetary surfaces and interplanetary dust grains may be generally defined as space weathering . In the inner solar system, from the asteroid belt inwards towards the Sun, the surface regolith structures of airless bodies are primarily determined by cumulative meteoritic impacts over billions of years, but the molecular composition to meters in depth can be substantially modified by irradiation effects. Plasma ions at eV to keV energies may both erode uppermost surfaces by sputtering, and implant or locally produce exogenic material, e.g. He-3 and H2O, while more energetic ions drive molecular change through electronic ionization. Galactic cosmic ray ions and more energetic solar ions can impact chemistry to meters in depth. High energy cosmic ray interactions produce showers of secondary particles and energetic photons that present hazards for robotic and human exploration missions but also enable detection of potentially useable resources such as water ice, oxygen, and many other elements. Surface sputtering also makes ejected elemental and molecular species accessible for in-situ compositional analysis by spacecraft with ion and neutral mass spectrometers. Modeling of relative impacts for these various space weathering processes requires knowledge of the incident species-resolved ion flux spectra at plasma to cosmic ray energies and as integrated over varying time scales. Although the main drivers for investigations of these processes come from NASA's planetary science and human exploration programs, the NASA heliophysics program provides the requisite data measurement and modeling resources to enable specification of the field & plasma and energetic particle irradiation environments for application to space weather and surface weathering investigations. The Virtual Heliospheric Observatory (VHO), Virtual Energetic Particle Observatory (VEPO), Lunar Solar Origins Exploration (LunaSOX), and Space Physics Data Facility (SPDF) services now provide a wide range of inner heliospheric spacecraft data that can be applied to space weathering of potential exploration destinations including the Moon, asteroids, and the moons of Mars, as well to radiation hazard assessment for the spacecraft and human explorers. For example, the new VEPO service for time-averaging of multi-source ion flux spectra enables the specification of composite flux spectra from a variety of ongoing and legacy missions for applications to surface interaction modeling. Apollo to Artemis data resources of LunaSOX enable specific space weathering investigations for the Moon, while VHO more generally covers the space field and plasma environments of the inner and outer solar system from the sunward-most perihelia of the twin Helios spacecraft to the ongoing heliosheath passages of the twin Voyagers. Composite multi-source spectra from VEPO can also be applied to the continuing compilation of accumulated 1-AU fluence spectra, mostly contributed by solar wind plasma and energetic particle events, for determination of time-averaged particle compositional and kinetic energy output from the Sun and for modeling of long-term irradiation impacts on planetary surfaces.

Cooper, John F.↗

NCC: A Multidisciplinary Design/Analysis Tool for Combustion Systems

A multi-disciplinary design/analysis tool for combustion systems is critical for optimizing the low-emission, high-performance combustor design process. Based on discussions between NASA Lewis Research Center and the jet engine companies, an industry-government team was formed in early 1995 to develop the National Combustion Code (NCC), which is an integrated system of computer codes for the design and analysis of combustion systems. NCC has advanced features that address the need to meet designer's requirements such as "assured accuracy", "fast turnaround", and "acceptable cost". The NCC development team is comprised of Allison Engine Company (Allison), CFD Research Corporation (CFDRC), GE Aircraft Engines (GEAE), NASA Lewis Research Center (LeRC), and Pratt & Whitney (P&W). This development team operates under the guidance of the NCC steering committee. The "unstructured mesh" capability and "parallel computing" are fundamental features of NCC from its inception. The NCC system is composed of a set of "elements" which includes grid generator, main flow solver, turbulence module, turbulence and chemistry interaction module, chemistry module, spray module, radiation heat transfer module, data visualization module, and a post-processor for evaluating engine performance parameters. Each element may have contributions from several team members. Such a multi-source multi-element system needs to be integrated in a way that facilitates inter-module data communication, flexibility in module selection, and ease of integration.

Liu, Nan-Suey↗

Overview of the NCC

A multi-disciplinary design/analysis tool for combustion systems is critical for optimizing the low-emission, high-performance combustor design process. Based on discussions between then NASA Lewis Research Center and the jet engine companies, an industry-government team was formed in early 1995 to develop the National Combustion Code (NCC), which is an integrated system of computer codes for the design and analysis of combustion systems. NCC has advanced features that address the need to meet designer's requirements such as "assured accuracy", "fast turnaround", and "acceptable cost". The NCC development team is comprised of Allison Engine Company (Allison), CFD Research Corporation (CFDRC), GE Aircraft Engines (GEAE), NASA Glenn Research Center (LeRC), and Pratt & Whitney (P&W). The "unstructured mesh" capability and "parallel computing" are fundamental features of NCC from its inception. The NCC system is composed of a set of "elements" which includes grid generator, main flow solver, turbulence module, turbulence and chemistry interaction module, chemistry module, spray module, radiation heat transfer module, data visualization module, and a post-processor for evaluating engine performance parameters. Each element may have contributions from several team members. Such a multi-source multi-element system needs to be integrated in a way that facilitates inter-module data communication, flexibility in module selection, and ease of integration. The development of the NCC beta version was essentially completed in June 1998. Technical details of the NCC elements are given in the Reference List. Elements such as the baseline flow solver, turbulence module, and the chemistry module, have been extensively validated; and their parallel performance on large-scale parallel systems has been evaluated and optimized. However the scalar PDF module and the Spray module, as well as their coupling with the baseline flow solver, were developed in a small-scale distributed computing environment. As a result, the validation of the NCC beta version as a whole was quite limited. Current effort has been focused on the validation of the integrated code and the evaluation/optimization of its overall performance on large-scale parallel systems.

Liu, Nan-Suey↗

XTE J1550-564: INTEGRAL Observations of a Failed Outburst

The well known black-hole X-ray binary transient XTE J1550-564 underwent an outburst during the spring of 2003 which was substantially underluminous in comparison to previous periods of peak activity in that source. In addition, our analysis shows that it apparently remained in the hard spectral state over the duration of that outburst. This is again in sharp contrast to major out-bursts of that source in 1998/1999 during which it exhibited an irregular light curve, multiple state changes and collimated outflows. This leads us to classify it as a failed outburst. We present the results of our study of the spring 2003 event including light curves based on observations from both INTEGRAL and RXTE. In addition, we studied the evolution of the high-energy 3-300 keV continuum spectrum using data obtained with three main instruments on INTEGRAL. These spectra are consistent with typical low-hard-state thermal Comptonization emission. We also consider the 2003 event in the context of a multi-source, multi-event period-peak luminosity diagram in which it is a clear outlyer. We then consider the possibility that the 2003 event was due to a discrete accretion event rather than a limit-cycle instability. In that context, apply model fitting to derive the timescale for viscous propagation in the disk, and infer some physical characteristics.

Sturner, S. J.↗

NASA Tech Briefs, January 2014

Topics include: Multi-Source Autonomous Response for Targeting and Monitoring of Volcanic Activity; Software Suite to Support In-Flight Characterization of Remote Sensing Systems; Visual Image Sensor Organ Replacement; Ultra-Wideband, Dual-Polarized, Beam-Steering P-Band Array Antenna; Centering a DDR Strobe in the Middle of a Data Packet; Using a Commercial Ethernet PHY Device in a Radiation Environment; Submerged AUV Charging Station; Habitat Demonstration Unit (HDU) Vertical Cylinder Habitat; Origami-Inspired Folding of Thick, Rigid Panels; A Novel Protocol for Decoating and Permeabilizing Bacterial Spores for Epifluorescent Microscopy; Method and Apparatus for Automated Isolation of Nucleic Acids from Small Cell Samples; Enabling Microliquid Chromatography by Microbead Packing of Microchannels; On-Command Force and Torque Impeding Devices (OC-FTID) Using ERF; Deployable Fresnel Rings; Transition-Edge Hot-Electron Microbolometers for Millimeter and Submillimeter Astrophysics; Spacecraft Trajectory Analysis and Mission Planning Simulation (STAMPS) Software; Cross Support Transfer Service (CSTS) Framework Library; Arbitrary Shape Deformation in CFD Design; Range Safety Flight Elevation Limit Calculation Software; Frequency-Modulated, Continuous-Wave Laser Ranging Using Photon-Counting Detectors; Calculation of Operations Efficiency Factors for Mars Surface Missions; GPU Lossless Hyperspectral Data Compression System; Robust, Optimal Subsonic Airfoil Shapes; Protograph-Based Raptor-Like Codes; Fuzzy Neuron: Method and Hardware Realization; Kalman Filter Input Processor for Boresight Calibration; Organizing Compression of Hyperspectral Imagery to Allow Efficient Parallel Decompression; and Temperature Dependences of Mechanisms Responsible for the Water-Vapor Continuum Absorption.

Source record↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

A systematic analytical framework for multi-source municipal solid waste characterization for energy recovery

Advancing municipal solid waste (MSW) management from disposal-oriented practices toward circular, value-driven systems requires standardized methodologies capable of identifying material composition and resource recoverable potential at the point of generation. Despite extensive research, MSW characterization remains fragmented due to inconsistences in sampling methodologies, waste sorting categories, and temporal coverage across previous studies which limit cross-site comparability, reproducibility, and constrain the reliable evaluation of potential resource recovery pathways. This lack of consistency has hindered the development of a unified framework for MSW characterization and resource assessment. This study introduces a standardized, field-validated protocol for MSW sampling and composition analysis that ensures consistent, traceable data across diverse waste sources. The protocol integrates randomized spatial sampling, systematic material sorting, and controlled subsampling for multi-site and multi-season field campaigns. Validation included MSW collection from residential, grocery, restaurant, and school MSW streams across five U.S. states, including Maryland, Idaho, Virginia, Ohio, and Mississippi, to demonstrate the protocol’s ability to identify source-based composition patterns relevant to resource recovery applications. Grocery and restaurant streams were dominated by food waste and high-moisture organics, while school waste contained higher paper content and residential waste showed greater heterogeneity. Aggregation into energy-relevant fractions highlighted practical recovery pathways via anaerobic digestion or gasification, supporting data-driven planning, policy, and circular economy strategies for sustainable waste management across waste sources.

09 BIOMASS FUELS↗

Global Ocean Evaporation Increases Since 1960 in Climate Reanalyses: How Accurate Are They?

AGCMs w/ Specified SSTs (AMIPs) GEOS-5, ERA-20CM Ensembles Incorporate best historical estimates of SST, sea ice, radiative forcing Atmospheric "weather noise" is inconsistent with specified SST. Instantaneous Sfc fluxes can be wrong sign (e.g. Indian Ocean Monsoon, high latitude oceans). Averaging over ensemble members helps isolate SST-forced signal. Reduced Observational Reanalyses: NOAA 20CR V2C, ERA-20C, JRA-55C Incorporate observed Sfc Press (20CR), Marine Winds (ERA-20C) and rawinsondes (JRA-55C) to recover much of true synoptic or weather w/o shock of new sat obs. Comprehensive Reanalyses (MERRA-2) Full suite of observational constraints- both conventional and remote sensing. But... substantial uncertainties owing to evolving satellite observing system. Multi-source Statistically Blended OAFlux, LargeYeager Blend reanalysis, satellite, and ocean buoy information. While climatological biases are removed, non-physical trends or variations in components remain. Satellite Retrievals GSSTF3, SeaFlux, HOAPS3... Global coverage. Retrieved near sfc wind speed, & humidity used with SST to drive accurate bulk aerodynamic flux estimates. Satellite inter-calibration, spacecraft pointing variations crucial. Short record ( late 1987-present). In situ Measurements ICOADS, IVAD, Res Cruises VOS and buoys offer direct measurements. Sparse data coverage (esp south of 30S. Changes in measurement techniques (e.g. shipboard anemometer height).

Climate↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Historical (1700–2012) Global Multi-Model Estimates of the Fire Emissions from the Fire Modeling Intercomparison Project (FireMIP)

Fire emissions are critical for carbon and nutrient cycles, climate, and air quality. Dynamic Global Vegetation Models (DGVMs) with interactive fire modeling provide important estimates for long-term and large-scale changes of fire emissions. Here we present the first multi-model estimates of global gridded historical fire emissions for 1700-2012, including carbon and 33 species of trace gases and aerosols. The dataset is based on simulations of nine DGVMs with different state-of-the-art global fire models that participated in the Fire Modeling Intercomparison Project (FireMIP), using the same and standardized protocols and forcing data, and the most up-to-date fire emission factor table from field and laboratory studies over various land cover types. We evaluate the simulations of present-day fire emissions by comparing them with satellite-based products. Evaluation results show that most DGVMs simulate present-day global fire emission totals within the range of satellite-based products, and can capture the high emissions over the tropical savannas, low emissions over the arid and sparsely vegetated regions, and the main features of seasonality. However, most of the models fail to simulate the interannual variability, partly due to a lack of modeling peat fires and tropical deforestation fires. Historically, all models show only a weak trend in global fire emissions before ~1850s, consistent with multi-source merged historical reconstructions. The long-term trends among DGVMs are quite different for the 20th century, with some models showing an increase and others a decrease in fire emissions, mainly as a result of the discrepancy in their simulated responses to human population density change and land-use and land-cover change (LULCC). Our study provides a basic dataset for developing regional and global multi-source merged historical reconstructions and merging methods, and analyzing historical changes of fire emissions and their uncertainties as well as their role in the Earth system. It also highlights the importance of accurately modeling the responses of fire emissions to LULCC and population density change in reducing uncertainties in historical reconstructions of fire emissions and providing more reliable future projections.

Li, Fang↗

Improved Hypoxia Modeling for Nutrient Control Decisions in the Gulf of Mexico

As required by the Harmful Algal Bloom and Hypoxia Research Control Act of 1998, the Mississippi River/Gulf of Mexico Watershed Nutrient Task Force issued the 2001 Gulf Hypoxia Action Plan (updated in 2008). In response to the Gulf Hypoxia Action Plan of 2001 (updated in 2008), the EPA Gulf of Mexico Hypoxia Modeling and Monitoring Project has established a detailed model for the Mississippi-Attchafalaya River Basin which provides a capability to forecast the multi-source nutrient loading to the Gulf and the subsequent bio-geochemical processes leading to hypoxic conditions and subsequent effects on Gulf habitats and fisheries. The primary purpose of the EPA model is to characterize the impacts of nutrient management actions, or proposed actions on the spatial and temporal characteristics of the Gulf hypoxic zone. The model is expected to play a significant role in determining best practices and improved strategies for incentivizing nutrient reduction strategies, including installation of on-farm structures to reduce sediment and nutrient runoff, use of cover crops and other agricultural practices, restoration of wetlands and riparian buffers, improved waste water treatment and decreased industrial nitrogen emissions. These decisions are currently made in a fragmented way by federal, state, and local agencies, using a variety of small scale models and limited data. During the past three years, EPA has collected an enormous amount of in-situ data to be used in the model. We believe that the use of NASA satellite data products in the model and for long term validation of the model has the potential to significantly increase the accuracy and therefore the utility of the model for the decision making described above. This proposal addresses the Gulf of Mexico Alliance (GOMA) priority issue of reductions in nutrient inputs to coastal ecosystem. It further directly relates to water quality for healthy beaches and shellfish beds and wetland and coastal conservation restoration.

Habib, Shaid↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Evaluating the feasibility of using downwind methods to quantify point source oil and gas emissions using continuously monitoring fence-line sensors

The dependable reporting of methane (CH 4 ) emissions from point sources, such as fugitive leaks from oil and gas infrastructure, is important for profit maximization (retaining more hydrocarbons), evaluating climate impacts, assessing CH 4 fees for regulatory programs, and validating CH 4 intensity in differentiated gas programs. Currently, there are disagreements between emissions reported by different quantification techniques for the same sources. It has been suggested that downwind CH 4 quantification methods using CH 4 measurements on the fence line of production facilities could be used to generate emission estimates from oil and gas operations at the site level, but it is currently unclear how accurate the quantified emissions are. To investigate the accuracy of downwind methods, this study uses fence-line simulated data collected during controlled-release experiments as input for a non-standard closed-path eddy covariance (EC), the Gaussian plume inverse model (GPIM), and the backward Lagrangian stochastic (bLs) model in a range of atmospheric conditions. This study's EC attempt was unsuccessful due to data collection and instrumentation issues, resulting in invalid results characterized by underestimated emissions, large negative fluxes, and cospectra/ogives that deviated from their ideal shapes. Consequently, the EC results could not be compared with the GPIM and bLS model. The bLs model demonstrated the highest accuracy for single-release single-point emissions, though it exhibited greater uncertainty than GPIM under multi-release conditions. Across the GPIM and bLs model, the most reliable quantification was achieved with 15 min averaging and a narrow 5° wind sector range. Although EC was limited in this context, future studies should consider employing a standard EC system and further optimizing GPIM and bLs approaches – particularly for complex multi-source scenarios – to enhance quantification accuracy and reduce uncertainty.

03 NATURAL GAS↗