Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multi-source data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

XTE J1550-564: INTEGRAL Observations of a Failed Outburst

The well known black-hole X-ray binary transient XTE J1550-564 underwent an outburst during the spring of 2003 which was substantially underluminous in comparison to previous periods of peak activity in that source. In addition, our analysis shows that it apparently remained in the hard spectral state over the duration of that outburst. This is again in sharp contrast to major out-bursts of that source in 1998/1999 during which it exhibited an irregular light curve, multiple state changes and collimated outflows. This leads us to classify it as a failed outburst. We present the results of our study of the spring 2003 event including light curves based on observations from both INTEGRAL and RXTE. In addition, we studied the evolution of the high-energy 3-300 keV continuum spectrum using data obtained with three main instruments on INTEGRAL. These spectra are consistent with typical low-hard-state thermal Comptonization emission. We also consider the 2003 event in the context of a multi-source, multi-event period-peak luminosity diagram in which it is a clear outlyer. We then consider the possibility that the 2003 event was due to a discrete accretion event rather than a limit-cycle instability. In that context, apply model fitting to derive the timescale for viscous propagation in the disk, and infer some physical characteristics.

Sturner, S. J.↗

NASA Tech Briefs, January 2014

Topics include: Multi-Source Autonomous Response for Targeting and Monitoring of Volcanic Activity; Software Suite to Support In-Flight Characterization of Remote Sensing Systems; Visual Image Sensor Organ Replacement; Ultra-Wideband, Dual-Polarized, Beam-Steering P-Band Array Antenna; Centering a DDR Strobe in the Middle of a Data Packet; Using a Commercial Ethernet PHY Device in a Radiation Environment; Submerged AUV Charging Station; Habitat Demonstration Unit (HDU) Vertical Cylinder Habitat; Origami-Inspired Folding of Thick, Rigid Panels; A Novel Protocol for Decoating and Permeabilizing Bacterial Spores for Epifluorescent Microscopy; Method and Apparatus for Automated Isolation of Nucleic Acids from Small Cell Samples; Enabling Microliquid Chromatography by Microbead Packing of Microchannels; On-Command Force and Torque Impeding Devices (OC-FTID) Using ERF; Deployable Fresnel Rings; Transition-Edge Hot-Electron Microbolometers for Millimeter and Submillimeter Astrophysics; Spacecraft Trajectory Analysis and Mission Planning Simulation (STAMPS) Software; Cross Support Transfer Service (CSTS) Framework Library; Arbitrary Shape Deformation in CFD Design; Range Safety Flight Elevation Limit Calculation Software; Frequency-Modulated, Continuous-Wave Laser Ranging Using Photon-Counting Detectors; Calculation of Operations Efficiency Factors for Mars Surface Missions; GPU Lossless Hyperspectral Data Compression System; Robust, Optimal Subsonic Airfoil Shapes; Protograph-Based Raptor-Like Codes; Fuzzy Neuron: Method and Hardware Realization; Kalman Filter Input Processor for Boresight Calibration; Organizing Compression of Hyperspectral Imagery to Allow Efficient Parallel Decompression; and Temperature Dependences of Mechanisms Responsible for the Water-Vapor Continuum Absorption.

Source record↗

Sequential Bayesian inference of transition rates in the hidden Markov model for multi-state system degradation

The more easily available system performance data and advances in data analytics have provided us with opportunities to optimize maintenance programs for engineered systems, for example nuclear power plants. One key task in maintenance optimization is to obtain an accurate model for system degradation. Here, in this research, we propose a Bayesian method to address this problem. Noting that systems usually exhibit multiple states and that the actual state of a system usually is not directly observable, in the method we first model the system degradation process and the observation process based on a hidden Markov model. Then we develop a sequential Bayesian inference algorithm based on importance sampling and the forward algorithm to infer the posterior distributions of the transition rates in the hidden Markov model based on available observations. The proposed Bayesian method allows us to take advantage of evidence from multiple sources, and also allows us to perform Bayesian inference sequentially, without the need to use the entire history of observations every time new observations are collected. We demonstrate the proposed method using both synthetic data for a nuclear power plant feedwater pump and realistic data for a nuclear power plant chemistry analytical device.

42 ENGINEERING↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

A systematic analytical framework for multi-source municipal solid waste characterization for energy recovery

Advancing municipal solid waste (MSW) management from disposal-oriented practices toward circular, value-driven systems requires standardized methodologies capable of identifying material composition and resource recoverable potential at the point of generation. Despite extensive research, MSW characterization remains fragmented due to inconsistences in sampling methodologies, waste sorting categories, and temporal coverage across previous studies which limit cross-site comparability, reproducibility, and constrain the reliable evaluation of potential resource recovery pathways. This lack of consistency has hindered the development of a unified framework for MSW characterization and resource assessment. This study introduces a standardized, field-validated protocol for MSW sampling and composition analysis that ensures consistent, traceable data across diverse waste sources. The protocol integrates randomized spatial sampling, systematic material sorting, and controlled subsampling for multi-site and multi-season field campaigns. Validation included MSW collection from residential, grocery, restaurant, and school MSW streams across five U.S. states, including Maryland, Idaho, Virginia, Ohio, and Mississippi, to demonstrate the protocol’s ability to identify source-based composition patterns relevant to resource recovery applications. Grocery and restaurant streams were dominated by food waste and high-moisture organics, while school waste contained higher paper content and residential waste showed greater heterogeneity. Aggregation into energy-relevant fractions highlighted practical recovery pathways via anaerobic digestion or gasification, supporting data-driven planning, policy, and circular economy strategies for sustainable waste management across waste sources.

09 BIOMASS FUELS↗

Global Ocean Evaporation Increases Since 1960 in Climate Reanalyses: How Accurate Are They?

AGCMs w/ Specified SSTs (AMIPs) GEOS-5, ERA-20CM Ensembles Incorporate best historical estimates of SST, sea ice, radiative forcing Atmospheric "weather noise" is inconsistent with specified SST. Instantaneous Sfc fluxes can be wrong sign (e.g. Indian Ocean Monsoon, high latitude oceans). Averaging over ensemble members helps isolate SST-forced signal. Reduced Observational Reanalyses: NOAA 20CR V2C, ERA-20C, JRA-55C Incorporate observed Sfc Press (20CR), Marine Winds (ERA-20C) and rawinsondes (JRA-55C) to recover much of true synoptic or weather w/o shock of new sat obs. Comprehensive Reanalyses (MERRA-2) Full suite of observational constraints- both conventional and remote sensing. But... substantial uncertainties owing to evolving satellite observing system. Multi-source Statistically Blended OAFlux, LargeYeager Blend reanalysis, satellite, and ocean buoy information. While climatological biases are removed, non-physical trends or variations in components remain. Satellite Retrievals GSSTF3, SeaFlux, HOAPS3... Global coverage. Retrieved near sfc wind speed, & humidity used with SST to drive accurate bulk aerodynamic flux estimates. Satellite inter-calibration, spacecraft pointing variations crucial. Short record ( late 1987-present). In situ Measurements ICOADS, IVAD, Res Cruises VOS and buoys offer direct measurements. Sparse data coverage (esp south of 30S. Changes in measurement techniques (e.g. shipboard anemometer height).

Climate↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Earth Science Technology Office (ESTO) New Observing Strategies (NOS) and NOS-Testbed (NOS-T)

With the advancement of space hardware technologies such as smaller spacecraft, component and instrument miniaturization and high performance space processors, and with the advancement of software technologies in artificial intelligence, big data analysis and autonomous decision making, Earth Science is looking at novel ways to observe phenomena that previously could not have been studied or would have been too expensive to study with traditional missions. In particular, the New Observing Strategies (NOS) component of the NASA Earth Science Technology Office (ESTO) Advanced Information Systems Technology (AIST) Program aims at leveraging these novel technologies as well as low cost and easy access to space to acquire multi-temporal or simultaneous multi-angular, multi-locations, multi-resolution and multi-spectral observations that will provide better multi-source measurements and will build a more dynamic and comprehensive picture of Earth Science phenomena that need to be studied and analyzed. For applications such as water resources management, air quality monitoring, biodiversity studies or disaster management, NOS will integrate the use of small instruments, small spacecraft, constellations of spacecraft and networks of sensors to design new missions that will provide the necessary measurements to improve future forecast and science modeling systems.Measurement acquisition will therefore be approached as a system of systems rather than on a mission basis, and a system of this complexity should not be expected to work without full integration and experimental characterization. Although most of the individual technologies enabling to link and coordinate multi-source observations are more or less mature, a few technologies need to be developed and all of them need to be integrated and tested as a system. In order for this validation to occur, the AIST Program is developing the NOS Testbed that includes 3 main goals:1.Validate novel NOS technologies, independently and as a system2.Demonstrate novel distributed operations concepts3.Socialize new Distributed Spacecraft Mission (DSM) and SensorWeb (SW) technologies and concepts to the science community by significantly retiring the risk of integrating these new technologies.The NOS Testbed will consist of multiple sensing nodes, simulated or actual, representing space, air and/or ground measurements, that are interconnected by a communications fabric (infrastructure that permits nodes to transmit and receive data between one another and interact with each other). Each node will be supported by hardware capabilities required to perform nodes monitoring and command & control, as well as intelligent "onboard" computing. The nodes will work together in a collaborative manner to demonstrate optimal science capabilities. The testbed will enable to validate technologies such as inter-node communication models, techniques and protocols; inter-node coordination; real-time data fusion and understanding; planning; sensor re-targeting; etc. Additionally, the testbed will have the capability to interact with various mission design tools, OSSEs and one or several forecast models. More details about the NOS Testbed will be presented at the confererence.

Earth Science missions; Advanced information Syste↗

Historical (1700–2012) Global Multi-Model Estimates of the Fire Emissions from the Fire Modeling Intercomparison Project (FireMIP)

Fire emissions are critical for carbon and nutrient cycles, climate, and air quality. Dynamic Global Vegetation Models (DGVMs) with interactive fire modeling provide important estimates for long-term and large-scale changes of fire emissions. Here we present the first multi-model estimates of global gridded historical fire emissions for 1700-2012, including carbon and 33 species of trace gases and aerosols. The dataset is based on simulations of nine DGVMs with different state-of-the-art global fire models that participated in the Fire Modeling Intercomparison Project (FireMIP), using the same and standardized protocols and forcing data, and the most up-to-date fire emission factor table from field and laboratory studies over various land cover types. We evaluate the simulations of present-day fire emissions by comparing them with satellite-based products. Evaluation results show that most DGVMs simulate present-day global fire emission totals within the range of satellite-based products, and can capture the high emissions over the tropical savannas, low emissions over the arid and sparsely vegetated regions, and the main features of seasonality. However, most of the models fail to simulate the interannual variability, partly due to a lack of modeling peat fires and tropical deforestation fires. Historically, all models show only a weak trend in global fire emissions before ~1850s, consistent with multi-source merged historical reconstructions. The long-term trends among DGVMs are quite different for the 20th century, with some models showing an increase and others a decrease in fire emissions, mainly as a result of the discrepancy in their simulated responses to human population density change and land-use and land-cover change (LULCC). Our study provides a basic dataset for developing regional and global multi-source merged historical reconstructions and merging methods, and analyzing historical changes of fire emissions and their uncertainties as well as their role in the Earth system. It also highlights the importance of accurately modeling the responses of fire emissions to LULCC and population density change in reducing uncertainties in historical reconstructions of fire emissions and providing more reliable future projections.

Li, Fang↗

Improved Hypoxia Modeling for Nutrient Control Decisions in the Gulf of Mexico

As required by the Harmful Algal Bloom and Hypoxia Research Control Act of 1998, the Mississippi River/Gulf of Mexico Watershed Nutrient Task Force issued the 2001 Gulf Hypoxia Action Plan (updated in 2008). In response to the Gulf Hypoxia Action Plan of 2001 (updated in 2008), the EPA Gulf of Mexico Hypoxia Modeling and Monitoring Project has established a detailed model for the Mississippi-Attchafalaya River Basin which provides a capability to forecast the multi-source nutrient loading to the Gulf and the subsequent bio-geochemical processes leading to hypoxic conditions and subsequent effects on Gulf habitats and fisheries. The primary purpose of the EPA model is to characterize the impacts of nutrient management actions, or proposed actions on the spatial and temporal characteristics of the Gulf hypoxic zone. The model is expected to play a significant role in determining best practices and improved strategies for incentivizing nutrient reduction strategies, including installation of on-farm structures to reduce sediment and nutrient runoff, use of cover crops and other agricultural practices, restoration of wetlands and riparian buffers, improved waste water treatment and decreased industrial nitrogen emissions. These decisions are currently made in a fragmented way by federal, state, and local agencies, using a variety of small scale models and limited data. During the past three years, EPA has collected an enormous amount of in-situ data to be used in the model. We believe that the use of NASA satellite data products in the model and for long term validation of the model has the potential to significantly increase the accuracy and therefore the utility of the model for the decision making described above. This proposal addresses the Gulf of Mexico Alliance (GOMA) priority issue of reductions in nutrient inputs to coastal ecosystem. It further directly relates to water quality for healthy beaches and shellfish beds and wetland and coastal conservation restoration.

Habib, Shaid↗

Desert Locust Cropland Damage Differentiated from Drought, with Multi-Source Remote Sensing in Ethiopia

In 2020, Ethiopia had the worst desert locust outbreak in 25 years, leading to food insecurity. Locust research has typically focused on predicting the paths and breeding grounds based on ground surveys and remote sensing of outbreak factors. In this study, we hypothesized that it is possible to detect desert locust cropland damage through the analysis of fine-scale (5–10 m) resolution satellite remote sensing datasets. We performed our analysis on 121 swarm point locations on croplands derived from the Food and Agriculture Organization (FAO) of the United Nations, and 94 ‘non-affected’ random cropland sample points generated for this study that are distributed within 20–25 km from the ‘center’ of swarm affected sample locations. Integrated Drought Condition Indices (IDCIs) and Vegetation Health Indices (VHIs) calculated for the affected sample locations for 2000–2020 were strongly correlated (R(exp 2) > 0.90) with that of the corresponding non-affected group of sample sites. Drought indices were strongly correlated with the evaluation Standardized Precipitation Evapotranspiration Indices (SPEIs) and showed that 2020 was the wettest year since 2000. In 2020, the NDVI and backscatter coefficient of cropland phenologies from the affected versus non-affected cropland sample sites showed a slightly wider, but significant gap in March (short growing season) and August-October (long growing season). Thus, slightly wider gaps in cropland phenologies between the affected and non-affected sites were likely induced from the locust damage, not drought, with fine scale data representing a larger gap.

Desert locust↗

Multi-source Estimates of Land / Ocean Moisture Transport Variability over the Satellite Era

It is widely appreciated that atmospheric transport of water from the world’s oceans is a process key to planetary energy balance as well as Earth's habitability. What is not yet clear is the extent of variability in moisture transports, the relative importance of interdecadal variability versus climate change signals, and importantly, our ability to quantify these changes. This work assesses variations in moisture transport variability during the satellite era (~1980 to present) by comparing several different estimates. (i) The most direct estimate is the vertically integrated flux convergence of moisture from reanalyses which use observed wind and moisture information. (ii) One alternative estimate comes from P-ET over land taken from global hydrologic models constrained with precipitation and near-surface meteorology. Here we use an ensemble of six models. An adjunct to this method is to employ satellite derive ET (e.g., GLEAM or DOLCE). (iii) Complementary to this is E-P over the global oceans derived from satellite estimates of P such as TRMM, GPM and GPCP and SeaFlux V3 or J-OFURO3 estimates of E, all relying heavily upon microwave measurements. Transport between land and oceans must essentially balance at monthly scales, i.e., vanish globally. (iv) a fourth perspective comes from estimate of terrestrial RO + storage rate, delta S. G-RUN Ensemble which uses observed streamflow and P measurements to calibrate a statistical model provides the former while GRACE, GRACE-FO provide total water storage anomalies used to calculate storage rate changes. GRACE REC uses GRACE data to train a precipitation-driven statistical model to extend storage estimates before the GRACE era. (All of these alternatives to reanalysis estimates also consider the small atmospheric column water vapor contribution.) We examine the transport changes from these three different methodologies, their relative accuracies and discuss the origin of their differences. Regional trends in moisture flux divergence and their role in multi-decadal trends are considered. Interannual variability arising in connection with ENSO variability is a dominant signal, driven largely by P changes. Trends since 1980 include reductions in moisture delivery to the western U.S., eastern Brazil, and central Africa with recovery of moisture convergence to the Sahel and parts of eastern North America.

Franklin Robertson↗

Multi-Sourced Collaboration for the Production and Refining of Rare Elements and Critical Metals (Final Technical Report)

The project objective was to develop a feasible and cost-effective method for recovering rare earth elements (REEs) and critical materials (CMs) from coal and coal byproducts, resulting in high-purity individually separated REEs and CMs. The targeted REEs included Y, Pr, Nd, Gd, Dy, and Sm, with a purity of over 99.5%, while the CMs included Co, Mn, Ga, Sr, Li, Ni, Zn, and Ge, with a purity of over 90%. The project aimed to design a prototype facility capable of producing 1-3 tonnes/day of high-purity REO mixes. The work was divided into four designated circuits: 1) REE extraction and concentration, 2) REE separation and purification, 3) RE metal production, and 4) CM production. To achieve these goals, the project involved 11 tasks, including technology reviews, research, process flow diagram development, mass balance estimation, and preliminary technical-economic analysis. The project team included researchers from the University of Kentucky, University of Alabama and Virginia Tech as well as process specialists from Argonne National Laboratory. MP Materials provided technical support regarding rare earth markets and processing while Alliance Coal performed resource assessment. The project included a market analysis for Nd/Pr, Tb, Dy, Gd, Y, Co, Mn, Li, Sr, Ga, Ni, Zn, and Ge. These analyses provided insights into the supply and demand trends as well as historic and future projections of market price relative to purity requirements for these elements. Two coal resources were selected for the project: the West Kentucky No. 13 (Baker) Seam and an undisclosed lignite resource in the Illinois coal basin. The estimated quantities of REEs in these resources were calculated based on production samples and drilling data. It was estimated that there is adequate supply for an operation producing one metric ton daily of higher purity mixed rare earth oxides (MREO) for approximately 20 years at a site located in western Kentucky. In Circuit 1, project data was obtained from a pilot heap leach and REE concentration facility. It was concluded that the existing circuit, which generated a MREO concentrate, two types of CM mixed products, and Li- and Sr-containing waters, would be suitable feed for circuits 2-4. Data from the first-of-its-kind coal coarse refuse heap leach pilot pad played a crucial role in estimating reliable elemental concentrations of the pregnant leaching solution (PLS). The average total REE concentration in the PLS was found to be 28.6 ppm. In Circuit 2, several concepts were explored including a novel process referred to as solvent-assisted chromatography (SAC). This concept involved a novel columnar reactor that incorporated multiple mixer/settlers, thereby enabling the operation of counter-flowing aqueous and organic phases. Unfortunately, due to project time constraints, a complete fundamental modeling analysis could not be completed to fully evaluate the technology. Molten salt electrowinning was considered as an alternative for circuit 3 following circuit 2 purification circuit utilizing the novel SAC process. A mass and energy balance of Nd reduction to metal in a fluoride containing molten salt electrolyte was conducted. Comparisons were made with the current state of Asian molten salt electrorefining, and potential improvements in siphoning rare earth metals (REM) from the reactor were presented. A cost estimate was performed for the production of 1 tonne per day, which yielded a total of $2.29 million for the nine electrowinning (EW) cells required. The selected option for circuits 2 and 3 was a plasma distillation process, which initially separates rare earth elements (REEs) from other elements. This is followed by selective electrowinning in various ionic liquids. The selection was made on the basis of thermodynamic modeling and experimental data previously published by a project partner. The combination offers an innovative approach to integrated refining and RE metal production. For Circuit 4, an extensive literature review was conducted for the processing of the CMs. The ultimate decision was to utilize a combined plasma and ionic liquid process as well to produce individual high-purity concentrates of Zn, Ni, Co, Mn, and Mg. A separate flowsheet for Li and Sr was recommended, which would yield carbonates of these elements. Due to the lack of suitable experimental data at this time, a process recommendation could not be provided but several methods have been proposed for consideration. Lastly, a techno-economic analysis (TEA) was conducted to assess the effectiveness of the proposed process for further investigation. The TEA results revealed a capital expense (CapEx) of $737 million and an annual operational expense (OpEx) of $220 million. Due to the selected elements, the hypothetical heap leach pad can produce 1 metric tonne per day of REO equivalent, but a conscious decision was made to only treat targeted REEs, resulting in the production of 0.4 metric tonne of REM. An estimated annual revenue of $90.87 million was projected based on standard market pricing information provided by the funding agency. During the TEA, ten different modules were evaluated for costing purposes. The precipitation circuit was identified as the largest single operational expense, followed by the Mg/Mn process due to the amount of treated metal. In terms of capital expenditures, the heap leach process incurred the highest cost, followed by the Mg/Mn process. The scalability of the plasma process is a crucial consideration since the reactors cannot be scaled beyond the largest demonstrated size due to their reliance on surface area of the slag and vapor phase. The purity estimate for the REEs are generally 98%±2% to produce a metal. The purity level being lower than the project objective was due to the lack of specific experimental data needed to tighten the tolerance of the estimates. Based on literature and previous experience, the CMs are estimated as follows; Ga (95%+, metal), Sr (95%+, carbonate), Li (95%+, carbonate), Ni (98%±2%, metal), Zn (95%+, metal sponge), Ge (95%+, metal), Co (98%±2%, metal), and Mn (98%±2%, metal).

01 COAL, LIGNITE, AND PEAT↗

Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonization

Objective: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. Methods: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. Results: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. Conclusions: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes.

Zhou, Doudou↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Evaluating the feasibility of using downwind methods to quantify point source oil and gas emissions using continuously monitoring fence-line sensors

The dependable reporting of methane (CH 4 ) emissions from point sources, such as fugitive leaks from oil and gas infrastructure, is important for profit maximization (retaining more hydrocarbons), evaluating climate impacts, assessing CH 4 fees for regulatory programs, and validating CH 4 intensity in differentiated gas programs. Currently, there are disagreements between emissions reported by different quantification techniques for the same sources. It has been suggested that downwind CH 4 quantification methods using CH 4 measurements on the fence line of production facilities could be used to generate emission estimates from oil and gas operations at the site level, but it is currently unclear how accurate the quantified emissions are. To investigate the accuracy of downwind methods, this study uses fence-line simulated data collected during controlled-release experiments as input for a non-standard closed-path eddy covariance (EC), the Gaussian plume inverse model (GPIM), and the backward Lagrangian stochastic (bLs) model in a range of atmospheric conditions. This study's EC attempt was unsuccessful due to data collection and instrumentation issues, resulting in invalid results characterized by underestimated emissions, large negative fluxes, and cospectra/ogives that deviated from their ideal shapes. Consequently, the EC results could not be compared with the GPIM and bLS model. The bLs model demonstrated the highest accuracy for single-release single-point emissions, though it exhibited greater uncertainty than GPIM under multi-release conditions. Across the GPIM and bLs model, the most reliable quantification was achieved with 15 min averaging and a narrow 5° wind sector range. Although EC was limited in this context, future studies should consider employing a standard EC system and further optimizing GPIM and bLs approaches – particularly for complex multi-source scenarios – to enhance quantification accuracy and reduce uncertainty.

03 NATURAL GAS↗