Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DATA CONVERSION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Unpaired image translation to mitigate domain shift in liquid argon time projection chamber detector responses

Deep learning algorithms often are developed and trained on a training dataset and deployed on test datasets. Any systematic difference between the training and a test dataset may severely degrade the final algorithm performance on the test dataset—what is known as the domain shift problem . This issue is prevalent in many scientific domains where algorithms are trained on simulated data but applied to real-world datasets. Typically, the domain shift problem is solved through various domain adaptation (DA) methods. However, these methods are often tailored for a specific downstream task, such as classification or semantic segmentation, and may not easily generalize to different tasks. This work explores the feasibility of using an alternative way to solve the domain shift problem that is not specific to any downstream algorithm. The proposed approach relies on modern Unpaired Image-to-Image (UI2I) translation techniques, designed to find translations between different image domains in a fully unsupervised fashion. In this study, the approach is applied to a domain shift problem commonly encountered in Liquid Argon Time Projection Chamber (LArTPC) detector research when seeking a way to translate samples between two differently distributed LArTPC detector datasets deterministically. This translation allows for mapping real-world data into the simulated data domain where the downstream algorithms can be run with much less domain-shift-related performance degradation. Conversely, using the translation from the simulated data to a real-world domain can increase the realism of the simulated dataset and reduce the magnitude of any systematic uncertainties. To evaluate the quality of the translations, we use both pixel-wise metrics and a downstream task to measure the effectiveness of UI2I methods for mitigating the domain shift problem. We adapted several popular UI2I translation algorithms to work on scientific data and demonstrated the viability of these techniques for solving the domain shift problem with LArTPC detector data. To facilitate further development of DA techniques for scientific datasets, the ‘Simple Liquid-Argon Track Samples’ dataset used in this study is also published.

97 MATHEMATICS AND COMPUTING↗

Evaluation of neutron dosimetry capabilities with the MC-15 portable multiplicity counter

This work proposes a preliminary neutron dose rate estimation method for a neutron multiplicity detector through measurement- and simulation-based analyses. Uncharacterized neutron-emitting sources may be encountered in situations such as nuclear emergency response, safeguards, and treaty verification. These circumstances may present irradiation risk to personnel conducting field assay, search, and characterization measurements. It is therefore of interest to provide a field neutron dosimetry capability with the existing neutron multiplicity counting (NMC) capabilities. To date, no commercially-available neutron detection systems are capable of both accurate NMC and real-time neutron dosimetry. This work will focus on estimating dose rate using input from a single fielded NMC called the MC-15. The energy-dependent neutron detection efficiency response of the MC-15 was quantified in monoenergetic neutron simulations and evaluated in response to two neutron-emitting sources and to a polyethylene-moderated source. The results were compared to existing neutron dosimeters and established the proof of concept for further investigation of the MC-15 for dose estimation. Measurement results were also replicated in simulations; additional simulations were then conducted to expand upon the limited empirical data. The initial empirical results provided a conversion factor appropriate for use when measuring 252 Cf neutrons that is independent of polyethylene shielding presence and thickness. The simulated data sets were then used to evaluate a fit equation allowing estimation of the neutron dose rate for less restricted geometric configurations and dependent only on the distance between the source and the detector. Additionally, the energy dependence of the efficiency response indicates that further empirical evaluations could provide energy-dependent conversion factors for broader neutron dosimetry capabilities with a wider range of neutron-emitting sources.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Zero-Power Analog Optical Processing

The motivation behind this research is the growing challenge of handling the massive amounts of data generated by modern imaging systems. Conventional digital image processing techniques are struggling to keep pace with the demands of high-resolution and high-speed imaging systems for remote sensing due to their high-power consumption and data storage requirements. We present a novel approach based on analog photonics to address this challenge. The proposed system utilizes a silicon-photonics-based image encoder positioned after image formation and initial optical-to-electrical conversion. The photonic encoder compresses image data using a passive disordered photonic structure to perform kernel-type random projections of the raw data. The compressed data is then processed by a back-end neural network, which reconstructs the original image with high fidelity (structural similarity exceeding 90%). Our proposed approach has the potential to compress images with ~ 1000X lower power consumption compared to digital approaches with data rates exceeding 1 terapixel/second.

97 MATHEMATICS AND COMPUTING↗

Tracking cropland transitions: A comparative analysis of U.S. land cover change data

There are a growing number of land cover data available for the conterminous United States, supporting various applications ranging from biofuel regulatory decisions to habitat conservation assessments. These datasets vary in their source information, frequency of data collection and reporting, land class definitions, categorical detail, and spatial scale and time intervals of representation. These differences limit direct comparison, contribute to disagreements among studies, confuse stakeholders, and hamper our ability to confidently report key land cover trends in the U.S. Here we assess changes in cropland derived from the Land Change Monitoring, Assessment, and Projection (LCMAP) dataset from the U.S. Geological Survey and compare them with analyses of three established land cover datasets across the coterminous U.S. from 2008-2017: (1) the National Resources Inventory (NRI), (2) a dataset Lark et al. 2020 derived from the Cropland Data Layer (CDL), and (3) a dataset from Potapov et al. 2022. LCMAP reports more stable cropland and less stable noncropland in all comparisons, likely due to its more expansive definition of cropland which includes managed grasslands (pasture and hay). Despite these differences, net cropland expansion from all four datasets was comparable (5.18-6.33 million acres), although the geographic extent and type of conversion differed. LCMAP projected the largest cropland expansion in the southern Great Plains, whereas other datasets projected the largest expansion in the northwestern and central Midwest. Most of the pixel-level disagreements (86%) between LCMAP and Lark et al. 2020 were due to definitional differences among datasets, whereas the remainder (14%) were from a variety of causes. Cropland expansion in the LCMAP likely reflects conversions of more natural areas, whereas cropland expansion in other data sources also captures conversion of managed pasture to cropland. The particular research question considered (e.g., habitat versus soil carbon) should influence which data source is more appropriate.

60 APPLIED LIFE SCIENCES↗

Electromagnetic Transient Modeling of Large Data Centers for Grid-Level Studies

The magnitude and complexity of electricity usage patterns from large data centers are having significant impacts on the operation and dynamics of the power grid; grid operators and planners require a range of specialized data center models to properly evaluate these impacts and specify technical solutions as needed. Towards addressing this need, Pacific Northwest National Laboratory (PNNL) has developed a library of electromagnetic transient (EMT) models for grid-level studies of data centers called the data center model library (DML). This report describes how the DML was created and how it may properly be used. The models present in the DML are generic models; subject matter expertise and additional technical data are needed to modify these models before they can represent any real data center. However, they will significantly reduce the level of effort required to develop site-specific models and can serve as a common starting point to guide industry towards a more refined consensus. Most of the models within DML are dedicated to representing the power electronics interfaces commonly used in modern data centers, such as double-conversion uninterruptible power supplies and single-phase power factor correction converters. These models are intended for use in grid-level studies and are a simplified aggregation of many small components. That said, background material on the physical and electrical design of large data centers is provided as companion material so that users can be aware of many of the details which have been omitted or streamlined as a matter of practical necessity. Additionally, guidance on the application of EMT analysis for data center interconnection studies is provided, which aids users in identifying when the DML is necessary and what sort of additional model development may be necessary for conducting real-world studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Electromagnetic Transient Modeling of Large Data Centers for Grid-Level Studies: Beta Release

The magnitude and complexity of electricity usage patterns from large data centers are having significant impacts on the operation and dynamics of the power grid; grid operators and planners require a range of specialized data center models to properly evaluate these impacts and specify technical solutions as needed. Towards addressing this need, Pacific Northwest National Laboratory (PNNL) has developed a library of electromagnetic transient (EMT) models for grid-level studies of data centers called the data center model library (DML). This report describes how the DML was created and how it may properly be used. This report details the DML’s beta release, completed in July 2026. This is a revision and expansion of the alpha release, which was made available in January 2026 The models present in the DML are generic models; subject matter expertise and additional technical data are needed to modify these models before they can represent any real data center. However, they will significantly reduce the level of effort required to develop site-specific models and can serve as a common starting point to guide industry towards a more refined consensus. Most of the models within DML are dedicated to representing the power electronics interfaces commonly used in modern data centers, such as double-conversion uninterruptible power supplies and single-phase power factor correction converters. These models are intended for use in grid-level studies and are a simplified aggregation of many small components. That said, background material on the physical and electrical design of large data centers is provided as companion material so that users can be aware of many of the details which have been omitted or streamlined as a matter of practical necessity. Additionally, guidance on the application of EMT analysis for data center interconnection studies is provided, which aids users in identifying when the DML is necessary and what sort of additional model development may be necessary for conducting real-world studies.

electromagnetic transients↗

Sulfate Conversion of Reillex HPQ Anion Exchange Resin for Disposal (Interim Report)

This report describes preliminary data to validate the Savannah River Plutonium Processing Facility’s (SRPPF) flowsheet for conversion of used Reillex HPQ anion exchange resin from the nitrate form to the sulfate form. The nitrate form is an oxidizer and therefore does not meet acceptance criteria for disposal at the Waste Isolation Pilot Plant (WIPP). The purpose of this study is to develop data to support acceptance for this disposition pathway. Due to the challenges characterizing the nitrate concentration on solid resin, the data developed to date are based upon indirect analysis of the ion exchange column effluent by ion chromatography. These challenges are discussed and two methods for quantification of nitrate directly on the resin are recommended for further development: TGA-MS and permanganate digestion followed by IC. The resin used for this work was provided in the chloride form; this is the form in which resin is supplied by the manufacturer. However, it had to be converted to the nitrate form, which is the form that will be used in SRPPF’s ion exchange process, prior to use in the sulfate conversion experiments. The chloride-form resin was characterized. A lab-scale procedure for the conversion of Reillex HPQ resin from the chloride to nitrate form was validated. The nitrate-form resin was assessed for particle size and chloride concentration to ensure it met SRPPF’s facility specifications. The baseline sulfate conversion flowsheet was tested. However, nitrate was still detectable in the effluent after approximately 10 bed volumes of 1 M sodium sulfate had been passed through the resin bed. Additional experiments were performed to assess the effect of increasing the feed volume, reducing the flowrate, the use of 2 M sulfuric acid instead of sodium sulfate, and the use of irradiated resin. The sulfuric acid test was the only one which provided a nondetectable nitrate concentration (<0.002 M) in the effluent. Detectable nitrate in the column effluent suggests that nitrate is still present on the resin itself.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Assessment of Thin Plastic Scintillation Detectors for Beta-Particle Measurements at the Advanced Test Reactor Critical Facility

The Fission Wire Measurement System is a custom measurement system designed in the 1960s to measure the beta-particle activity of irradiated uranium-aluminum fission wires. This measurement is conducted to determine the fission rate profile of the Advanced Reactor Test Critical facility. The Advanced Test Reactor Critical facility is an open-pool, low-power test reactor used to qualify experiment configurations and verify core models prior to full-power experiment irradiations in the Advanced Test Reactor. Power distribution measurements in ATR-C use uranium-aluminum wires that are distributed throughout the core to validate simulation and modeling results. These measurements require from 340 to 1500 wires to be irradiated and measured within a 12-hour window. The system consists of 4 measurement channels and one reference channel, each with a 2-pi proportional gas flow detector and the measurement channels each have an automated sample changer. The gas flow detectors are of a custom design for this detector system that use methane gas with a large anode wire compared to modern proportional counters. These detectors, which are nearly 60 years old are irreplaceable. The measurements from these gas detectors are affected by the gas flow rate, atmospheric and line pressure, and are very sensitive to the applied high voltage. Recent improvements have been made to the control and data acquisition system, but the detectors have remained the same. The nature of the measurement of the fission product decay activity is such that the energy spectrum of the signal is changing with time. Thin, 250-um thick, plastic scintillators were commercially obtained as a potential replacement for the gas flow detectors. The original calibration of the uranium-aluminum fission wires was conducted in 1965 using a series of irradiations of gold foils and the wires in a well-characterized thermal neutron field. These measurements provided a time-dependent fission rate conversion factor from the gold foil data to calibrate the fission wires based on the response from the 2-pi proportional gas detectors. Transitioning to the new detectors requires qualification and testing. The sensitivity of the scintillators to changes in the energy spectrum of the fission wires and translation of the calibration factor have been completed. These measurements indicated that the sensitivity of the scintillators over time changes at a different rate than the sensitivity of the gas flow detectors. However, the inverse activity of measurements of both detector types is linear with time. Initial results indicate that the scintillator detectors will be a sufficient replacement for the gas detectors with minor adjustments to the fission rate conversion factor. Replacement of the detectors will improve the fission wire measurements and provide a more stable and reliable measurement system.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Physicochemical and Molecular Insights into the Boundary Layer and Free Troposphere Aerosol Interactions over the Southern Great Plains

Ambient aerosols’ vertical profiles are critical for evaluating the role of aerosols in atmospheric chemistry and radiative transfer, but limited data on these profiles hinders our ability to fully assess their impact on the Earth's radiative balance. Here, in this study, we investigated the size-, time-, and altitude resolved composition of individual particles and bulk molecular composition of particle samples collected by an uncrewed aerial system–ArcticShark over the Southern Great Plains. Single particle microanalysis shows that, the free tropospheric (FT) samples are dominated (56-66%) by carbonaceous sulfate particles, while boundary layer (BL) samples are dominated (57-74%) by carbonaceous particles. Back trajectory simulations suggest that FT particles are likely influenced by long-range transport and have undergone aqueous-phase processing. Conversely, in-situ size distribution data shows evidence of particle growth in the upper BL and just below the FT. This observation may indicate vertical transport of particles from an elevated aerosol layer in the FT, possibly linked to a new particle formation event. This observation is further supported by high resolution molecular composition data, which reveals particle volatility increasing with increasing size, which aligns with the growth event. This study aids in fundamental understanding of the compositional and molecular specificity of vertically resolved organic aerosols to provide insights into particle size evolution for future atmospheric models.

ArcticShark↗

Low-dimensional carbon materials decorated FAPbI 3 for carbon-based perovskite solar cells

Carbon nanomaterials are at the forefront of research in perovskite solar cells (PSCs) due to their exceptional electrical, optical, and stability properties. Their diverse applications include serving as interfacial layers, additives, hole and electron transport materials, and back electrodes. While the influence of various low-dimensional carbon nanomaterial structures on crystallinity, optical and electrical performance, and overall device efficiency has been a topic of interest, it has not been thoroughly explored until now. In this study, we effectively integrated carbon quantum dots (CQDs), multi-walled carbon nanotubes (MWCNTs), and graphene into the FAPbI 3 photoactive layer using a two-step sequential deposition method. Our experiments revealed marked improvements in the photovoltaic performance of PSCs that incorporated all three types of carbon nanomaterials. In particular, the data shows significant enhancements in power conversion efficiency, demonstrating the effectiveness of these materials in optimizing device functionality. Notably, MWCNTs distinguished themselves by exhibiting a remarkable potential for enhancing long-term stability. This finding underscores the importance of selecting the right carbon nanomaterials for future PSC developments, paving the way for more reliable and efficient solar energy solutions. As a result, our research highlights the critical role of carbon nanomaterials in advancing perovskite solar technology.

14 SOLAR ENERGY↗

WO 3 /CuWO 4 Ratio Controls Open-Circuit Photovoltage and Photocurrent in Type II Heterojunction Solar Fuel Photoelectrodes

WO 3 /CuWO 4 photoelectrodes for the oxygen evolution reaction benefit from a type II heterojunction for charge separation. However, the impact of the WO 3 /CuWO 4 ratio on the photocurrent and the photovoltage is not clear. To probe the effect of composition, Cu x W 1-x O y thin films with variable W:Cu ratio were prepared on FTO by reactive magnetron co-sputtering of W and Cu, followed by air annealing at 500ºC. EDS, XRD, Rietveld refinement, and Raman spectroscopy confirm the presence of crystalline WO 3 and CuWO 4 in the W rich films and increasing amounts of amorphous copper oxides in the Cu rich films. Bandgaps were determined by optical absorption spectroscopy, surface photovoltage spectroscopy (SPS), and photoaction spectra and are found to decrease from 2.7 eV to 1.2 eV with increasing copper oxide content. SPS reveals n-type semiconductor photoanode behavior for WO 3 /CuWO 4 samples and p-type photocathode behavior for CuO x rich films. Photoelectrochemical experiments confirm stable water oxidation with Faraday efficiency near unity for all W rich films and photocurrents that are increasing with CuWO 4 content. Optimal performance is seen for WO 3 /CuWO 4 mixed phases containing 47-75 mass% CuWO 4 . These compositions maximize charge separation at the type II heterojunction interface between the two materials. Additionally, according to incident photon to current efficiency (IPCE) data, the WO 3 improves photon conversion below 350 nm, while CuWO 4 improves conversion at 450-525 nm. Overall, this work shows for the first time how the WO 3 /CuWO 4 ratio controls the photovoltage and the photocurrent in type II heterojunction solar fuel photoelectrodes, and how copper oxides in the copper rich films severely degrade the performance. Furthermore, these results are useful in the context of bulk-heterojunction electrodes for the conversion of solar energy into fuels.

CuWO4↗

Measurements of ϒ states production in 𝑝 + 𝑝 collisions at $\sqrt{s}$ = 500 GeV with STAR: Cross sections, ratios, and multiplicity dependence

We report measurements of ϒ⁡(1⁢𝑆), ϒ⁡(2⁢𝑆) and ϒ⁡(3⁢𝑆) production in 𝑝 + 𝑝 collisions at $\sqrt{s}$ =500 GeV by the STAR experiment in year 2011, corresponding to an integrated luminosity ℒ int = 13 pb −1 . The results provide precise cross sections, transverse momentum (𝑝 T ) and rapidity (𝑦) spectra, as well as cross section ratios for 𝑝 T < 10 GeV/c and |𝑦| < 1. The dependence of the ϒ yield on charged particle multiplicity has also been measured, offering new insights into the mechanisms of quarkonium production. The data are compared to various theoretical models: the color evaporation model (CEM) accurately describes the ϒ⁡(1⁢𝑆) production, while the color glass condensate+nonrelativistic quantum chromodynamics (CGC+NRQCD) model overestimates the data, particularly at low 𝑝 T . Conversely, the color singlet model (CSM) underestimates the rapidity dependence. These discrepancies highlight the need for further development in understanding the production dynamics of heavy quarkonia in high-energy hadronic collisions. The trend in the multiplicity dependence is consistent with CGC/saturation and string percolation models or ϒ production happening in multiple parton interactions modeled by PYTHIA 8.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Simple and Scalable Process for Nanocellulose Production from Residues and Waste (CRADA Final Report)

This project successfully developed an innovative, cost-effective, and scalable method for converting low-cost agricultural waste feedstocks into nanocellulose—a sustainable material with wide-ranging industrial applications. It addressed the critical challenges of high production costs and limited global supply, which have hindered the widespread adoption of nanocellulose in industry. Through this work, significant advancements were made in optimizing resource use, improving process efficiency, and reducing costs. Notable achievements included a 50% reduction in water usage, 30% lower chemical consumption, and 20% energy savings, all while maintaining high-quality product standards. The technical feasibility of the process was further validated at a 10L scale in collaboration with the Advanced Biofuels and Bioproducts Development Unit (ABPDU), demonstrating scalability and replicability. Key challenges in reducing nanocellulose production costs were addressed by utilizing low-cost biomass feedstocks and implementing low-temperature conversion reactions, leading to lower capital and operating expenses. The project also mitigated financial risk by generating critical data on the feasibility, adaptability, and scalability of the conversion technology, paving the way for its commercial implementation. Public benefits include advancing the circular bioeconomy, reducing environmental impact, fostering job creation, and enabling a shift toward biobased materials as sustainable alternatives to fossil-based products. These outcomes align closely with national goals to reduce greenhouse gas emissions and promote sustainable technological innovation.

09 BIOMASS FUELS↗

Precision Plant Biomass Characterization in Agriculture: Harnessing Machine Learning and Hyperspectral Imaging [Slides]

Efficient Biomass Separation Object detection of anatomical parts (Cob, Stalk, Husk) in IR images enables precise separation, improving preprocessing (e.g., drying, grinding) for biofuel production. Detailed Biomass Characterization with Hyperspectral Data Hyperspectral imaging captures spectral signatures of biomass, allowing for the identification of specific traits like moisture content, lignin levels, and nutrient composition, leading to optimized treatments for each biomass part. Enhanced Feedstock Quality By leveraging hyperspectral data, feedstock can be processed based on its chemical composition, improving conversion efficiency and biofuel yield. Automation for Large-Scale Operations Automated object detection and hyperspectral data analysis reduce manual labor, ensuring accurate sorting and faster processing, making large-scale biofuel production more efficient. Maximized Biomass Utilization Accurate identification of biomass properties minimizes waste and ensures that each part is processed according to its highest biofuel potential.

09 BIOMASS FUELS↗

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY25-Q4

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY26-Q1

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗