Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “community data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

xCDAT: A Python Package for Simple and Robust Analysis of Climate Data

xCDAT (Xarray Climate Data Analysis Tools) is an open-source Python package that extends Xarray (Hoyer & Hamman, 2017) for climate data analysis on structured grids. xCDAT streamlines analysis of climate data by exposing common climate analysis operations through a set of straightforward APIs. Some of xCDAT’s key features include spatial averaging, temporal averaging, and regridding. These features are inspired by the Community Data Analysis Tools (CDAT) library (Dean N. Williams et al., 2009) (D. N. Williams, 2014) (Doutriaux et al., 2019) and leverage powerful packages in the Xarray ecosystem including xESMF (Zhuang et al., 2023), xgcm (Abernathey et al., 2022), and CF xarray (Cherian et al., 2023). To ensure general compatibility across various climate models, xCDAT operates on datasets that are compliant with the Climate and Forecast (CF) metadata conventions (Hassell et al., 2017).

54 ENVIRONMENTAL SCIENCES↗

Status of the International Criticality Safety Benchmark Evaluation Project

The International Criticality Safety Benchmark Evaluation Project (ICSBEP) has continued its work generating evaluations of new and historical benchmark experiments since the last update to the nuclear criticality safety (NCS) community at the 12th International Conference on Nuclear Criticality Conference held in 2023. One additional version of the ICSBEP Handbook has been published since that update, and the Technical Review Group (TRG) held two in-person meetings to review and approve additional benchmarks. The 2022 and 2023 editions of the handbook were combined into one release (published in November 2024) and contained 13 new evaluations with 46 different configurations and two major revisions to existing evaluations. The 2024 version of the handbook, currently under publication review, will contain two new evaluations with 15 new configurations and one major revision to HEU-MET-FAST-028, the evaluation of Flattop with a uranium core. The ICSBEP TRG met again in person in April 2025 to review benchmarks for the 2025 ICSBEP Handbook and final comment resolution is currently ongoing. Many of the new benchmarks represent contemporaneous experiments that have been specifically optimized to provide validation cases relevant to the NCS community. One major area of focus for new critical experiments is to target the sparsely populated intermediate energy (or resonance) region. Another focus of many of the new benchmarks is to provide experiments sensitive to different materials, such as chlorine, hafnium, tantalum, titanium, molybdenum, chromium, and polymethyl methacrylate (PMMA, or Lucite). The ICSBEP continues to deliver high-quality, peer reviewed evaluations of integral experiments relevant to the nuclear data community.

HEU-MET-FAST-028↗

A baseline structure inventory with critical attribution for the US and its territories

Leveraging high performance computing, remote sensing, geographic data science, machine learning, and computer vision, Oak Ridge National Laboratory has partnered with Federal Emergency Management Agency (FEMA) to build a baseline structure inventory covering the US and its territories to support disaster preparedness, response, and recovery. The dataset contains more than 125 million structures with critical attribution, and is ready to be used by federal agencies, local government and first responders to accelerate on-the-ground response to disasters, further identify vulnerable areas, and develop strategies to enhance the resilience of critical structures and communities. Data can be freely and openly accessed through Figshare data repository, ESRI’s Living Atlas or FEMA’s Geodata platform.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Evaluation of Physical Microphysical Property Retrieval Algorithms During the 2020 IMPACTS Field Campaign

The NASA Investigation of Microphysics and Precipitation for Atlantic Coast Threatening Snowstorms (IMPACTS) field campaign provides high-quality, high-altitude aircraft lidar (532 nm), radar (W-band) and in-cloud microphysical aircraft data taken during wintertime storm events impacting the United States. This study evaluates two mass-dimensional relationships (Brown and Francis (1995, BF95); Heymsfield (2014, H14) and two lidar-radar microphysical retrieval algorithms (Cloudsat and CALIPSO Ice Cloud Property Product (2C-ICE); VarPy (a variational method derived from the satellite lidar-radar data community)) to estimate aircraft-retrieved volume extinction coefficient (σ), ice water content (IWC), and effective radius (r e ) during the 2020 IMPACTS deployment. BF95 and H14 have a close 1:1 correlation (R 2 = 0.98) with in-situ observations of σ. However, only BF95 displays a linear, consistent, and almost temperature-independent low bias for IWC and r e , which likely arises from the environmental conditions used to determine each. Unlike the field-campaign-derived BF95 and H14 relationships, VarPy and 2C-ICE directly ingest the aircraft-based lidar and radar data to simulate σ, IWC, and r e . For all three microphysical parameters, VarPy and 2C-ICE retrieval errors became notably more pronounced around the dendritic growth zone (-15°C to -10°C) and near freezing (≥-5°C), which suggests that both algorithms experience difficulty addressing riming and aggregation processes and with larger particles (dendrites and plates) due in part to their simplified ice particle assumptions. However, the mean-melt diameter ice-particle assumption did yield more accurate IWC estimates, which led to slightly better overall results for VarPy.

54 ENVIRONMENTAL SCIENCES↗

Rhizosphere Microbiome Diversity Potentially Supports Robust Nature of Field Pennycress ( Thlaspi arvense L.) in Dryland Cropping Systems of Eastern Washington

ABSTRACT Field pennycress ( Thlaspi arvense L.) is an annual in the Brassicaceae family and is currently being developed as an oilseed intermediate crop suitable for renewable biodiesel and jet fuel. It displays many desirable characteristics for this role including cold tolerance, a rapid life cycle, and a seed fatty acid profile conducive to bioenergy generation. These traits make field pennycress favorable for winter oilseed cultivation in the inland Pacific Northwest (iPNW). Simultaneously, intermediate crops are an increasingly recognized component of both agronomic sustainability and soil health management. Intermediate crops enhance soil microbial diversity, which benefits both soil and plant health. To understand the impact of field pennycress on soil microbial diversity, two natural accessions and seven experimental accessions were grown at three sites in Eastern Washington. Aboveground biomass and rhizosphere soil were then collected. Soil genomic DNA was extracted from rhizosphere samples and used to generate an amplicon library for bacterial (16S) and fungal (ITS) rRNA sequences. The resulting libraries were analyzed in QIIME2, which revealed that not only did the fad2 deficient line from the Spring32‐10 background have significantly increased aboveground biomass production compared to other pennycress genotypes, but also displayed significantly higher β‐diversity in the rhizosphere community specifically at the site experiencing the driest conditions. ANCOM analysis showed that multiple sequences similar to beneficial plant and soil health enhancing organisms such as Trichoderma spirale , Pseudomonas spp., and Methylobacterium goesingense were found to be enriched in the microbiome of the fad2 Spring32‐10 background also at that site. To add additional context to rhizosphere community data, root exudates from two pennycress genotypes were captured in magenta boxes and analyzed using HPLC. Future work will expand our understanding of the mechanisms by which field pennycress creates diversity in the rhizosphere, thus expanding our ability to cultivate this crop in the iPNW.

54 ENVIRONMENTAL SCIENCES↗

CROCUS Urban Canyons - Space Science and Engineering Center (SPARC) Doppler lidar data

This is the netCDF format output from the Halo Photonics Streamline XR Doppler lidar that was deployed next to the Space Science and Engineering Center (SPARC) trailer at the University of Illnois-Chicago greenhouse parking lot during CROCUS Urban Canyons. The purpose of collecting this dataset is to provide vertical and horizontal wind profiles for studying the characteristics of turbulence over the urban canyon of Chicago. This data contains the radial velocity, intensity, and backscatter from the vertical profile, range height indicator, and sector scans that were performed over both Intensive Operating Period 1 and 2 of CROCUS Urban Canyons. There are four different types of files: * The Range Height Indicator (RHI) files contain scans that are along a constant azimuth, spanning the entire hemisphere of elevation values above the surface. * The Velocity Azimuth Display (VAD) files contain the raw radial velocity data from the 6-beam, 60 degree scans. * The User1 files contain stacked Plan Position Indicator scans over a 45 degree quadrant over downtown Chicago. * The Stare files contain vertically pointing scans. These are standard netCDF files that can be opened using xarray. The VAD scans can be processed from their raw radial velocities to horizontal wind speeds with the Atmospheric data Community Toolkit (https://arm-doe.github.io/ACT/).

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC WINDS↗

BNF C-band Scanning ARM Precipitation Radar 2nd Generation (CSAPR-2) Extracted Radar Columns and In-Situ Sensors (RadCLss)

Corrected Moments in Antenna Coordinates (CMAC) calculates quantitative precipitation estimates (QPE) from empirical relationships based on equivalent radar reflectivity factor, specific differ- ential phase, and specific attenuation. To evaluate these empirical relationships, the Extracted Radar Columns and In-Situ Sensors (RADclss) product was developed. Utilizing Py-ART, RAD- clss extracts CMAC radar columns above various ARM and partner locations. These columns are then spatiotemporally synced with in-situ observations at the surface utilizing the Atmospheric data Community Toolkit (ACT; Theisen et al. 2025), allowing direct comparison of radar parameters with rain gauges and laser disdrometers for further investigation.

bankhead↗

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES↗

VA Community Determinants of Health Data Curation Documentation FY26-Q1

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

VA Community Determinants of Health Data Curation Documentation FY26-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1 km grid) to another (e.g., U.S. Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, U.S. Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1 km grids. Some economic data may only be available at the ZIP code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., U.S. Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗

A path to intelligent watersheds: coordinating the data to decision pipeline

Operations of multi-reservoir systems are challenged in-part by the interplay of complex physical processes functioning within the watershed. The employment of intelligent systems can be of aid by linking environmental sensing, information technology, data analytics, simulation and decision support to achieve a data-to-decision flow of information. A further challenge is that watershed resources are managed for multiple purposes requiring some level of coordination among numerous resource managers, asset operators and users. System intelligence in this context relies on shared community platforms (data portals, community models), and coordinated communication between decision makers. Opportunities to enrich watershed intelligence has been the subject of a roadmapping exercise for the Department of Energy’s Water Power Technologies Office which has relied on broad stakeholder engagement. Initial phases of engagement involved personal interviews and a series of virtual group meetings, which focused on identifying opportunities to improve the intelligence of the physical infrastructure within our watersheds—examples of feedback include improved sensing of snowpack and runoff, data standards for facilitated data sharing, and better forecasting tools. The latter phase of engagement involved the conduct of a case study in the Upper Colorado River basin where key stakeholders were interviewed to map how their decisions are informed by intelligence from other basin stakeholders. Our presentation will highlight the interdisciplinary flow of information in complex watershed systems and identify physical and institutional opportunities toward the strategic operation of water infrastructure.

Colorado River↗

Community Choice Aggregation(CCA) Data Collection Webinar for Status and Trends in the Voluntary Market Report (2024 Data) [Slides]

We have subcontracted LEAN Energy US, to help us improve our CCA data collection effort for the Annual Voluntary Energy Markets Data Report. LEAN Energy US (Local Energy Aggregation Network) is a national 501(c)3 non-profit organization dedicated to accelerating the country's transition to clean and renewable power, supporting competition and customer choice in the energy sector, and maintaining affordable electricity rates. We work in partnership with a range of organizations to actively support the formation and operational success of Community Choice Aggregation (CCA) programs around the country. This webinar, hosted in partnership with LEAN Energy US, is intended to introduce their members to our data collection effort and encourage CCAs in their network to participate.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Validation of Modern Nuclear Data Processing in SCALE

The nuclear data (ND) community is continuously developing more accurate, diversified, and comprehensive data for radiation transport modeling to support the nuclear science community. As these community efforts progress, it is crucial that ND processing tools like AMPX (used for SCALE [1] ND) also be developed in parallel to incorporate these new data into transport codes and actually deliver those data to end users. AMPX is a mature, well-tested code that was developed by many people at Oak Ridge National Laboratory (ORNL) over the course of the past few decades. A large portion of the AMPX codebase, however, was outdated, difficult to maintain, and incompatible with modern code development tools. Some of the most important parts of the AMPX code have now been replaced with modern C++ code that can be maintained more cost-effectively and can be tested more rigorously.

AMPX↗

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences↗

VA Community Determinants of Health Data Curation Documentation FY25-Q4

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗