Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data gap”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Filling in Subsurface Storage Open Data Gaps - Updates to CCS Data Availability on EDX and EDX Spatial (FWP-1022465)

There is a need to preserve and efficiently access data resources to drive the next generation of research and development while ensuring compliance with DOE regulations. Over the last 10+ years, there has been ongoing efforts by the DOE Carbon Storage Program to ensure that there is effective data curation and preservation of DOE funded research leveraging the NETL-FECM data repository, the Energy Data eXchange (EDX). This talk presents updates about ongoing efforts to continue to support the mission of ensuring that carbon storage data is findable, accessible, interoperable, and reusable to the carbon storage stakeholder community through EDX and EDX Spatial. Presented at the NETL Carbon Management Review Meeting, Pittsburgh, 2024.

Morkner, Paige

Offshore Geologic Carbon Storage Data Collection and Data Gaps Analysis

This is a TRS documenting the Offshore Geologic Carbon Storage Data Collection. It describes the Data Collection web application and its creation as well as an accompanying Data Gaps Assessment. We present an interactive data collection and data gaps analysis to aggregate, understand, and disseminate the data that are publicly available to support offshore GCS in the United States. This data collection and data gaps analysis can be leveraged by stakeholders to understand where GCS may be viable offshore, create GCS project analogs, and address challenges to GCS in offshore environments.

58 GEOSCIENCES

A Proxy Method to Bridge LCA Data Gaps Using Automated Material Classification and Probabilistic Under-Specification

Life cycle assessments (LCAs) are essential for understanding the environmental impacts of material production. However, gaps in life cycle inventory (LCI) data for material and chemical inputs present a key challenge for LCA practitioners, especially in the early design stages. Strategies for filling in these gaps require additional time and expertise, which can hinder the LCA’s completion. This study combined automatic material classification and probabilistic under-specification to create a time-efficient method to fill material LCI data gaps. To illustrate the proposed method, proxy environmental impact distributions were generated using publicly available material LCI data classified into the ChemOnt chemical taxonomy using the open-source chemical classification software ClassyFire. Input materials with data gaps were then classified into the same taxonomy, where proxy environmental impact values could be selected from the available distributions to quickly fill in any data gaps. Although these methods were applied to classify material production processes available in the Federal LCA Commons and Ecoinvent databases, they can be applied to any LCA database. This study shows that classifying materials by their chemical structure produces taxonomies with increased granularity relative to industrial classification, improving the ability of under-specified proxy data to be used for differentiating the environmental impacts of competing designs.

biological databases

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the Geological Society of America Connects 2024 Annual Meeting in Anaheim, California, 22-25 September 2024.

Creason, Christopher

Life Cycle Inventories and Data Gap Analysis for Rare Earth Elements: Neodymium and Dysprosium from Mining to Magnets

The United States demand for Neodymium-Iron-Boron (NdFeB) magnets, produced from rare earth elements (REEs) such as (Nd) and Dysprosium (Dy), far exceeds its nascent domestic production capacity, rendering it reliant on vulnerable global supply chains dominated by China. To guide research and development investments in securing U.S. REE supply, defensible benchmark metrics across environmental, economic, and social dimensions are needed. In this study, we built globally-representative, process-based cradle-to-cradle life cycle inventories for Nd and Dy in NdFeB magnets lifecycles, encompassing primary material acquisition, beneficiation, smelting and refining, metal processing, specialty alloy and chemical transformation, subcomponent manufacturing, consumer application (use phase) and end-of-life management. We carried out detailed literature review, and applied process engineering principles to build industry-representative upscaled life cycle inventories for both metals. We used these models to conduct bottom-up literature review and gap analysis on existing literature, compilation of data sources for each life cycle stage (and transformations where necessary), and a preliminary technoeconomic analysis (TEA)/life cycle costing analysis (LCCA). Findings from this work emphasize the need for metal specific, representative REE LCIs to establish robust benchmarks for advancing sustainable REE technologies and guiding R&D in REE supply chains.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Data Centers Gap Analysis [Slides]

Data centers and other large loads are a significant driver of unprecedented, near-term demand growth in the United States. Power system planners, utilities, regulators, and other stakeholders are grappling with how to integrate data centers on the system without comprising reliability, resiliency, and energy affordability. NLR is pursuing work to develop a siting and decision-making tool that would draw on power systems modeling expertise to achieve granular representation of trade-offs involved in data center sitting and development. This slide deck supports the same workstream by reviewing the literature to identify mitigation options to facilitate near-term integration of large loads and by presenting options for pursuing data development and/or modeling projects to improve representation of siting options.

29 ENERGY PLANNING, POLICY, AND ECONOMY

An Open-Source Framework for Characterizing Urban Energy Models: Integrating Top-Down and Bottom-Up Methods to Predict Residential Buildings Characteristics: Preprint

Bottom-up urban energy models are crucial for understanding current energy use patterns and informing design strategies. However, accurately characterizing these models to represent different communities remains a challenge due to the extensive data needed for simulating existing energy use behavior. This data includes information related to human activities and building characteristics, all of which correlate with socioeconomic factors. To overcome this challenge, we developed an automated framework that utilizes both top-down and bottom-up data, to predict unknown building and occupant characteristics that are needed for more accurate and equitable modeling and analytics. Our framework, integrated into the URBANopt district energy modeling platform, uses statistical data models from ResStock. URBANopt models co-located buildings and neighborhoods. At this scale there are data gaps in building characteristic data, such as materials, insulation, occupancy, income, and energy usage of the buildings. To address this data gap, we use ResStock data, representative at the census tract scale, and develop machine-learning and deeplearning techniques to disaggregate it to individual buildings. By mapping unique occupant, building and economic properties to URBANopt energy models, we gain detailed insights into the variability of building energy use across different neighborhoods. This insight helps deploy technologies for co-located buildings and supports targeted upgrades for communities with unique economic and demographic characteristics, ensuring energy equity. Accurate characterization of energy models allows us to develop equitable strategies tailored to diverse neighborhoods, whether underserved or affluent. Our automated framework streamlines energy modeling and provides a reliable tool for building energy characterization.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION

An Integrated ML/AI Framework for Digitizing, Structuring and Searching DOE U-TRU-Fuels Data with Gap Analysis of Non-DOE Records

The U.S. Department of Energy (DOE) Advanced Fuels Campaign (AFC) is advancing transmutation fuel technologies to reduce long-lived radioactive waste by converting minor actinides into shorter-lived or stable elements through irradiation in sodium-cooled fast reactors. Key experiments such as AFC-1, AFC-2, FUels for the transmutation of Trans-URanium elements In phéniX (FUTURIX)-Fortes Teneurs en Actinides (FTA), and Experimental Breeder Reactor-II (EBR-II) X501 have provided fuel fabrication, irradiation, and performance data on various transuranic-bearing fuel forms. This report documents the creation of an artificial-intelligence assisted database, which has consolidated all DOE-owned data related to Transuranic (TRU)-bearing fuel experiments and stored across it across both the Idaho National Laboratory (INL) Nuclear Data Management and Analysis System and the INL high performance computing (HPC) infrastructure. A dedicated webpage, hosted on the INL HPC system, has been developed to support role-based access and data interaction. The database architecture allows researchers to navigate large, heterogeneous archives with far greater speed and accuracy than manual search and lays the foundation for future expansion into multimodal nuclear materials analysis environments. The database represents a major step towards a nationally integrated fuels database utilizing artificial intelligence tools.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Bridging the Gap on Data and Analysis for Distribution System Planning: Information That Utilities Can Provide Regulators, State Energy Offices and Other Stakeholders

Electric utilities conduct planning annually to ensure their distribution system meets technical standards, policies, and regulations; addresses forecasted grid conditions; satisfies customer needs; and advances utility priorities. The plan identifies grid deficiencies, analyzes potential solutions, and prioritizes capital investments and other expenditures. About 20 U.S. states and jurisdictions require regulated utilities to file some type of distribution system plan with the public utility commission for review. Requirements for sharing distribution system data and analyses vary widely, from few specific requirements to a detailed list of information that must be provided. While utilities conduct extensive analysis to develop distribution system plans, in most jurisdictions regulators and stakeholders do not know what data are available and how the utility uses the data in planning and investing. This report aims to bridge the gap by increasing understanding of the types of data and analyses utilities employ to develop distribution system plans and how the information affects their decision-making. The report describes information that states and stakeholders can ask for related to 11 data categories: -Forecasting loads and distributed energy resources (DERs) -Scenario analysis -Worst-performing circuits -Asset management strategy -Hosting capacity analysis -Value of DERs -Grid needs assessment -Cost-effectiveness framework for investments -Distribution system investment strategy and implementation -Geotargeted programs -Non-wires alternatives procurements.

24 POWER TRANSMISSION AND DISTRIBUTION

RHOD Site - NOAA PSL Wind Retrievals WINDoe / Derived Data

This dataset contains daily NetCDF files with horizontal wind profiles retrieved with the WINDoe retrieval (Gebauer and Bell 2024) at Rhode Island (RHOD). WINDoe retrievals datasets are also available at Nantucket Island (NANT, nant.windoe.z01.c1) and Block Island (BLOC, bloc.windoe.z01.c1). WINDoe is an optimal estimation algorithm to retrieve wind profiles combining multiple instruments. The code is available in this github repository (https://github.com/OAR-atmospheric-observations/WINDoe/tree/main) and the retrieval is described by Gebauer and Bell (2024). WINDoe allows combining the individual datasets and outputs into one profile taking into account the information and uncertainties of each dataset. The use of WINDoe minimizes data gaps and maximizes data availability, compared to using wind profiles from only one of the instruments. The regular height grid eases comparisons to numerical weather prediction models. Code modifications have been made that include reading in WFIP3 specific instruments, averaging Doppler lidar radial velocities at various azimuth angles to avoid overfitting, and allowing the user to define a height grid by the user in the vipfile. The instruments used as input to the retrieval are a radar wind profiler (low- and high resolution mode) providing data in and above the boundary layer, a scanning Doppler lidar usually providing data throughout the boundary layer, a profiling lidar providing data from 50 to 200 m at BLOC and NANT, and from 10 to 280 m at Rhode Island, and a surface tower (4 m at NANT and RHOD and 10 m at BLOC). From the scanning lidars, we used radial velocity measurements at 60 deg elevation angle at six different azimuth angles with a resolution of approximately 30 m along the line of sight and the lowest range gate at approximately 70 m. The wind profiles are retrieved with WINDoe up to 3.74 km with 10 m vertical resolution. The profiles are retrieved every 15 min at BLOC and NANT and every 60 min at RHOD.

17 WIND ENERGY

BLOC Site - NOAA PSL Wind Retrievals WINDoe / Derived Data

This dataset contains daily netcdf files with horizontal wind profiles retrieved with the WINDoe retrieval (Gebauer and Bell 2024) at Block Island (BLOC). WINDoe retrievals datasets are also available at Nantucket Island (NANT, nant.windoe.z01.c1) and Rhode Island (RHOD, rhod.windoe.z01.c1). WINDoe is an optimal estimation algorithm to retrieve wind profiles combining multiple instruments. The code is available in this github repository (https://github.com/OAR-atmospheric-observations/WINDoe/tree/main), and the retrieval is described by Gebauer and Bell (2024). WINDoe allows combining the individual datasets and outputs into one profile taking into account the information and uncertainties of each dataset. The use of WINDoe minimizes data gaps and maximizes data availability, compared to using wind profiles from only one of the instruments. The regular height grid eases comparisons to numerical weather prediction models. Code modifications have been made that include reading in WFIP3 specific instruments, averaging Doppler lidar radial velocities at various azimuth angles to avoid overfitting, and allowing the user to define a height grid by the user in the vipfile. The instruments used as input to the retrieval are a radar wind profiler (low- and high resolution mode) providing data in and above the boundary layer, a scanning Doppler lidar usually providing data throughout the boundary layer, a profiling lidar providing data from 50 to 200 m at BLOC and NANT, and from 10 to 280 m at Rhode Island, and a surface tower (4 m at NANT and RHOD and 10 m at BLOC). From the scanning lidars, we used radial velocity measurements at 60 deg elevation angle at six different azimuth angles with a resolution of approximately 30 m along the line of sight and the lowest range gate at approximately 70 m. The wind profiles are retrieved with WINDoe up to 3.74 km with 10 m vertical resolution. The profiles are retrieved every 15 min at BLOC and NANT and every 60 min at RHOD.

17 WIND ENERGY

NANT Site - NOAA PSL Wind Retrievals WINDoe / Derived Data

This dataset contains daily NetCDF files with horizontal wind profiles retrieved with the WINDoe retrieval (Gebauer and Bell 2024) at Nantucket Island (NANT). WINDoe retrievals datasets are also available at Block Island (BLOC, bloc.windoe.z01.c1) and Rhode Island (RHOD, rhod.windoe.z01.c1). WINDoe is an optimal estimation algorithm to retrieve wind profiles combining multiple instruments. The code is available in this github repository (https://github.com/OAR-atmospheric-observations/WINDoe/tree/main), and the retrieval is described by Gebauer and Bell (2024). WINDoe allows combining the individual datasets and outputs into one profile taking into account the information and uncertainties of each dataset. The use of WINDoe minimizes data gaps and maximizes data availability, compared to using wind profiles from only one of the instruments. The regular height grid eases comparisons to numerical weather prediction models. Code modifications have been made that include reading in WFIP3 specific instruments, averaging Doppler lidar radial velocities at various azimuth angles to avoid overfitting, and allowing the user to define a height grid by the user in the vipfile. The instruments used as input to the retrieval are a radar wind profiler (low- and high resolution mode) providing data in and above the boundary layer, a scanning Doppler lidar usually providing data throughout the boundary layer, a profiling lidar providing data from 50 to 200 m at BLOC and NANT, and from 10 to 280 m at Rhode Island, and a surface tower (4 m at NANT and RHOD and 10 m at BLOC). From the scanning lidars, we used radial velocity measurements at 60 deg elevation angle at six different azimuth angles with a resolution of approximately 30 m along the line of sight and the lowest range gate at approximately 70 m. The wind profiles are retrieved with WINDoe up to 3.74 km with 10 m vertical resolution. The profiles are retrieved every 15 min at BLOC and NANT and every 60 min at RHOD.

17 WIND ENERGY

Filling the Gaps: A Bayesian Mixture Model for Imputing Missing Soil Water Content Data

ABSTRACT Soil water content (SWC) data are central to evaluating how soil moisture varies over time and space and influences critical plant and ecosystem functions, especially in water‐limited drylands. However, sensors that record SWC at high frequencies often malfunction, leading to incomplete timeseries and limiting our understanding of dryland ecosystem dynamics. We developed an analytical approach to impute missing SWC data, which we tested at six eddy flux tower sites along an elevation gradient in the southwestern United States. We impute missing data as a mixture of linearly interpolated SWC between the observed endpoints of a missing data gap and SWC simulated by an ecosystem water balance model (SOILWAT2). Within a Bayesian framework, we allowed the relative utility (mixture weight) of each component (linearly interpolated vs. SOILWAT2) to vary by depth, site and gap characteristics. We explored “fixed” weights versus “dynamic” weights that vary as a function of cumulative precipitation, average temperature, and time since the start of the gap. Both models estimated missing SWC data well ( R 2 = 0.70–0.88 vs. 0.75–0.91 for fixed vs. dynamic weights, respectively), but the utility of linearly interpolated versus SOILWAT2 values depended on site and depth. SOILWAT2 was more useful for more arid sites, shallower depths, longer and warmer gaps and gaps that received greater precipitation. Overall, the mixture model reliably gap‐fills SWC, while lending insight into processes governing SWC dynamics. This approach to impute missing data could be adapted to accommodate more than two mixture components and other types of environmental timeseries.

Ogle, Kiona [School of Informatics, Computing, and

Identifying Controlling Variables for Mercury Vapors in Alpha-4 at Y-12: Two Year Data Collection Update

Multiple sensor packages were deployed at Alpha-4 by SRNL, in collaboration with United Cleanup Oak Ridge LLC (UCOR), to monitor mercury vapor concentrations and meteorological parameters. These sensors collected data, inside and outside of the legacy-use facility, for approximately two years. Though data gaps still exist, particularly in colder months, several controlling variables were identified that govern mercury vapor concentrations within Alpha-4. Temperature, barometric pressure gradients, humidity, and wind speed have been identified as controlling variables. A strong positive correlation was seen between mercury vapor concentrations and temperature which generally followed diurnal fluctuations. Temperatures below approximately 10 degrees Celsius did not show any spikes above the PEL, indicating more work can be performed at any time during the winter months – more data should be collected to confirm consistency in this finding. Additionally, spikes in mercury vapor concentrations above that of the permissible exposure limit (PEL; 100 µg/m 3 ) occurred primarily in late afternoon or evening/overnight hours (between 3 PM and 6 AM), which suggests D&D operations might be best scheduled during morning or daytime hours prior to the late afternoon. However, a limited number of spikes did occur outside of the identified window, although this may be attributed to disturbances in air flow and mercury vapor release from work activities performed inside of the Alpha-4 building. The analysis conducted allows for a strong predictive capability for estimating mercury vapor concentrations based upon accurate meteorological parameters. Still, additional data collection, particularly in the winter months, could help to strengthen the predictive power and validate the identified data trends. Further, increased temporal resolution could also help to better characterize the incipient stages of the increases and decreases in the mercury vapor concentration. Continued monitoring support by SRNL at Y-12 is underway at Alpha-4 to further close remaining data gaps and support deactivation and decommissioning work. Within a collaborative effort with UCOR, the SRNL team is collecting mercury vapor data to study the efficacy of a novel mercury suppressant, FerroBlack® which was recently deployed at Alpha-4. In addition, modifications to the current monitoring setup to increase measurement resolution is also being investigated.

54 ENVIRONMENTAL SCIENCES

Filling data analysis gaps in time-resolved crystallography by machine learning

There is a growing understanding of the structural dynamics of biological molecules fueled by x-ray crystallography experiments. Time-resolved serial femtosecond crystallography (TR-SFX) with x-ray Free Electron Lasers allows the measurement of ultrafast structural changes in proteins. Nevertheless, this technique comes with some limitations. One major challenge is the quality of data from TR-SFX measurements, which often faces issues like data sparsity, partial recording of Bragg reflections, timing errors, and pixel noise. To overcome these difficulties, conventionally, large volumes of data are collected and grouped into a few temporal bins. The data in each bin are then averaged and paired with the mean of their corresponding jittered timestamps. This procedure provides one structure per bin, resulting in a limited number of averaged structures for the entire time interval spanned by the experiment. Therefore, the information on ultrafast structural dynamics at high temporal resolution is lost. This has initiated research for advanced methods of analyzing experimental TR-SFX data beyond the standard binning and averaging method. To address this problem, we use a machine learning algorithm called Nonlinear Laplacian Spectral Analysis (NLSA), which has emerged as a promising technique for studying the dynamics of complex systems. In this work, we demonstrate the power of this algorithm using synthetic x-ray diffraction snapshots from a protein with significant data incompleteness, timing uncertainties, and noise. Our study confirms that NLSA is a suitable approach that effectively mitigates the effects of these artifacts in TR-SFX data and recovers accurate structural dynamics information hidden in such data.

Trujillo, Justin (ORCID:0000000285505360)

Bridging the Gap on Data, Metrics, and Analyses for Grid Resilience to Weather Events: Information that utilities can provide regulators, state energy offices, and other stakeholders

A growing number of states require regulated utilities to file resilience plans to improve the electric grid’s ability to anticipate, withstand, adapt to and recover from increasingly severe weather events. This report aims to help state regulators identify and request data, metrics, and analyses from utilities and use it in decisions on utility resilience plans and investments. The report reviews state requirements and utility plans focused on overall grid resilience, climate change resilience and vulnerabilities, infrastructure modernization, storm protection, and wildfire mitigation. It details types of data, metrics, and analyses across five categories--and provides examples of each from the utility plans. The first category is vulnerability assessments, or evaluations of the susceptibility of systems, communities, or assets to potential harm from identified hazards. The second is data on hazards and the exposure of utility assets and customers to these hazards. The third is attribute metrics, or system characteristics that contribute to or describe the resilience of a system. The fourth is performance metrics, which are impacts of resilience investments on system performance--typically a reduction of negative impacts from hazard events. Finally, evaluation and prioritization are analyses that utilities conduct to estimate impacts from resilience measures (evaluation) and prioritize measures based on costs and estimated impacts (prioritization). The report concludes with examples of key trends and emerging best practices for states and utilities, and identifies areas for further research.

24 POWER TRANSMISSION AND DISTRIBUTION