Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data gap analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Offshore Geologic Carbon Storage Data Collection and Data Gaps Analysis

This is a TRS documenting the Offshore Geologic Carbon Storage Data Collection. It describes the Data Collection web application and its creation as well as an accompanying Data Gaps Assessment. We present an interactive data collection and data gaps analysis to aggregate, understand, and disseminate the data that are publicly available to support offshore GCS in the United States. This data collection and data gaps analysis can be leveraged by stakeholders to understand where GCS may be viable offshore, create GCS project analogs, and address challenges to GCS in offshore environments.

58 GEOSCIENCES↗

Geothermal Data Gap Analysis Over the Western US

NREL, as part of the Play Fairway Analysis Retrospective, compiled and mapped publicly available geologic and geophysical data in relation to the 2008 USGS geothermal potential analysis. Included in this submission are maps displaying the publicly available data for LIDAR coverage, aeromagnetic coverage, gravity station locations, and geologic map coverage over the Western United States.

15 GEOTHERMAL ENERGY↗

Life Cycle Inventories and Data Gap Analysis for Rare Earth Elements: Neodymium and Dysprosium from Mining to Magnets

The United States demand for Neodymium-Iron-Boron (NdFeB) magnets, produced from rare earth elements (REEs) such as (Nd) and Dysprosium (Dy), far exceeds its nascent domestic production capacity, rendering it reliant on vulnerable global supply chains dominated by China. To guide research and development investments in securing U.S. REE supply, defensible benchmark metrics across environmental, economic, and social dimensions are needed. In this study, we built globally-representative, process-based cradle-to-cradle life cycle inventories for Nd and Dy in NdFeB magnets lifecycles, encompassing primary material acquisition, beneficiation, smelting and refining, metal processing, specialty alloy and chemical transformation, subcomponent manufacturing, consumer application (use phase) and end-of-life management. We carried out detailed literature review, and applied process engineering principles to build industry-representative upscaled life cycle inventories for both metals. We used these models to conduct bottom-up literature review and gap analysis on existing literature, compilation of data sources for each life cycle stage (and transformations where necessary), and a preliminary technoeconomic analysis (TEA)/life cycle costing analysis (LCCA). Findings from this work emphasize the need for metal specific, representative REE LCIs to establish robust benchmarks for advancing sustainable REE technologies and guiding R&D in REE supply chains.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Filling data analysis gaps in time-resolved crystallography by machine learning

There is a growing understanding of the structural dynamics of biological molecules fueled by x-ray crystallography experiments. Time-resolved serial femtosecond crystallography (TR-SFX) with x-ray Free Electron Lasers allows the measurement of ultrafast structural changes in proteins. Nevertheless, this technique comes with some limitations. One major challenge is the quality of data from TR-SFX measurements, which often faces issues like data sparsity, partial recording of Bragg reflections, timing errors, and pixel noise. To overcome these difficulties, conventionally, large volumes of data are collected and grouped into a few temporal bins. The data in each bin are then averaged and paired with the mean of their corresponding jittered timestamps. This procedure provides one structure per bin, resulting in a limited number of averaged structures for the entire time interval spanned by the experiment. Therefore, the information on ultrafast structural dynamics at high temporal resolution is lost. This has initiated research for advanced methods of analyzing experimental TR-SFX data beyond the standard binning and averaging method. To address this problem, we use a machine learning algorithm called Nonlinear Laplacian Spectral Analysis (NLSA), which has emerged as a promising technique for studying the dynamics of complex systems. In this work, we demonstrate the power of this algorithm using synthetic x-ray diffraction snapshots from a protein with significant data incompleteness, timing uncertainties, and noise. Our study confirms that NLSA is a suitable approach that effectively mitigates the effects of these artifacts in TR-SFX data and recovers accurate structural dynamics information hidden in such data.

Trujillo, Justin (ORCID:0000000285505360)↗

Data Centers Gap Analysis [Slides]

Data centers and other large loads are a significant driver of unprecedented, near-term demand growth in the United States. Power system planners, utilities, regulators, and other stakeholders are grappling with how to integrate data centers on the system without comprising reliability, resiliency, and energy affordability. NLR is pursuing work to develop a siting and decision-making tool that would draw on power systems modeling expertise to achieve granular representation of trade-offs involved in data center sitting and development. This slide deck supports the same workstream by reviewing the literature to identify mitigation options to facilitate near-term integration of large loads and by presenting options for pursuing data development and/or modeling projects to improve representation of siting options.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

An Integrated ML/AI Framework for Digitizing, Structuring and Searching DOE U-TRU-Fuels Data with Gap Analysis of Non-DOE Records

The U.S. Department of Energy (DOE) Advanced Fuels Campaign (AFC) is advancing transmutation fuel technologies to reduce long-lived radioactive waste by converting minor actinides into shorter-lived or stable elements through irradiation in sodium-cooled fast reactors. Key experiments such as AFC-1, AFC-2, FUels for the transmutation of Trans-URanium elements In phéniX (FUTURIX)-Fortes Teneurs en Actinides (FTA), and Experimental Breeder Reactor-II (EBR-II) X501 have provided fuel fabrication, irradiation, and performance data on various transuranic-bearing fuel forms. This report documents the creation of an artificial-intelligence assisted database, which has consolidated all DOE-owned data related to Transuranic (TRU)-bearing fuel experiments and stored across it across both the Idaho National Laboratory (INL) Nuclear Data Management and Analysis System and the INL high performance computing (HPC) infrastructure. A dedicated webpage, hosted on the INL HPC system, has been developed to support role-based access and data interaction. The database architecture allows researchers to navigate large, heterogeneous archives with far greater speed and accuracy than manual search and lays the foundation for future expansion into multimodal nuclear materials analysis environments. The database represents a major step towards a nationally integrated fuels database utilizing artificial intelligence tools.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Carbon Storage Technical Viability Approach (CS TVA) Database

The Carbon Storage Technical Viability Approach (CS TVA) database was developed to support the implementation of the CS TVA Matrix to a national data availability assessment for technically viable carbon storage. This database leverages the efforts of multiple adjacent and overlapping databases by non-redundantly combining the databases into a single database along with additionally providing tags facilitating the CS TVA. The non-redundant aspect of the database permits an accurate assessment of the concentration of available data, aiding in spatial and categorical data gaps analysis relative to the individual CS TVA Matrix Components. Version 2.0 of the database is an expansion of Version 1.0. Version 2.0 was created to include additional data gathered to fill gaps in the existing data set. Downloading the CS TVA v2.0 database will result in two separate databases, the version 1.0 original .gdb, and a second addendum .gdb with the new data gathered, together these two databases make up v2.0. Please see the ReadMe file below for full details, metadata information, use disclaimer, and attributions.

Coal↗

Application of Multi-Criteria Decision Analysis Techniques for Informing Select Agent Designation and Decision Making

The Centers for Disease Control and Prevention (CDC) Select Agent Program establishes a list of biological agents and toxins that potentially threaten public health and safety, the procedures governing the possession, utilization, and transfer of those agents, and training requirements for entities working with them. Every 2 years the Program reviews the select agent list, utilizing subject matter expert (SME) assessments to rank the agents. In this study, we explore the applicability of multi-criteria decision analysis (MCDA) techniques and logic tree analysis to support the CDC Select Agent Program biennial review process, applying the approach broadly to include non-select agents to evaluate its generality. We conducted a literature search for over 70 pathogens against 15 criteria for assessing public health and bioterrorism risk and documented the findings for archiving. The most prominent data gaps were found for aerosol stability and human infectious dose by inhalation and ingestion routes. Technical review of published data and associated scoring recommendations by pathogen-specific SMEs was found to be critical for accuracy, particularly for pathogens with very few known cases, or where proxy data (e.g., from animal models or similar organisms) were used to address data gaps. Analysis of results obtained from a two-dimensional plot of weighted scores for difficulty of attack (i.e., exposure and production criteria) vs. consequences of an attack (i.e., consequence and mitigation criteria) provided greater fidelity for understanding agent placement compared to a 1-to-n ranking and was used to define a region in the upper right-hand quadrant for identifying pathogens for consideration as select agents. A sensitivity analysis varied the numerical weights attributed to various properties of the pathogens to identify potential quantitative (x and y) thresholds for classifying select agents. The results indicate while there is some clustering of agent scores to suggest thresholds, there are still pathogens that score close to any threshold, suggesting that thresholding “by eye” may not be sufficient. The sensitivity analysis indicates quantitative thresholds are plausible, and there is good agreement of the analytical results with select agent designations. A second analytical approach that applied the data using a logic tree format to rule out pathogens for consideration as select agents arrived at similar conclusions.

60 APPLIED LIFE SCIENCES↗

Bridging the Gap on Data and Analysis for Distribution System Planning: Information That Utilities Can Provide Regulators, State Energy Offices and Other Stakeholders

Electric utilities conduct planning annually to ensure their distribution system meets technical standards, policies, and regulations; addresses forecasted grid conditions; satisfies customer needs; and advances utility priorities. The plan identifies grid deficiencies, analyzes potential solutions, and prioritizes capital investments and other expenditures. About 20 U.S. states and jurisdictions require regulated utilities to file some type of distribution system plan with the public utility commission for review. Requirements for sharing distribution system data and analyses vary widely, from few specific requirements to a detailed list of information that must be provided. While utilities conduct extensive analysis to develop distribution system plans, in most jurisdictions regulators and stakeholders do not know what data are available and how the utility uses the data in planning and investing. This report aims to bridge the gap by increasing understanding of the types of data and analyses utilities employ to develop distribution system plans and how the information affects their decision-making. The report describes information that states and stakeholders can ask for related to 11 data categories: -Forecasting loads and distributed energy resources (DERs) -Scenario analysis -Worst-performing circuits -Asset management strategy -Hosting capacity analysis -Value of DERs -Grid needs assessment -Cost-effectiveness framework for investments -Distribution system investment strategy and implementation -Geotargeted programs -Non-wires alternatives procurements.

24 POWER TRANSMISSION AND DISTRIBUTION↗

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance↗

Analysis of Selected Publicly Available Geothermal Exploration Data Gaps

As part of a United States Department of Energy (DOE) supported retrospective analysis of DOE's Play Fairway Analysis (PFA) projects, the National Renewable Energy Laboratory (NREL) compiled and analyzed publicly available geothermal exploration datasets to identify and highlight data gaps in areas prospective for hosting geothermal resources. The analysis was intended to understand the existing geographic coverage of selected datasets commonly utilized both by the PFA projects and geothermal developers during resource assessments including geologic mapping, temperature gradient drilling, and aeromagnetic, gravimetric, and lidar surveys. Results indicate that broad areas of the western United States estimated to have geothermal potential lack sufficient geologic and geophysical coverage necessary for even regional resource exploration. The study directly informed the recent Geoscience Data Acquisition for Western Nevada, or GeoDAWN - which united DOE's Geothermal Technologies Office (GTO) with the U.S. Geological Survey (USGS) of the U.S. Department of the Interior to assist U.S. needs for energy and critical minerals. The study also has the potential to inform public investment in further data acquisition for characterization of the Earth both for geothermal and other natural resource assessments.

data↗

Machine Learning Assisted Gap-Filled Discharge Data for the East River Community Watershed, Colorado, for Water Years 2014-2021

This dataset contains a collection of machine learning assisted gap-filled discharge data created for all discharge stations across the East River Watershed, Colorado. This data was generated by using raw discharge data collected by Rosemary Carroll, and conducting a random forest machine learning analysis to gap-fill discharge data across all years at the hourly time level. Discharge data with gaps creates problems for analysis of measured and modeled fluxes of carbon and nitrogen exported out of each sub-watershed. Gap-filled data is also required as an input to surface water models, which helps to address our main research question related to how snowmelt timing impacts the timing and magnitude of nitrogen exports. Data is provided in one csv file.

54 ENVIRONMENTAL SCIENCES↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the Geological Society of America Connects 2024 Annual Meeting in Anaheim, California, 22-25 September 2024.

Creason, Christopher↗

Street context of various demographic groups in their daily mobility

Abstract We present an urban science framework to characterize phone users’ exposure to different street context types based on network science, geographical information systems (GIS), daily individual trajectories, and street imagery. We consider street context as the inferred usage of the street, based on its buildings and construction, categorized in nine possible labels. The labels define whether the street is residential, commercial or downtown, throughway or not, and other special categories. We apply the analysis to the City of Boston, considering daily trajectories synthetically generated with a model based on call detail records (CDR) and images from Google Street View. Images are categorized both manually and using artificial intelligence (AI). We focus on the city’s four main racial/ethnic demographic groups (White, Black, Hispanic and Asian), aiming to characterize the differences in what these groups of people see during their daily activities. Based on daily trajectories, we reconstruct most common paths over the street network. We use street demand (number of times a street is included in a trajectory) to detect each group’s most relevant streets and regions. Based on their street demand, we measure the street context distribution for each group. The inclusion of images allows us to quantitatively measure the prevalence of each context and points to qualitative differences on where that context takes place. Other AI methodologies can further exploit these differences. This approach presents the building blocks to further studies that relate mobile devices’ dynamic records with the differences in urban exposure by demographic groups. The addition of AI-based image analysis to street demand can power up the capabilities of urban planning methodologies, compare multiple cities under a unified framework, and reduce the crudeness of GIS-only mobility analysis. Shortening the gap between big data-driven analysis and traditional human classification analysis can help build smarter and more equal cities while reducing the efforts necessary to study a city’s characteristics.

Salgado, Ariel (ORCID:0000000177015372)↗

Application of multi-criteria decision analysis techniques and decision support framework for informing select agent designation for agricultural animal pathogens

The United States Department of Agriculture (USDA), Division of Agricultural Select Agents and Toxins (DASAT) established a list of biological agents and toxins (Select Agent List) that potentially threaten agricultural health and safety, the procedures governing the transfer of those agents, and training requirements for entities working with them. Every 2 years the USDA DASAT reviews the Select Agent List, using subject matter experts (SMEs) to perform an assessment and rank the agents. To assist the USDA DASAT biennial review process, we explored the applicability of multi-criteria decision analysis (MCDA) techniques and a Decision Support Framework (DSF) in a logic tree format to identify pathogens for consideration as select agents, applying the approach broadly to include non-select agents to evaluate its robustness and generality. We conducted a literature review of 41 pathogens against 21 criteria for assessing agricultural threat, economic impact, and bioterrorism risk and documented the findings to support this assessment. The most prominent data gaps were those for aerosol stability and animal infectious dose by inhalation and ingestion routes. Technical review of published data and associated scoring recommendations by pathogen-specific SMEs was found to be critical for accuracy, particularly for pathogens with very few known cases, or where proxy data (e.g., from animal models or similar organisms) were used to address data gaps. The MCDA analysis supported the intuitive sense that select agents should rank high on the relative risk scale when considering agricultural health consequences of a bioterrorism attack. However, comparing select agents with non-select agents indicated that there was not a clean break in scores to suggest thresholds for designating select agents, requiring subject matter expertise collectively to establish which analytical results were in good agreement to support the intended purpose in designating select agents. The DSF utilized a logic tree approach to identify pathogens that are of sufficiently low concern that they can be ruled out from consideration as a select agent. In contrast to the MCDA approach, the DSF rules out a pathogen if it fails to meet even one criteria threshold. Both the MCDA and DSF approaches arrived at similar conclusions, suggesting the value of employing the two analytical approaches to add robustness for decision making.

60 APPLIED LIFE SCIENCES↗

Development of a Pre-Combustion CO 2 Capture Process Using High-Temperature PBI Hollow-Fiber Membranes

The overall objective of this project was to evaluate the advantages of transformational polybenzimidazole (PBI) polymer hollow-fiber membrane (HFM)-based, carbon dioxide (CO 2 ) capture and purification technology at bench-scale using an actual coal-derived syngas stream from a coal gasification facility. The project was carried out over two budget periods. The technical objectives in Budget Period 1 (BP1) included preparing HFs and modules and upgrading the available skid for field testing. The technical objectives for BP2 were to field-test the skid unit with actual coal-derived syngas from an oxygen-blown gasifier to obtain performance data, update the Techno-Economic Analysis (TEA) that would assist with future process scale-up, and provide information on the design of a small pilot-scale test unit. The goal was to advance the PBI-HFM CO 2 capture and gas separation system for pre-combustion applications beyond second-generation economic performance predictions and make progress toward meeting overall fossil energy performance goals of CO 2 capture with 95% CO 2 purity at a cost of electricity (COE) 30% less than baseline capture approaches. The research program was designed with progressive technical tasks leading to both dynamic and steady-state testing of the PBI-HFM skid with actual coal-derived syngas. The work plan was to: (1) fabricate sufficient Generation-2 (GEN-2) fibers for module fabrication; (2) upgrade the fiber skid to accommodate large fiber modules for bench-scale field testing; (3) conduct dynamic and steady-state testing with coal-derived syngas from an oxygen-blown gasifier and obtain system performance data; (4) perform a TEA and environmental, health, and safety (EH&S) assessment; (5) update the State-Point Data Table, Technology Gap Analysis (TGA), and Technology Maturation Plan (TMP); (6) uninstall and return the test skid to the Recipient’s facilities; and (7) submit a Final Report that describes the results and analysis of the project research effort.

03 NATURAL GAS↗

Marine energy converters: Potential acoustic effects on fishes and aquatic invertebrates

The potential effects of underwater anthropogenic sound and substrate vibration from offshore renewable energy development on the behavior, fitness, and health of aquatic animals is a continuing concern with increased deployments and installation of these devices. Initial focus of related studies concerned offshore wind. However, over the past decade, marine energy devices, such as a tidal turbines and wave energy converters, have begun to emerge as additional, scalable renewable energy sources. Because marine energy converters (MECs) are not as well-known as other anthropogenic sources of potential disturbance, their general function and what is known about the sounds and substrate vibrations that they produce are introduced. Furthermore, while most previous studies focused on MECs and marine mammals, this paper considers the potential of MECs to cause acoustic disturbances affecting nearshore and tidal fishes and invertebrates. In particular, the focus is on particle motion and substrate vibration from MECs because these effects are the most likely to be detected by these animals. Finally, an analysis of major data gaps in understanding the acoustics of MECs and their potential impacts on fishes and aquatic invertebrates and recommendations for research needed over the next several years to improve understanding of these potential impacts are provided.

16 TIDAL AND WAVE POWER↗