Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data gap analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

Myriad World Baseline: Global Geodemographic Estimates

The LandScan Myriad World Baseline (MWB) method produces global, residential (nighttime/home-location) gridded geodemographic estimates based on 5-year age/gender cohorts—at 30-arcsecond (≈1 km) resolution. MWB is designed to fill gaps where detailed, georeferenced survey data (e.g., Demographic and Health Surveys (DHS)) are missing or outdated, and to provide a baseline that can support human security analysis, including consequence assessment, “patterns of life” modeling, and scenario-based population futures. MWB’s workflow spatializes household-level age/gender characteristics from the GLOPOP-S dataset by conflating household and gridded expected relative wealth adapted from Global Gridded Relative Deprivation Index (GRDI), then adjusts them to a target year of interest. Age/gender estimates are then applied to harmonize lowest-administrative-level statistics with LandScan residential counts, yielding final geodemographic estimates. Two validation case studies are presented: Ghana (2021) and Tokyo/Kanagawa, Japan (2020), illustrating spatial variability in demographic cohorts and comparing MWB outputs to official gridded statistics. Results show close overall alignment relative to validation criteria including population pyramids and age-dependency ratios.

Tuccillo, Joe [ORNL] (ORCID:0000000259300943)↗

Empirical Validation of UBEM: An Assessment of Bias in Urban Building Energy Modeling for Chicago

Residential and commercial buildings currently account for 30% of total global final energy consumption. Urban-scale building energy modeling (UBEM) can enable scalable investments and unlock building improvements by quantifying energy, demand, emissions, and cost reductions of specific measures or packages for building-specific technologies in large geographic regions. While the sophistication of UBEM data sources and technologies have increased dramatically in the past decade, there remains a knowledge gap for empirical validation and sources of bias between building-specific energy models and measured data at varying geographic scales.As UBEM continues to develop, systemic analysis of accuracy, bias, and limitations of the resulting models is necessary to inform best practices and move toward standardization. These are characterized for the Automatic Building Energy Modeling (AutoBEM) software suite with an initial case study involving metered electricity consumption data from 247,188 buildings in Chicago, Illinois, USA - averaged across years 2019-2021 - compared to the following datasets: (1) the AutoBEM-generated nation-scale Model America version 2 (MAv2) data for 596,064 buildings, (2) tax assessor data for 579,829 buildings, (3) tax assessor data filled with MAv2, and (4) 102 representative dynamic archetypes. The accuracy is reported for every building type and vintage combination, along with multiple sources of bias for unique building descriptors. The AutoBEM simulation workflow produced energy consumption estimates that closely match aggregated metered electricity consumption data for different types of buildings constructed during various time periods at the city scale - with initial normalized mean bias error of 10.9%, and 1.1% after removing outliers. Contribution of statistically significant factors including building type, land use, age, and size to variance in UBEM bias is quantified.

Garg, Ankur↗

A bi-level data-driven framework for fault-detection and diagnosis of HVAC systems

Long-term operation of heating, ventilation, and air conditioning (HVAC) systems will eventually lead to a range of HVAC system failures, resulting in excessive energy consumption and maintenance costs. Here, to avoid HVAC malfunctioning, fault detection diagnostic (FDD) is utilized as a common practice. Machine learning methods have lately received considerable interest for FDD analysis of HVAC systems due to their high detection accuracy. Meanwhile, HVAC malfunctions are regarded as rare occurrences, hence normal operating data samples are much more accessible than data samples in faulty and malfunctioning conditions. The dominating frequency of normal operation in HVAC datasets has also led to heavily biased classification algorithms within the literature. Moreover, the focus of previous literature has been on increasing the accuracy of the models which leads to a high number of false positives (misleading alarms) in the system. In order to enhance the performance of diagnostic procedures and fill the mentioned gaps, this study proposes a novel data-driven framework. A bi-level machine learning framework is developed for diagnosing faults in air handling units (AHUs) and rooftop units (RTUs) based on principal component analysis (PCA), time series anomaly detection, and random forest (RF). It is shown that PCA can reduce the dataset dimension with one principal component accounting for 95% of data variance. Also, the random forest could classify the faults with 89% precision for single-zone AHU, 85% precision for RTU, and 79% for multi-zone AHU. By proposing this framework, three persistent challenges are addressed: (I) minimizing false positives; (II) accounting for data imbalance; and (III) normal condition monitoring of equipment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance↗

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

Enhancing Electron Microscopy Image Classification Using Data Augmentation

Manual labeling for machine learning tasks such as image classification is tedious and labor-intensive; as a result, scientific datasets suitable for deep learning applications are scarce and limited. While data augmentation techniques have shown promise for extending image datasets, very little work has been done to understand the impact of combining multiple augmentation methods sequentially or the limits of their effectiveness when combined. Our work addresses this gap by examining how standard and combinatorial data augmentation affects the performance of machine learning models when trained on small datasets for label classification tasks. For our analysis, we generate single, double and quadruple-augmented datasets for a microscopy image classification task using six standard augmentation methods, and compare the resultant improvements observed in binary classification accuracy with three standard image classification models (DenseNet169, MobileNetV2, ResNet101V2). Our experiments show a non-monotonic relationship between the number of simultaneous augmentation methods and classification accuracy, indicating that there is a trade-off between the degree of augmentation and the model performance. These findings suggest that the optimal number of augmentation methods will vary by domain and use case. We also find that the order in which augmentation methods are applied to a limited dataset matters when combining augmentation schemes, with our use case showing performance differences up to 2.6% when the augmentation order is reversed for double-augmented datasets. Our work offers insights to the limits of data augmentation when working on image classification tasks with limited datasets.

Welsman, Jordan A↗

Fire, insect and disease‐caused tree mortalities increased in forests of greater structural diversity during drought

Abstract Structural diversity is an emerging dimension of biodiversity that accounts for size variations in organs among individuals in a community. Previous studies show significant effects of structural diversity on forest growth, but its effects on forest mortality are not known, particularly at a large scale. To address this knowledge gap, we quantified structural diversity using stem structural diversity (SSD) based on both tree diameter and height. We obtained U.S. Forest Service Forest Inventory and Analysis (FIA) data from over 2400 plots across southcentral U.S. forests that have suffered a recent drought. Using data from multiple sampling times, we calculated SSD and compared the relative importance of SSD, species diversity, functional diversity and other stand attributes in determining tree mortalities caused by fire, insects and diseases. We also used FIRETEC, a physics‐based fire model, to test the effect of SSD on canopy consumption by fire. Our results showed that (1) SSD was positively associated with tree mortalities caused by all three disturbances; (2) species richness was negatively associated with insect‐ and disease‐caused mortalities; (3) functional diversity was negatively associated with fire‐ and disease‐caused mortalities and (4) more phylogenetically related species had more similar mortality rates by insect and disease but not fire. Moreover, the FIRETEC model showed increasing canopy consumption by fire in stands with greater SSD. Together, the different tree mortalities during drought associated with SSD more consistently than the other biodiversity metrics were evaluated. Synthesis . Our results suggest that SSD could be considered in modelling forest dynamics and planning management to sustain forest health under disturbances.

54 ENVIRONMENTAL SCIENCES↗

GeoTGo: AI/ML software for development of community geothermal resources

For effective and equitable outcomes in achieving the national goal of net-zero carbon emissions, communities must be not only included, but even lead the implementation of innovative green-energy technologies. Collaborations with communities should happen through informed decision-making, community-centered research and engagement of stakeholders at the local, state, and regional levels. Community-led research and implementation are fundamental to achieving success. These collaborations include rule makers, environmental regulators, clean energy industries, and technology researchers and developers. Unfortunately, many green infrastructure initiatives still adhere to a top-down and expert-driven process of site selection and design without awareness and acknowledgment of public engagement needs. This can lead to costly delays, including lawsuits, and ultimately less than desired or lacking outcomes as well as missed opportunities1. Geothermal, like many new technologies whose social and economic impacts are not fully understood, often cause disproportionately high adverse effects on disadvantaged communities. These effects can be related to human health, environmental, climate, and other cumulative impacts, as well as the accompanying economic challenges of these impacts. We are focusing our work on the needs of the New Mexico Native American Pueblos and Tribes (NMP&T). To address these needs, we are developing a novel web-based interactive software and user friendly interface called GeoTGO (https://geotgo.com) that provides everything that is needed for communities to better understand and develop their geothermal resources. We will bridge the gap between technology advancements and community needs by facilitating the interactions between the geothermal industry, regulators, stakeholders, and end-users. GeoTGO will merge data, software (including data analysis, text mining, artificial intelligence, and modeling tools), knowledge, expertise, and experience to provide fast processing and dissemination of the latest information about cutting-edge geothermal technologies to users and communities. More information about the project is available at https://envitrace.com/projects/geotgo.html.

15 GEOTHERMAL ENERGY↗

Identification of high-dielectric constant compounds from statistical design

Abstract The discovery of high-dielectric materials is crucial to increasing the efficiency of electronic devices and batteries. Here, we report three previously unexplored materials with very high dielectric constants (69 < ϵ < 101) and large band gaps (2.9 < E g (eV) < 5.5) obtained by screening materials databases using statistical optimization algorithms aided by artificial neural networks (ANN). Two of these new dielectrics are mixed-anion compounds (Eu 5 SiCl 6 O 4 and HoClO) and are shown to be thermodynamically stable against common semiconductors via phase diagram analysis. We also uncovered four other materials with relatively large dielectric constants (20 < ϵ < 40) and band gaps (2.3 < E g (eV) < 2.7). While the ANN training-data are obtained from the Materials Project, the search-space consists of materials from the Open Quantum Materials Database (OQMD)—demonstrating a successful implementation of cross-database materials design. Overall, we report the dielectric properties of 17 materials calculated using ab initio calculations, that were selected in our design workflow. The dielectric materials with high-dielectric properties predicted in this work open up further experimental research opportunities.

36 MATERIALS SCIENCE↗

Survey of Modeling and Simulation Techniques for Advanced Manufacturing Technologies Volume II – Predicting Material Performance from Material Microstructure

This report describes the current state of modeling and simulation techniques for predicting the properties of materials fabricated with advanced manufacturing techniques, given the initial microstructure of the material. The report includes a literature survey and a gap analysis outlining and prioritizing key issues in applying these modeling and simulation techniques to nuclear reactor structural materials. The discussion covers both physics-based and data-driven modeling techniques and includes a broad range of manufacturing techniques and materials that may have future nuclear applications. This report is the second in a two-part series, with the first report covering modeling and simulation methods for predicting the initial, as-manufactured structure of advanced manufacturing materials, given a description of the process. Both reports focus on a set of manufacturing technologies likely to be applied to reactor structural components. Taken together, the two reports provide a complete summary of the current state of processing-structure-properties models for advanced manufacturing as well as a survey of applications to reactor structural materials

42 ENGINEERING↗

Energy-Storing Cryogenic Carbon Capture™ for Utility and Industrial-scale Processes (Final Report)

This project leverages work completed under funding from previous U.S. Department of Energy (DOE) projects and other sources of funding around the Cryogenic Carbon Capture Process™ (CCC). The purpose of this project is to further develop a perturbation of the CCC process that stores energy called the Cryogenic Carbon Capture™ Energy Storing (CCC-ES) process. The successful completion of the tasks in this project has prepared the CCC-ES process for scaling and developed initial plans and a feasibility analysis for an engineering-scale or small-commercial scale system that can store at least 10 MWh of energy. The specific areas of work in this project are: (1) Develop a conceptual, site-specific integrated process flow diagram: Sustainable Energy Solutions (SES) developed a conceptual, site-specific plan including a process flow diagram for integrating CCC-ES at a specific location. The proposed system was based on a typical engineering-scale CCC system designed to capture 30-50 metric tonnes of CO 2 per day (TPD), with modifications and additional changes and equipment put in place to add the energy storage element to the system. The 30-50 TPD scale corresponds to the required 10 MWh of stored energy. This task includes updating simulation software for transient analysis. SES also coordinated with the Jim Bridger Power Plant in Point of Rocks, Wyoming to get site-specific data. Using these data, SES completed site-specific preliminary process flow diagrams, estimated capital and operating costs for the CCC-ES system, and used the upgraded software package to estimate the primary figures of merit. (2) Complete a technoeconomic analysis (TEA): SES completed a full-scale TEA of the CCC-ES system to meet specific success criteria. This TEA used process information from the host site selected above. (3) Technology and data gap assessment for the CCC-ES process: As part of the technology gap assessment, SES performed an assessment of the best alternative energy storage solutions including figures of merit and operational strengths and weaknesses. SES then compared the proposed system with a discussion of how it addresses the weaknesses of alternative systems, showing for example a cost of stored energy at less than $50/MWh and with a roundtrip efficiency of greater than 95%. This assessment includes technical and other risk analyses and technology gaps with research and development needs to commercialize the process by 2030. Finally, work includes a commercialization roadmap/development pathway outlining the major milestones to move to a commercially operating CCC-ES system. (4) Phase II pre-FEED project plan: SES completed a project plan for the Phase II pre-FEED analysis using the same host site as in the Phase I conceptual study.

03 NATURAL GAS↗

Quantifying market volume sensitivity to material property modifications in polyhydroxybutyrate: A parametric analysis approach

Polyhydroxybutyrate (PHB), a biodegradable biopolymer, represents a promising alternative to petroleum-based thermoplastics. However, despite consistent market growth, PHB faces persistent commercialization challenges that limit widespread adoption. Existing research has focused predominantly on optimizing PHB production processes, leaving a critical gap in understanding which material property modifications would most effectively enhance market competitiveness. This study addresses this gap by systematically analyzing the relationship between polymer material properties and market performance using U.S. market data from 2008 to 2021 for 21 thermoplastic polymers across 19 material properties. We employed principal component regression to identify property modifications that could maximize market volume while reducing CO 2 emissions. Our parametric analysis revealed that two specific material properties – Hardness Shore A and Sheet Extrusion Temperature – significantly influence PHB marketability across different price points. Market simulations demonstrated that a 10% increase in Hardness Shore A could increase PHB market volume by 431.5 million kg while reducing emissions by 188.7 kg CO 2 . A similar 10% increase to Sheet Extrusion Temperature could yield a 297.5 million kg volume increase and a 99.2 kg CO 2 reduction in emissions. Critically, this approach is agnostic to the specific methods required to achieve these property changes, instead providing material scientists with quantitative, data-driven targets for R&D prioritization. Here, this framework offers a novel methodology for evaluating biopolymer competitiveness and supporting strategic decisions to accelerate PHB market adoption and contribute to decarbonization of the plastics industry.

09 BIOMASS FUELS↗

CHESS 2025: Leaf Area Index (LAI) for meadow, shrub, tree, and understory vegetation

This dataset contains Leaf Area Index (LAI) measurements made as part of the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Data were collected in the Upper Gunnison Basin, Colorado, across three study domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). Field observations of LAI were collected within 72 hours of airborne data collection by the National Ecological Observatory Network’s Aerial Observation Platform (NEON AOP). The NEON AOP collected waveform LiDAR (Light Detection and Ranging) and imaging spectrometer data in 426 spectral bands from the visible to shortwave infrared. LAI measurements were collected using the LICOR LAI-2200C Plant Canopy Analyzer following protocols outlined in the instrument manual (LI-COR 2019). Sampling targeted four distinct vegetation types: meadows, shrubs, trees, and aspen forest understory. We have archived data separately by site type because different field methods were used for each. At meadow sites, measurements were made at the four corners of 1m x 1m plots, with the instrument moving inward toward the center of the plot. At shrub sites, we measured the canopies of individual shrubs. At tree sites, we made measurements within a 10m x 10m subplot centered around a focal tree, with 30 observations taken on a regular grid. At aspen understory sites, we measured overstory trees following the tree protocol and understory herbaceous vegetation following the meadow protocol. All measurements included above-canopy (A) and below-canopy (B) readings, with specific protocols for scattering correction measurements in direct-sun conditions. Data were processed using the R package `rlai` (Worsham 2025). This package includes functions to calculate LAI, gap fraction, apparent clumping factor (Ω), scattering correction, and other canopy metrics. Package contents: Full file descriptions appear in ‘flmd.csv’. Files named according to the convention ‘lai_*_summary_data_cleaned.csv’ contain summary values of LAI, apparent clumping factor (Ωapp), and scattering correction factors for each site. These are the analysis-ready products that most data users will work with. Files named ‘lai_*_metadata_cleaned.csv’ contain additional site-level observations made during field collection. We have also archived intermediate and supplementary data for users who wish to check our processing approach or apply alternative methods. ‘raw_lai_2200C.zip’ contains the raw files as read from the LI-COR instrument, with no processing applied, in TXT format. The zip archive contains subdirectories by site type, which are further subdivided by sampling area. Filenames correspond to the sampling site number. ‘intermediate_results.zip’ contains detailed output from the processing routines, in JSON format. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘scattering_correction_logs.zip’ contains logfiles from the implementation of Kobayashi et al.'s (2013) scattering correction algorithm. The logfiles report values of several parameters at each iteration of the algorithm, as the model converges toward a stable solution. They are intended for users who want to verify scattering correction performance. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘spot_checks.csv’ reports LAI and other values for a small number of files processed with LI-COR FV2200 software (LI-COR 2013) using the same control parameters as in our R-based approach. Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). All zip files can be expanded with common archive utilities. TXT, CSV, and JSON files can be ingested into R or Python computing environments or read in common text editor utilities. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. * Todorov and Worsham are co–first authors.

2018 NEON and 2025 CHESS Campaigns↗

A Pattern-Recognition-Based Ensemble Data Imputation Framework for Sensors from Building Energy Systems

Building operation data are important for monitoring, analysis, modeling, and control of building energy systems. However, missing data is one of the major data quality issues, making data imputation techniques become increasingly important. There are two key research gaps for missing sensor data imputation in buildings: the lack of customized and automated imputation methodology, and the difficulty of the validation of data imputation methods. In this paper, a framework is developed to address these two gaps. First, a validation data generation module is developed based on pattern recognition to create a validation dataset to quantify the performance of data imputation methods. Second, a pool of data imputation methods is tested under the validation dataset to find an optimal single imputation method for each sensor, which is termed as an ensemble method. The method can reflect the specific mechanism and randomness of missing data from each sensor. The effectiveness of the framework is demonstrated by 18 sensors from a real campus building. The overall accuracy of data imputation for those sensors improves by 18.2% on average compared with the best single data imputation method.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Wildfire-Power Grid Interactions: Feedback, Impacts, Monitoring, Modeling, and Mitigation Strategies

Wildfires are increasingly interacting with electric power systems through a two-way hazard chain: fires damage grid assets and trigger cascading outages, while grid faults can ignite new fires under hot, dry, and windy conditions. This review synthesizes the state of knowledge across five domains: (i) physical impacts of flames, heat, and smoke on lines, towers, insulators, and substations; (ii) power-infrastructure-initiated ignitions via conductor clash, high-impedance faults, and corona discharge; (iii) widespread blackouts and disproportionate societal impacts; (iv) multi-scale monitoring spanning laboratory tests, in-situ and grid-integrated sensors, and Earth observation; (v) coupled modeling that links fire behavior with grid operations; and (vi) technological and strategic mitigation pathways spanning prevention, response, and recovery. We integrate these domains into a novel 'feedback-aware' socio-technical framework. Through a longitudinal analysis (2005-2025) of global incidents, we identify that while vegetation contact remains the most frequent ignition source, aging infrastructure failure has emerged as a critical driver of catastrophic 'mega-fires'. We further identify persistent gaps, including limited interoperability of high-frequency grid and environmental data, scarce real-time data assimilation, and under-developed equity metrics for outage management. We conclude by outlining a research agenda to (1) deploy interoperable sensing architectures, (2) advance feedback-coupled fire-grid simulations, and (3) evaluate mitigation portfolios through techno-economic and fairness lenses. Recognizing wildfire-grid interactions as coupled socio-technical systems is essential for protecting infrastructure and communities and for ensuring reliable, sustainable electricity in a changing world.

24 POWER TRANSMISSION AND DISTRIBUTION↗

QA/QC-ed Groundwater Level Time Series in PLM-1 and PLM-6 Monitoring Wells, East River, Colorado (2016-2022)

This data set contains QA/QC-ed (Quality Assurance and Quality Control) water level data for the PLM1 and PLM6 wells. PLM1 and PLM6 are location identifiers used by the Watershed Function SFA project for two groundwater monitoring wells along an elevation gradient located along the lower montane life zone of a hillslope near the Pumphouse location at the East River Watershed, Colorado, USA. These wells are used to monitor subsurface water and carbon inventories and fluxes, and to determine the seasonally dependent flow of groundwater under the PLM hillslope. The downslope flow of groundwater in combination with data on groundwater chemistry (see related references) can be used to estimate rates of solute export from the hillslope to the floodplain and river. QA/QC analysis of measured groundwater levels in monitoring wells PLM-1 and PLM-6 included identification and flagging of duplicated values of timestamps, gap filling of missing timestamps and water levels, removal of abnormal/bad and outliers of measured water levels. The QA/QC analysis also tested the application of different QA/QC methods and the development of regular (5-minute, 1-hour, and 1-day) time series datasets, which can serve as a benchmark for testing other QA/QC techniques, and will be applicable for ecohydrological modeling. The package includes a Readme file, one R code file used to perform QA/QC, a series of 8 data csv files (six QA/QC-ed regular time series datasets of varying intervals (5-min, 1-hr, 1-day) and two files with QA/QC flagging of original data), and three files for the reporting format adoption of this dataset (InstallationMethods, file level metadata (flmd), and data dictionary (dd) files).QA/QC-ed data herein were derived from the original/raw data publication available at Williams et al., 2020 (DOI: 10.15485/1818367). For more information about running R code file (10.15485_1866836_QAQC_PLM1_PLM6.R) to reproduce QA/QC output files, see README (QAQC_PLM_readme.docx). This dataset replaces the previously published raw data time series, and is the final groundwater data product for the PLM wells in the East River. Complete metadata information on the PLM1 and PLM6 wells are available in a related dataset on ESS-DIVE: Varadharajan C, et al (2022). https://doi.org/10.15485/1660962. These data products are part of the Watershed Function Scientific Focus Area collection effort to further scientific understanding of biogeochemical dynamics from genome to watershed scales. 2022/09/09 Update: Converted data files using ESS-DIVE’s Hydrological Monitoring Reporting Format. With the adoption of this reporting format, the addition of three new files (v1_20220909_flmd.csv, V1_20220909_dd.csv, and InstallationMethods.csv) were added. The file-level metadata file (v1_20220909_flmd.csv) contains information specific to the files contained within the dataset. The data dictionary file (v1_20220909_dd.csv) contains definitions of column headers and other terms across the dataset. The installation methods file (InstallationMethods.csv) contains a description of methods associated with installation and deployment at PLM1 and PLM6 wells. Additionally, eight data files were re-formatted to follow the reporting format guidance (er_plm1_waterlevel_2016-2020.csv, er_plm1_waterlevel_1-hour_2016-2020.csv, er_plm1_waterlevel_daily_2016-2020.csv, QA_PLM1_Flagging.csv, er_plm6_waterlevel_2016-2020.csv, er_plm6_waterlevel_1-hour_2016-2020.csv, er_plm6_waterlevel_daily_2016-2020.csv, QA_PLM6_Flagging.csv). The major changes to the data files include the addition of header_rows above the data containing metadata about the particular well, units, and sensor description. 2023/01/18 Update: Dataset updated to include additional QA/QC-ed water level data up until 2022-10-12 for ER-PLM1 and 2022-10-13 for ER-PLM6. Reporting format specific files (v2_20230118_flmd.csv, v2_20230118_dd.csv, v2_20230118_InstallationMethods.csv) were updated to reflect the additional data. R code file (QAQC_PLM1_PLM6.R) was added to replace the previously uploaded HTML files to enable execution of the associated code. R code file (QAQC_PLM1_PLM6.R) and ReadMe file (QAQC_PLM_readme.docx) were revised to clarify where original data was retrieved from and to remove local file paths.

54 ENVIRONMENTAL SCIENCES↗

Cybersecurity Standards for Distributed Energy Resources: Gaps and Harmonization Strategy

This report examines cybersecurity standards for Distributed Energy Resources (DERs) in light of their rapid growth and increasing integration into energy systems. It identifies critical gaps in existing frameworks, including inadequate coverage of DER-specific challenges, complexities in implementing comprehensive standards, integration issues with legacy systems, adoption hurdles for newer standards, and a lack of harmonization across regulatory landscapes. The analysis highlights vulnerabilities such as data integrity risks, unauthorized device control, and denial-of-service attacks across various DER technologies like solar PV, wind turbines, energy storage systems, and hydrogen fuel cells. The report proposes a harmonization strategy to address these deficiencies by developing unified cybersecurity requirements, certification programs, and training resources while fostering collaboration among stakeholders such as government agencies, industry groups, DER operators, manufacturers, and research institutions. A phased roadmap is outlined to refine and implement these measures through pilot testing and widespread adoption. Ultimately, the report underscores the urgent need for coordinated efforts to enhance DER cybersecurity and ensure the reliable operation of future energy systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗