Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data gap analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Development and Preliminary Analysis of a U.S. Geothermal Heat Pump Installation Database

This paper seeks to addresses the significant gap in the literature regarding the installation and adoption of geothermal heat pump (GHP) systems in the United States. While the "2021 U.S. Geothermal Power Production and District Heating Market Report" published by the National Renewable Energy Laboratory (NREL) focused on direct-use geothermal district heating systems, it did not include an analysis of GHP installations (Robins et al. 2021). To bridge this gap, NREL has compiled a novel database currently containing 70,470 records of GHP installations, primarily sourced from state well permits and small-scale studies. Our methodology emphasizes the collection, cleaning, and standardization of data, addressing challenges such as inconsistent reporting formats and privacy concerns. Despite limitations in data on capacity, costs, and performance, our preliminary geospatial analysis reveals insights into the distribution of GHP systems across urban and rural areas and climate zones. The paper highlights the importance of publicly accessible data for advancing GHP technology adoption with a discussion of existing data sources and their limitations, advocating for improved collaboration between NREL and industry stakeholders.

data collection↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

Myriad World Baseline: Global Geodemographic Estimates

The LandScan Myriad World Baseline (MWB) method produces global, residential (nighttime/home-location) gridded geodemographic estimates based on 5-year age/gender cohorts—at 30-arcsecond (≈1 km) resolution. MWB is designed to fill gaps where detailed, georeferenced survey data (e.g., Demographic and Health Surveys (DHS)) are missing or outdated, and to provide a baseline that can support human security analysis, including consequence assessment, “patterns of life” modeling, and scenario-based population futures. MWB’s workflow spatializes household-level age/gender characteristics from the GLOPOP-S dataset by conflating household and gridded expected relative wealth adapted from Global Gridded Relative Deprivation Index (GRDI), then adjusts them to a target year of interest. Age/gender estimates are then applied to harmonize lowest-administrative-level statistics with LandScan residential counts, yielding final geodemographic estimates. Two validation case studies are presented: Ghana (2021) and Tokyo/Kanagawa, Japan (2020), illustrating spatial variability in demographic cohorts and comparing MWB outputs to official gridded statistics. Results show close overall alignment relative to validation criteria including population pyramids and age-dependency ratios.

Tuccillo, Joe [ORNL] (ORCID:0000000259300943)↗

Lifetime Assessment of the NEXT Ion Thruster

Ion thrusters are low thrust, high specific impulse devices with required operational lifetimes on the order of 10,000 to 100,000 hr. The NEXT ion thruster is the latest generation of ion thrusters under development. The NEXT ion thruster currently has a qualification level propellant throughput requirement of 450 kg of xenon, which corresponds to roughly 22,000 hr of operation at the highest throttling point. Currently, a NEXT engineering model ion thruster with prototype model ion optics is undergoing a long duration test to determine wear characteristics and establish propellant throughput capability. The NEXT thruster includes many improvements over previous generations of ion thrusters, but two of its component improvements have a larger effect on thruster lifetime. These include the ion optics with tighter tolerances, a masked region and better gap control, and the discharge cathode keeper material change to graphite. Data from the NEXT 2000 hr wear test, the NEXT long duration test, and further analysis is used to determine the expected lifetime of the NEXT ion thruster. This paper will review the predictions for all of the anticipated failure mechanisms. The mechanisms will include wear of the ion optics and cathode s orifice plate and keeper from the plasma, depletion of low work function material in each cathode s insert, and spalling of material in the discharge chamber leading to arcing. Based on the analysis of the NEXT ion thruster, the first failure mode for operation above a specific impulse of 2000 sec is expected to be the structural failure of the ion optics at 750 kg of propellant throughput, 1.7 times the qualification requirement. An assessment based on mission analyses for operation below a specific impulse of 2000 sec indicates that the NEXT thruster is capable of double the propellant throughput required by these missions.

VanNoord, Jonathan L.↗

Empirical Validation of UBEM: An Assessment of Bias in Urban Building Energy Modeling for Chicago

Residential and commercial buildings currently account for 30% of total global final energy consumption. Urban-scale building energy modeling (UBEM) can enable scalable investments and unlock building improvements by quantifying energy, demand, emissions, and cost reductions of specific measures or packages for building-specific technologies in large geographic regions. While the sophistication of UBEM data sources and technologies have increased dramatically in the past decade, there remains a knowledge gap for empirical validation and sources of bias between building-specific energy models and measured data at varying geographic scales.As UBEM continues to develop, systemic analysis of accuracy, bias, and limitations of the resulting models is necessary to inform best practices and move toward standardization. These are characterized for the Automatic Building Energy Modeling (AutoBEM) software suite with an initial case study involving metered electricity consumption data from 247,188 buildings in Chicago, Illinois, USA - averaged across years 2019-2021 - compared to the following datasets: (1) the AutoBEM-generated nation-scale Model America version 2 (MAv2) data for 596,064 buildings, (2) tax assessor data for 579,829 buildings, (3) tax assessor data filled with MAv2, and (4) 102 representative dynamic archetypes. The accuracy is reported for every building type and vintage combination, along with multiple sources of bias for unique building descriptors. The AutoBEM simulation workflow produced energy consumption estimates that closely match aggregated metered electricity consumption data for different types of buildings constructed during various time periods at the city scale - with initial normalized mean bias error of 10.9%, and 1.1% after removing outliers. Contribution of statistically significant factors including building type, land use, age, and size to variance in UBEM bias is quantified.

Garg, Ankur↗

A bi-level data-driven framework for fault-detection and diagnosis of HVAC systems

Long-term operation of heating, ventilation, and air conditioning (HVAC) systems will eventually lead to a range of HVAC system failures, resulting in excessive energy consumption and maintenance costs. Here, to avoid HVAC malfunctioning, fault detection diagnostic (FDD) is utilized as a common practice. Machine learning methods have lately received considerable interest for FDD analysis of HVAC systems due to their high detection accuracy. Meanwhile, HVAC malfunctions are regarded as rare occurrences, hence normal operating data samples are much more accessible than data samples in faulty and malfunctioning conditions. The dominating frequency of normal operation in HVAC datasets has also led to heavily biased classification algorithms within the literature. Moreover, the focus of previous literature has been on increasing the accuracy of the models which leads to a high number of false positives (misleading alarms) in the system. In order to enhance the performance of diagnostic procedures and fill the mentioned gaps, this study proposes a novel data-driven framework. A bi-level machine learning framework is developed for diagnosing faults in air handling units (AHUs) and rooftop units (RTUs) based on principal component analysis (PCA), time series anomaly detection, and random forest (RF). It is shown that PCA can reduce the dataset dimension with one principal component accounting for 95% of data variance. Also, the random forest could classify the faults with 89% precision for single-zone AHU, 85% precision for RTU, and 79% for multi-zone AHU. By proposing this framework, three persistent challenges are addressed: (I) minimizing false positives; (II) accounting for data imbalance; and (III) normal condition monitoring of equipment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

BIO-Plex Information System Concept

This paper describes a suggested design for an integrated information system for the proposed BIO-Plex (Bioregenerative Planetary Life Support Systems Test Complex) at Johnson Space Center (JSC), including distributed control systems, central control, networks, database servers, personal computers and workstations, applications software, and external communications. The system will have an open commercial computing and networking, architecture. The network will provide automatic real-time transfer of information to database server computers which perform data collection and validation. This information system will support integrated, data sharing applications for everything, from system alarms to management summaries. Most existing complex process control systems have information gaps between the different real time subsystems, between these subsystems and central controller, between the central controller and system level planning and analysis application software, and between the system level applications and management overview reporting. An integrated information system is vitally necessary as the basis for the integration of planning, scheduling, modeling, monitoring, and control, which will allow improved monitoring and control based on timely, accurate and complete data. Data describing the system configuration and the real time processes can be collected, checked and reconciled, analyzed and stored in database servers that can be accessed by all applications. The required technology is available. The only opportunity to design a distributed, nonredundant, integrated system is before it is built. Retrofit is extremely difficult and costly.

Jones, Harry↗

Measure sea level air pressure from space to improve knowledge and forecasting of the atmospheric state

Modern numerical weather prediction (NWP) and analysis models require globally-observed meteorological data including sea-level pressure (SLP) for accurate operations. Up until now, SLP has only been measured by in situ instruments from ships, buoys, and ocean platforms. These measurements are sparse with large gaps, leaving models starved of this critical information to constrain the atmospheric state. Recent advancements in differential absorption radar (DAR) provide a path to close this critical observation gap through spaceborne observations in the coming decade, improving the analysis models relied upon for atmospheric research and the weather forecasts depended upon daily for public safety and commerce.

Matthew L Walker McLinden↗

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance↗

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

Enhancing Electron Microscopy Image Classification Using Data Augmentation

Manual labeling for machine learning tasks such as image classification is tedious and labor-intensive; as a result, scientific datasets suitable for deep learning applications are scarce and limited. While data augmentation techniques have shown promise for extending image datasets, very little work has been done to understand the impact of combining multiple augmentation methods sequentially or the limits of their effectiveness when combined. Our work addresses this gap by examining how standard and combinatorial data augmentation affects the performance of machine learning models when trained on small datasets for label classification tasks. For our analysis, we generate single, double and quadruple-augmented datasets for a microscopy image classification task using six standard augmentation methods, and compare the resultant improvements observed in binary classification accuracy with three standard image classification models (DenseNet169, MobileNetV2, ResNet101V2). Our experiments show a non-monotonic relationship between the number of simultaneous augmentation methods and classification accuracy, indicating that there is a trade-off between the degree of augmentation and the model performance. These findings suggest that the optimal number of augmentation methods will vary by domain and use case. We also find that the order in which augmentation methods are applied to a limited dataset matters when combining augmentation schemes, with our use case showing performance differences up to 2.6% when the augmentation order is reversed for double-augmented datasets. Our work offers insights to the limits of data augmentation when working on image classification tasks with limited datasets.

Welsman, Jordan A↗

Fire, insect and disease‐caused tree mortalities increased in forests of greater structural diversity during drought

Abstract Structural diversity is an emerging dimension of biodiversity that accounts for size variations in organs among individuals in a community. Previous studies show significant effects of structural diversity on forest growth, but its effects on forest mortality are not known, particularly at a large scale. To address this knowledge gap, we quantified structural diversity using stem structural diversity (SSD) based on both tree diameter and height. We obtained U.S. Forest Service Forest Inventory and Analysis (FIA) data from over 2400 plots across southcentral U.S. forests that have suffered a recent drought. Using data from multiple sampling times, we calculated SSD and compared the relative importance of SSD, species diversity, functional diversity and other stand attributes in determining tree mortalities caused by fire, insects and diseases. We also used FIRETEC, a physics‐based fire model, to test the effect of SSD on canopy consumption by fire. Our results showed that (1) SSD was positively associated with tree mortalities caused by all three disturbances; (2) species richness was negatively associated with insect‐ and disease‐caused mortalities; (3) functional diversity was negatively associated with fire‐ and disease‐caused mortalities and (4) more phylogenetically related species had more similar mortality rates by insect and disease but not fire. Moreover, the FIRETEC model showed increasing canopy consumption by fire in stands with greater SSD. Together, the different tree mortalities during drought associated with SSD more consistently than the other biodiversity metrics were evaluated. Synthesis . Our results suggest that SSD could be considered in modelling forest dynamics and planning management to sustain forest health under disturbances.

54 ENVIRONMENTAL SCIENCES↗

GeoTGo: AI/ML software for development of community geothermal resources

For effective and equitable outcomes in achieving the national goal of net-zero carbon emissions, communities must be not only included, but even lead the implementation of innovative green-energy technologies. Collaborations with communities should happen through informed decision-making, community-centered research and engagement of stakeholders at the local, state, and regional levels. Community-led research and implementation are fundamental to achieving success. These collaborations include rule makers, environmental regulators, clean energy industries, and technology researchers and developers. Unfortunately, many green infrastructure initiatives still adhere to a top-down and expert-driven process of site selection and design without awareness and acknowledgment of public engagement needs. This can lead to costly delays, including lawsuits, and ultimately less than desired or lacking outcomes as well as missed opportunities1. Geothermal, like many new technologies whose social and economic impacts are not fully understood, often cause disproportionately high adverse effects on disadvantaged communities. These effects can be related to human health, environmental, climate, and other cumulative impacts, as well as the accompanying economic challenges of these impacts. We are focusing our work on the needs of the New Mexico Native American Pueblos and Tribes (NMP&T). To address these needs, we are developing a novel web-based interactive software and user friendly interface called GeoTGO (https://geotgo.com) that provides everything that is needed for communities to better understand and develop their geothermal resources. We will bridge the gap between technology advancements and community needs by facilitating the interactions between the geothermal industry, regulators, stakeholders, and end-users. GeoTGO will merge data, software (including data analysis, text mining, artificial intelligence, and modeling tools), knowledge, expertise, and experience to provide fast processing and dissemination of the latest information about cutting-edge geothermal technologies to users and communities. More information about the project is available at https://envitrace.com/projects/geotgo.html.

15 GEOTHERMAL ENERGY↗

A preliminary assessment of the accuracy of selected meteorological parameters determined from Nimbus 6 satellite profile data

Published rms errors in rawinsonde data and discrepancies between satellite and rawinsonde profile data for temperature, dewpoint temperature, mixing ratio, and wind speed. Satellite rms errors were found to be 2 to 3 times as large as those for rawinsonde data. Gradients of the preceding parameters were computed for both rawinsonde and satellite data and compared with means and near extreme values computed from the AVE 2 and AVE 4 experiments. In all cases, it was found that satellite data can be used to determine with relatively good accuracy the near extreme gradients but not those whose value does not exceed the average. Synoptic charts were prepared to show that patterns of temperature could be determined with relatively good accuracy, while those of dew point were not as good as those for temperature. Winds represented by cloud motion vectors (satellite winds) were compared with rawinsonde winds, and it was found that large gaps exist in satellite values for a given pressure level and that errors in the satellite determined concluded that satellite profile data are very useful in synoptic analysis, particularly in data sparse regions as well as regions where near extreme gradients exist in the measured parameters.

Scoggins, J. R.↗

Advanced Rainbow Solar Photovoltaic Arrays

Photovoltaic arrays of the rainbow type, equipped with light-concentrator and spectral-beam-splitter optics, have been investigated in a continuing effort to develop lightweight, high-efficiency solar electric power sources. This investigation has contributed to a revival of the concept of the rainbow photovoltaic array, which originated in the 1950s but proved unrealistic at that time because the selection of solar photovoltaic cells was too limited. Advances in the art of photovoltaic cells since that time have rendered the concept more realistic, thereby prompting the present development effort. A rainbow photovoltaic array comprises side-by-side strings of series-connected photovoltaic cells. The cells in each string have the same bandgap, which differs from the bandgaps of the other strings. Hence, each string operates most efficiently in a unique wavelength band determined by its bandgap. To obtain maximum energy-conversion efficiency and to minimize the size and weight of the array for a given sunlight input aperture, the sunlight incident on the aperture is concentrated, then spectrally dispersed onto the photovoltaic array plane, whereon each string of cells is positioned to intercept the light in its wavelength band of most efficient operation. The number of cells in each string is chosen so that the output potentials of all the strings are the same; this makes it possible to connect the strings together in parallel to maximize the output current of the array. According to the original rainbow photovoltaic concept, the concentrated sunlight was to be split into multiple beams by use of an array of dichroic filters designed so that each beam would contain light in one of the desired wavelength bands. The concept has since been modified to provide for dispersion of the spectrum by use of adjacent prisms. A proposal for an advanced version calls for a unitary concentrator/ spectral-beam-splitter optic in the form of a parabolic curved Fresnel-like prism array with panels of photovoltaic cells on two sides (see figure). The surface supporting the solar cells can be adjusted in length or angle to accommodate the incident spectral pattern. An unoptimized prototype assembly containing ten adjacent prisms and three photovoltaic cells with different bandgaps (InGaP2, GaAs, and InGaAs) was constructed to demonstrate feasibility. The actual array will consist of a lightweight thin-film silicon layer of prisms curved into a parabolic shape. In an initial test under illumination of 1 sun at zero airmass, the energy-conversion efficiency of the assembly was found to be 20 percent. Further analysis of the data from this test led to a projected energy conversion efficiency as high as 41 percent for an array of 6 cells or strings (GaP, AlGaAs, InGaP2, GaAs, and two different InGaAs cells or strings).

Mardesich, Nick↗

Security Vulnerability Profiles of Mission Critical Software: Empirical Analysis of Security Related Bug Reports

While some prior research work exists on characteristics of software faults (i.e., bugs) and failures, very little work has been published on analysis of software applications vulnerabilities. This paper aims to contribute towards filling that gap by presenting an empirical investigation of application vulnerabilities. The results are based on data extracted from issue tracking systems of two NASA missions. These data were organized in three datasets: Ground mission IVV issues, Flight mission IVV issues, and Flight mission Developers issues. In each dataset, we identified security related software bugs and classified them in specific vulnerability classes. Then, we created the security vulnerability profiles, i.e., determined where and when the security vulnerabilities were introduced and what were the dominating vulnerabilities classes. Our main findings include: (1) In IVV issues datasets the majority of vulnerabilities were code related and were introduced in the Implementation phase. (2) For all datasets, around 90 of the vulnerabilities were located in two to four subsystems. (3) Out of 21 primary classes, five dominated: Exception Management, Memory Access, Other, Risky Values, and Unused Entities. Together, they contributed from 80 to 90 of vulnerabilities in each dataset.

Goseva-Popstojanova, Katerina↗

Developing A Continuous Ozone Record Through the SAGE and Aura Missions With NASA Reanalysis Products

During the last quarter of the 20th century, the Stratospheric Aerosol and Gas Experiment (SAGE) missions were crucial in monitoring the loss and the subsequent recovery of the stratospheric ozone layer. Due to the employed solar occultation and self-calibration method, the SAGE monitors have produced stable data throughout the lifetime of each instrument. However, over ten years passed between the end of the SAGE II and SAGE III/M3M missions in 2005 and the launch of SAGE III/ISS instrument in 2017, leaving a gap in the data that much be bridged in order to assess the trends in the ozone record. Reanalysis products, such as the Modern-Era Retrospective analysis for Research and Applications, version 2 (MERRA-2), are attractive candidates for trend analysis due to the statistically optimized combination of multiple observing systems and the regular temporal and spatial coverage. In this study, we explore using the SAGE records to develop a stable reanalysis data product, suitable for trend analysis, from the start of the SAGE II record in 1984 through the present. Changes in the assimilated observation systems can introduce discontinuities within the MERRA-2 ozone record, such as in 2004 when the MERRA-2 system shifted from assimilating ozone retrievals collected by SBUV instruments to those collected by instruments onboard the Aura satellite. We follow the radiative transfer procedure outlined by Wargan et al. (2018) to address discontinuities in the MERRA-2 ozone dataset at the 2004 transition and during the Aura record. SAGE II ozone profiles are used to address discontinuities in upper stratospheric ozone associated with changes in the MERRA-2 meteorological observing system in 1998 and 1995. Lastly, we will use the resulting bias-corrected MERRA-2 ozone fields to assess the relative performance of the data from different SAGE sensors.

SAGE↗