Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data gap analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Domain Adaptation for Measurements of Strong Gravitational Lenses

Upcoming surveys are predicted to discover galaxy-scale strong lenses on the order of 10\textsuperscript{5}, making deep learning methods necessary in lensing data analysis. Currently, there is insufficient real lensing data to train deep learning algorithms, but the alternative of training only on simulated data results in poor performance on real data. Domain Adaptation may be able to bridge the gap between simulated and real datasets. We utilize domain adaptation for the estimation of Einstein radius ($\Theta_E$) in simulated galaxy-scale gravitational lensing images with different levels of observational realism. We evaluate two domain adaptation techniques - Domain Adversarial Neural Networks (DANN) and Maximum Mean Discrepancy (MMD). We train on a source domain of simulated lenses and apply it to a target domain of lenses simulated to emulate noise conditions in the Dark Energy Survey (DES). We show that both domain adaptation techniques can significantly improve the model performance on the more complex target domain dataset. This work is the first application of domain adaptation for a regression task in strong lensing imaging analysis. Our results show the potential of using domain adaptation to perform analysis of future survey data with a deep neural network trained on simulated data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Domain Adaptation for Measurements of Strong Gravitational Lenses

Upcoming surveys are predicted to discover galaxy-scale strong lenses on the order of $10^5$, making deep learning methods necessary in lensing data analysis. Currently, there is insufficient real lensing data to train deep learning algorithms, but the alternative of training only on simulated data results in poor performance on real data. Domain Adaptation may be able to bridge the gap between simulated and real datasets. We utilize domain adaptation for the estimation of Einstein radius ($\Theta_E$) in simulated galaxy-scale gravitational lensing images with different levels of observational realism. We evaluate two domain adaptation techniques - Domain Adversarial Neural Networks (DANN) and Maximum Mean Discrepancy (MMD). We train on a source domain of simulated lenses and apply it to a target domain of lenses simulated to emulate noise conditions in the Dark Energy Survey (DES). We show that both domain adaptation techniques can significantly improve the model performance on the more complex target domain dataset. This work is the first application of domain adaptation for a regression task in strong lensing imaging analysis. Our results show the potential of using domain adaptation to perform analysis of future survey data with a deep neural network trained on simulated data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Mobility Gaps between Low-Income and Not Low-Income Households: A Case Study in New York State

Understanding the travel challenges faced by low-income residents has always been and continues to be one of the most important transportation equity topics. This study aims to explore the mobility gaps between low-income households (HHs) and not low-income HHs, and how the gaps vary within different socio-demographic population groups in New York State (NYS). The latest National Household Travel Survey data was used as the primary data source for the analysis. The study first employed the K-prototype clustering algorithm to categorize the HHs in NYS based on their socio-demographic attributes. Five population groups were identified based on nine different household (HH) features such as HH size, vehicle ownership, and elderly status of its members. Then, the mobility differences, measured by trip frequency, trip distance, travel time, and person miles traveled, were examined among the five population groups. Results suggest that the individuals in low-income HHs consistently took fewer trips and made shorter trips compared to their not low-income counterparts in NYS. The travel distance gaps were most obvious among white HHs with more vehicles than drivers. In addition, while the population from low-income HHs made shorter trips on average (2.7 mi shorter per trip), they experienced longer travel time than those from not low-income HHs (1.8 min longer per trip). These key findings provide a deeper understanding of the travel behavior disparities between low-income and not low-income households. The findings could also support policymakers and transportation planners in addressing the critical needs of residents in low-income households in NYS and provide inputs for designing a more equitable transportation system.

Liu, Yuandong↗

Data-Informed Synthetic Networks of Water Distribution Systems for Resilience Analysis in Puerto Rico

The increasing potential of infrastructure disruptions calls for high-quality infrastructure models to be used in resilience analysis and decision making. Unfortunately, many utilities and communities do not have access to accurate and detailed models due to a lack of data and resources. Furthermore, security restrictions on sharing infrastructure models present roadblocks to research, analysis, and decision making. Recent advances in the development of synthetic water distribution models provide a potential solution to this problem. There is an opportunity to improve these methods by leveraging incomplete pipe datasets to aid synthetic network generation. To address this gap, we developed a methodology for synthetic network generation that incorporates partial pipe data using a modification of the minimum cost flow algorithm for network generation and pipe sizing. This methodology demonstrates how partial pipe data can be leveraged to improve site-specific synthetic network generation. For the study area of Mayagüez, Puerto Rico, a synthetic model generated using 50% of real pipe data matches the pressure of the validation system with an average error of 23.5 m of head, which improves upon the average error of 31.6 m of head produced by a synthetic model generated using no data of the real pipes. Additionally, synthetic networks are shown to replicate the pressure response under a disruption scenario of the validation network, suggesting potential use in resilience analysis.

resilience analysis↗

Defining and Measuring Forest Dependence in the United States: Operationalization and Sensitivity Analysis

This manuscript helps bridge a gap between theoretical work that advocates for a broad view of forest dependence, and empirical work that has focused narrowly on economic measures. Background: Forest dependence has been widely recognized as a valuable concept for understanding human communities’ well-being and vulnerability to shocks and changes. Past theoretical literature has highlighted the importance of recognizing various types of dependence—environmental, economic, and social—yet past empirical literature on the topic in the United States has almost exclusively relied on measures of economic dependence such as employment and earnings from the traditional forest products sector. Objective and Methods: As a first step to bridge the gap between the theoretical and empirical, we reviewed the existing, publicly available, reliable, wall-to-wall data sources to identify alternate proxy measures for forest dependence. Data availability made the analysis feasible only at the county level—the administrative subdivisions of the state—or higher. Results and Conclusions: We created environmental, economic, and social criteria based on threshold levels of the following proxy variables: forest area, earnings, employment, and indigenous population. Using these criteria, we identified 524 counties to be potentially forest-dependent of 3140 total counties in the United States. The largest concentration was in the Pacific Northwest and Southeast regions, and a higher proportion were non-metro counties than metro. Varying the threshold levels significantly changes the number of counties identified but does not alter the overall geographic trends.

54 ENVIRONMENTAL SCIENCES↗

A Comparative Analysis of Infrastructure-Based Perception Sensors for Intelligent Transportation Systems

The rise of privatized and public investment in smart city infrastructure and intelligent transportation systems has generated a heightened demand for perception sensors that effectively track and detect objects while being reliable in diverse weather and lighting conditions. This growing demand for perception sensors has accelerated their development and enhanced their capabilities. With these new capabilities, it is challenging to determine the most suitable sensing unit to use in each situation. Therefore, it is essential to have a comprehensive understanding of the benefits and limitations of each sensing unit to effectively leverage their capabilities. The purpose of this paper is to provide a detailed evaluation of various perception sensors. Additionally, this paper will demonstrate the benefits of combining multiple perception sensors, which complement each other by addressing data gaps inherent to single-sensor systems, to facilitate the creation of a digital twin that models the real world. The Infrastructure, Perception, and Control (IPC) team will conduct data analysis using data collected through field testing at traffic intersections in Colorado Springs, Colorado, to make comparisons between sensors. This research aims to provide clear and concise information about modern perception systems, which will support the development of intelligent transportation systems.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning BPS spectra and the gap conjecture

We explore statistical properties of Bogomol’nyi-Prasad-Sommerfield q-series for strongly coupled supersymmetric theories that correspond to a particular family of three-manifolds. We discover that gaps between exponents in the -series are statistically more significant at the beginning of the -series compared to gaps that appear in higher powers of. Our observations are obtained by calculating saliencies of -series features used as input data for principal component analysis, which is a standard example of an explainable machine learning technique that allows for a direct calculation and a better analysis of feature saliencies.

97 MATHEMATICS AND COMPUTING↗

Characterization and Quantification of Radiation-Induced Clusters/Precipitates in RPV Steels Using STEM-EDS and Machine Learning

Over the operational lifespan of a nuclear reactor, reactor pressure vessel (RPV) steels are subjected to significant neutron irradiation, resulting in complex microstructural changes and the consequent degradation of mechanical properties. Various physically motivated correlation models have been developed to predict neutron irradiation-induced embrittlement of RPVs under different irradiation conditions. However, the efficient and accurate characterizations and quantification of radiation-induced clusters in RPVs are still challenging, which will affect the precision of the predictive models for embrittlement of RPV components. In the DOE Visiting Faculty Program (VFP) research work at Oak Ridge National Lab (ORNL), I integrate machine learning to aid Scanning Transmission Electron Microscopy – Energy Dispersive X-ray Spectroscopy (STEM-EDS) analyses, which improve the characterization and quantification of radiation-induced clusters in RPV steels, thereby enabling more accurate predictions of material behavior under irradiation. The surveillance base- and welded- RPV steels were annealed at various temperatures of 340 °C, 450 °C and 500 °C for up to 168 hours, respectively. Afterwards, I have characterized radiation-induced clusters using advanced STEM-EDS techniques and subsequently applying machine learning algorithms to analyze and refine STEM-EDS datasets, enhancing the quantification of clusters compositions and distributions. In the end, an efficient workflow for integrating STEM-EDS data analysis with machine learning to address challenges including noise reduction has been developed. The completion of this VFP work will support bridge critical gaps in the accurate quantification of radiation-induced clusters in RPV steels using STEM-EDS and support the development of more precise models for predicting RPV embrittlement in the Light Water Reactor Sustainability program supported by Department of Energy and enhancing the collaboration between ORNL and Alred University. The outcome of the VFP project will leverage a few research papers submission to peer-reviewed journals in the relevant scientific field and a few oral presentations at national and international conferences.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Electrode Erosion and Prefire Studies Towards Fusion Scale Pulsed Power

This study presents a comprehensive investigation of electrode erosion and discharge behavior in spark gap switches over long switching cycle lifetimes. Brass, copper–tungsten (CuW), and stainless steel electrodes are tested under controlled conditions to quantify material degradation, debris accumulation, and changes in breakdown voltage. High-resolution imaging and statistical analysis of spark channel locations and gap breakdown voltages reveal how surface evolution influences long-term performance and reliability. These results provide essential data for lifetime modeling and inform design strategies for pulsed power systems in emerging applications such as private sector fusion energy and large-scale facilities like Sandia’s Z Machine and proposed ZX upgrades, where high repetition reliability and predictable behavior are critical.

electrical breakdown↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Axolotl: a scalable genomics library based on Apache Spark (Axolotl) v1.0.0

Axolotl is a Python library for scalable distributed genome and metagenome data analysis. Existing tools and systems that we rely on are struggling to keep up with the rapid explosion of genomic data. Compounding this issue, developing scalable solutions require a steep learning curve in parallel programming, which presents a barrier to academic researchers. While we do have scalable solutions for specific tasks, we lack comprehensive, end-to-end solutions. It's this gap in our toolkit that we aim to address with Axolotl. The Axolotl library is built for easy parallel processing, efficiently handling multiple tasks or large datasets simultaneously, and scaling up to meet the demands of extensive genomic data analysis.

Wang, Zhong↗

OzDES Reverberation Mapping Program: Stacking analysis with Hβ, Mg ii , and C iv

ABSTRACT Reverberation mapping is the leading technique used to measure direct black hole masses outside of the local Universe. Additionally, reverberation measurements calibrate secondary mass-scaling relations used to estimate single-epoch virial black hole masses. The Australian Dark Energy Survey (OzDES) conducted one of the first multi-object reverberation mapping surveys, monitoring 735 AGN up to z ∼ 4, over 6 years. The limited temporal coverage of the OzDES data has hindered recovery of individual measurements for some classes of sources, particularly those with shorter reverberation lags or lags that fall within campaign season gaps. To alleviate this limitation, we perform a stacking analysis of the cross-correlation functions of sources with similar intrinsic properties to recover average composite reverberation lags. This analysis leads to the recovery of average lags in each redshift-luminosity bin across our sample. We present the average lags recovered for the Hβ, Mg ii, and C iv samples, as well as multiline measurements for redshift bins where two lines are accessible. The stacking analysis is consistent with the Radius–Luminosity relations for each line. Our results for the Hβ sample demonstrate that stacking has the potential to improve upon constraints on the R–L relation, which have been derived only from individual source measurements until now.

79 ASTRONOMY AND ASTROPHYSICS↗

A global urban heat island intensity dataset: Generation, comparison, and analysis

The urban heat island (UHI) effect, a phenomenon of local warming over urban areas, is the most well-known impact of urbanization on climate. Globally consistent estimates of the UHI intensity (UHII) are crucial for examining this phenomenon across time and space. However, publicly available UHII datasets are limited and have several constraints: (1) they are for clear-sky surface UHII, not all-sky surface UHII and canopy (air temperature) UHII; (2) the estimation methods often neglect anthropogenic disturbance, introducing uncertainties in the estimated UHII. To address these issues, this study proposes a new dynamic equal-area (DEA) method that can minimize the influence of various confounding factors on UHII estimates through a dynamic cyclic process. Utilizing the DEA method and leveraging various gridded temperature data, we develop a global-scale (>10,000 cities), long-term (over 20 years by month), and multi-faceted (clear-sky surface, all-sky surface, and canopy) UHII dataset. Further, based on these estimates, we provide a comprehensive analysis of the UHII and its trends in global cities. The UHII is found to be greater than zero in >80% of cities, with global annual average magnitudes around 1.0 °C (day) and 0.8 °C (night) for surface UHII, and close to 0.5 °C for canopy UHII. Furthermore, an interannual upward trend in UHII is observed in >60% of cities, with global annual average trends exceeding 0.1 °C/decade (day) and over 0.06 °C/decade (night) for surface UHII, and slightly surpassing 0.03 °C/decade for canopy UHII. Notably, there exists a positive correlation between the magnitude and trend of UHII, suggesting that cities with stronger UHII tend to experience faster growth in UHII. Additionally, discrepancies in UHII are found between different temperature data, stemming not only from distinctions in data types (surface or air temperature) but also from differences in data acquisition times (Terra or Aqua), weather conditions (clear-sky or all-sky), and processing methodologies (with or without gap filling). Overall, our proposed method, dataset, and analysis results have the potential to provide valuable insights for future urban climate studies. The UHII dataset is publicly available at https://doi.org/10.6084/m9.figshare.24821538.

54 ENVIRONMENTAL SCIENCES↗

Methanol adsorption and dissociation on GaP(110) studied by ambient pressure X-ray photoelectron spectroscopy

Ambient pressure X-ray photoelectron spectroscopy (AP-XPS) was used to investigate methanol (CH 3 OH) adsorption and reaction on the GaP(110) surface. Exposure of CH 3 OH to GaP(110) at room temperature led to the formation of at least four different surface species as indicated by analysis of C 1s and O 1s XPS features. By combining AP-XPS data with density functional theory calculations, the surface species were identified as methoxy (CH 3 O*), formaldehyde (CH 2 O*), and paired methanol (p-CH 3 O*H) and methoxy (p-CH 3 O*) species, where “paired” means that they belong to a hydrogen-bonded methoxy-methanol complex. Asterisk * here indicates an adsite. The formation of CH 2 O* via the dehydrogenation of CH 3 O* was shown to be limited by the availability of vacant phosphorus (P) sites on GaP(110). With an increase in CH 3 OH pressure, the fractional coverage of CH 3 O* species reached 0.55, and the surface P sites were completely saturated with hydrogen. Under a constant CH 3 OH pressure of 0.5 Torr, the surface concentration of the paired species and of CH 2 O* remained constant until 400 K. At higher temperatures, thermally driven reactions led to a significant increase in the concentration of surface CH x * species, which suggests that C-O bond cleavage of the CH 3 O group is the dominant decomposition mechanism on GaP(110). In conclusion, based on the reactivity of GaP(110) toward CH 3 OH dehydrogenation, elevated temperatures and CH 3 OH pressures may be used to functionalize this surface.

36 MATERIALS SCIENCE↗

An independent analysis of bias sources and variability in wind plant pre-construction energy yield estimation methods

The wind resource assessment community has long had the goal of reducing the bias between wind plant pre-construction energy yield assessment (EYA) and the observed annual energy production (AEP). This comparison is typically made between the 50% probability of exceedance (P50) value of the EYA and the long-term corrected operational AEP (hereafter OA P50), and is known as the P50 bias. The industry has critically lacked an independent analysis of bias reduction investigated across multiple consultants to identify the greatest sources of uncertainty and variance in the EYA process and the best opportunities for uncertainty reduction. The present study addresses this gap by benchmarking consultant methodologies against each other and against operational data at a scale not seen before in industry collaborations. We consider data from 10 wind plants and evaluate discrepancies between eight consultancies in the steps taken from estimates of gross to net energy. Consultants tend to overestimate the gross energy produced at the turbines and then compensate by further overestimating downstream losses, leading to a mean P50 bias near zero, still with significant variability among the individual wind plants. Within our data sample, we find that consultant estimates of all loss categories, except environmental losses, tend to reduce the project-to-project variability of the P50 bias. The disagreement between consultants, however, remains flat throughout the addition of losses. Finally, we find that differences in consultants’ estimates of project performance can lead to differences up to $10/MWh in the levelized cost of energy for a wind plant.

Todd, Austin C.↗

Development and Preliminary Analysis of a U.S. Geothermal Heat Pump Installation Database

This paper seeks to addresses the significant gap in the literature regarding the installation and adoption of geothermal heat pump (GHP) systems in the United States. While the "2021 U.S. Geothermal Power Production and District Heating Market Report" published by the National Renewable Energy Laboratory (NREL) focused on direct-use geothermal district heating systems, it did not include an analysis of GHP installations (Robins et al. 2021). To bridge this gap, NREL has compiled a novel database currently containing 70,470 records of GHP installations, primarily sourced from state well permits and small-scale studies. Our methodology emphasizes the collection, cleaning, and standardization of data, addressing challenges such as inconsistent reporting formats and privacy concerns. Despite limitations in data on capacity, costs, and performance, our preliminary geospatial analysis reveals insights into the distribution of GHP systems across urban and rural areas and climate zones. The paper highlights the importance of publicly accessible data for advancing GHP technology adoption with a discussion of existing data sources and their limitations, advocating for improved collaboration between NREL and industry stakeholders.

data collection↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗