Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Transit Rider/Travel Behavior Inventory Survey - Minneapolis-St. Paul Metro - 2010

The Metropolitan Council administered a system-wide transit on-board survey between September and November 2010 as part of the 2010 Travel Behavior Inventory to provide detailed transit usage patterns and rider information to support modeling and planning efforts. The study’s focus was to capture only the most current and reliable data necessary to determine future public transportation needs in the Minneapolis region. The study was structured to collect detailed transit ridership data for different routes during different times of day to develop a disaggregate transit trip table that will support advanced travel demand modeling. The study was designed to leverage existing data sources, such as the 2005 on-board survey, to provide the best quality data to update the Travel Behavior Inventory in the Minneapolis-Saint Paul metropolitan area.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Transforming Agricultural Productivity with AI-Driven Forecasting: Innovations in Food Security and Supply Chain Optimization

Global food security is under significant threat from climate change, population growth, and resource scarcity. This review examines how advanced AI-driven forecasting models, including machine learning (ML), deep learning (DL), and time-series forecasting models like SARIMA/ARIMA, are transforming regional agricultural practices and food supply chains. Through the integration of Internet of Things (IoT), remote sensing, and blockchain technologies, these models facilitate the real-time monitoring of crop growth, resource allocation, and market dynamics, enhancing decision making and sustainability. The study adopts a mixed-methods approach, including systematic literature analysis and regional case studies. Highlights include AI-driven yield forecasting in European hydroponic systems and resource optimization in southeast Asian aquaponics, showcasing localized efficiency gains. Furthermore, AI applications in food processing, such as plasma, ozone and Pulsed Electric Field (PEF) treatments, are shown to improve food preservation and reduce spoilage. Key challenges—such as data quality, model scalability, and prediction accuracy—are discussed, particularly in the context of data-poor environments, limiting broader model applicability. The paper concludes by outlining future directions, emphasizing context-specific AI implementations, the need for public–private collaboration, and policy interventions to enhance scalability and adoption in food security contexts.

99 GENERAL AND MISCELLANEOUS↗

Over three decades, and counting, of near-surface turbulent flux measurements from the Atmospheric Radiation Measurement (ARM) user facility

Processes mediating the coupling of terrestrial, aquatic, biospheric, and atmospheric systems influence weather, climate, and ecosystem dynamics via transfer of energy, momentum, water, and carbon (or other species). These exchange processes are quantified by measurements of near-surface turbulent fluxes. Understanding processes at these interfaces provides insight toward understanding and predicting current and future states within the Earth system. The Atmospheric Radiation Measurement (ARM) user facility has been conducting measurements of near-surface turbulent fluxes since the early 1990s at long-term fixed locations and shorter-term mobile deployments across the Earth. ARM has utilized two established methods for conducting these measurements: energy balance Bowen ratio (EBBR) and eddy covariance (EC). Primary measurements from the former include sensible and latent heat flux, while the latter also measures fluxes of momentum and carbon (primarily carbon dioxide, with methane fluxes measured at two locations to date). The EBBR systems have been deployed at 22 locations, and, to date, the EC systems have been deployed at over 50 sites, with plans for additional novel site locations in the future. Herein, the history, evolution, and key aspects of these instrument systems are documented, along with information on data quality assurance and post-processing, as well as best use practices. Additionally, three data validation experiments were recently conducted, and their key findings are summarized. Finally, ancillary datasets acquired by ARM, which can contextualize and aid interpretation of the near-surface turbulent flux measurements, are discussed. The datasets described herein include the eddy correlation flux measurement system: 30ECOR (https://doi.org/10.5439/1879993, Sullivan et al., 1997), 30QCECOR (https://doi.org/10.5439/1097546, Gaustad, 2003), ECORSF (https://doi.org/10.5439/1494128, Sullivan et al., 2019a), and associated AmeriFlux and Methane Value-Added Product, AMCMETHANE (https://doi.org/10.5439/1508268, Billesbach, 2011); the energy balance Bowen ratio system: 30EBBR (https://doi.org/10.5439/1023895, Sullivan et al., 1993) and 30BAEBBR (https://doi.org/10.5439/1027268, Gaustad and Xie, 1993); and the carbon dioxide flux measurement system: CO2FLX (https://doi.org/10.5439/1287574, https://doi.org/10.5439/1287575, https://doi.org/10.5439/1287576, Koontz et al., 2015a, b, c; https://doi.org/10.5439/1989774, https://doi.org/10.5439/1989776, https://doi.org/10.5439/1992202, Biraud and Chan, 2002a, b, c). These data can be found by searching the above data stream names at https://adc.arm.gov/discovery/#/results/ (last access: 8 September 2025).

Sullivan, Ryan C. [Argonne National Laboratory (AN↗

DEPRECATED AI-Batt-OS (Autonomous Identification of Battery Life Models - Open Source) [SWR 21-17]

DEPRECATED. This repository was archived by the owner on Jun 30, 2026. It is now read-only. Open source implementation of some of the methods utilized by AI-Batt, a battery lifetime modeling and analysis toolkit provided by the National Laboratory of the Rockies (NLR). This software demonstrates the use of bi-level optimization and symbolic regression techniques to semi-autonomously identify algebraic models predicting the capacity fade of lithium-ion batteries during calendar aging. Modeling the degradation of batteries is a complex task, due to the difficulty in separating the time-dependent and time-independent factors impacting cell level degradation, across multiple data series with different numbers of measurements and/or data quality. Bi-level optimization enables model parameters to be optimized to either the entire data set or to individual data series, allowing statistical disambiguation of global behaviors (data series independent) and local behaviors (data series dependent). Symbolic regression is used to automatically search for optimal low-dimesional models predicting the variation of locally optimized parameters versus time-independent experimental variables from millions of possible models, resulting in a more accurate and repeatable model identification process than is possible by a manual search. The provided tools also implement cross-validation and bootstrap resampling schemes, empowering statistical model comparison/selection and quantification of model uncertainties. An example script replicates the results from the manuscript "Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning", submitted to ECS. All code is written in MATLAB. Requires the Statistics and Machine Learning Toolbox. Contact Dr. Paul Gasper at Paul.Gasper@nlr.gov for any questions.

Gasper, Paul [National Renewable Energy Lab. (NREL↗

The ARM Precipitation Best Estimate (PrecipBE) Value-Added Product Report

The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s Precipitation Best-Estimate (PrecipBE) Value-Added Product integrates multiple precipitation datastreams, accounting for data quality and instrument limitations, to deliver comprehensive per-precipitation event properties alongside ancillary ARM data set data. PrecipBE bundles all valid surface rainfall samples into artificial intelligence (AI)-ready tabular and time-series formats, reporting bundle means and uncertainty ranges. This per-event structure provides an insightful and easy-to-use resource for researchers analyzing precipitation characteristics.

54 ENVIRONMENTAL SCIENCES↗

Large-momentum effective theory’s asymptotic extrapolation vs the inverse problem

Large-momentum effective theory is a physics-guided systematic expansion to calculate light-cone parton distributions, including collinear (PDFs) and transverse-momentum-dependent ones, at any fixed momentum fraction 𝑥 within a range of [𝑥 min , 𝑥 max ]. It theoretically solves the ill-posed inverse problem that afflicts other theoretical approaches to collinear PDFs, such as short-distance factorizations. Recently, Dutrieux et al. raised practical concerns about whether current or even future lattice data will have sufficient precision in the subasymptotic correlation region to support an error-controlled extrapolation—and if not, whether it becomes an inverse problem where the relevant uncertainties cannot be properly quantified. While we agree that not all current lattice data have the desired precision to qualify for an asymptotic extrapolation, some calculations do, and more are expected in the future. We comment on the analysis and results in Dutrieux et al. and argue that a physics-based systematic extrapolation still provides the most reliable error estimates, even when the data quality is not ideal. In contrast, reframing the long-distance asymptotic extrapolation as a data-driven-only inverse problem with ad hoc mathematical conditioning could lead to unnecessarily conservative errors.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Measured Long-Term Solar Irradiance for Climate Studies

Solar radiation is primarily measured using high-quality radiometers (e.g., pyranometers and pyrheliometers). These instruments need to be calibrated regularly (every two years at a minimum). They are also susceptible to various sources of uncertainties. Therefore, rigorous data quality assessment is required to obtain high-confidence data from these radiometers. This is especially true when the data is used to understand climatic trends and/or extreme weather events. In this study, we attempted to create a continuous and reliable dataset by correcting the underlying data which we believe contains bias due to reference pyranometers swap procedures. These biases can be significant which can be up to two percentage points thereby influencing the interpretation of climate changes and/or extreme events. This study elaborates on the biases and methods to correct those biases.

data integrity↗

Feasibility of Correlation-Aware Inference and Universal Precision Scaling in Bonse–Hart Ultra-Small-Angle Neutron Scattering

Bonse–Hart ultra-small-angle neutron scattering (USANS) provides access to micrometre-scale structure, but useful measurements often require long counting times. In this work, we test whether the expected smoothness of the scattering profile can be exploited to improve data quality at lower counting statistics. We apply a Gaussian-process-based method to Bonse–Hart USANS data and evaluate its performance on pseudo-measurements generated from high-statistics experiments under Poisson statistics. This provides a stringent test of how well the underlying I(Q) profile can be reconstructed when the available counts are substantially reduced. We further show that, in the counting-limited regime, the reconstruction error follows a universal scaling behaviour that differs from the usual independent-counting expectation. At higher counts, the improvement crosses over to a resolution-limited regime set by analyser-angle discretization and rocking-curve width. These results clarify when correlation-aware inference is useful in USANS and provide a practical basis for improving measurement efficiency and beam-time usage.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)↗

A Public Data Set of Auto-Generated Geotagged PV Site Equipment, Generated via Deep Learning

In this research, we present a data set over 100 photovoltaic (PV) sites in TX, which have been automatically geotagged via a fully autonomous deep learning (DL) pipeline. Specifically, locations of inverters, tracker/fixed tilt rows, batteries, and substations are labeled algorithmically. To ensure high data quality, all systems have been reviewed manually and any deep learning errors have been corrected. This public data set, as well as the open-sourced pipeline used to generate it, is valuable for site planning, modelling, and insurance purposes. Given time and resources, we hope to extend the data set to additional states/regions in the US.

14 SOLAR ENERGY↗

Water intensity of photovoltaic module manufacturing at the terawatt scale

As the U.S. ramps photovoltaic (PV) manufacturing to the terawatt scale and emphasizes re-shoring manufacturing, potential regional impacts on the U.S. water supply should be considered, particularly since many PV companies rely almost exclusively on public water supplies for manufacturing. This work surveys the academic literature and PV manufacturer reports to estimate the water intensity of monocrystalline silicon, multicrystalline silicon, and cadmium telluride modules manufactured at the terawatt scale, determining that on average, cadmium telluride manufacturing is less water intensive on a per megawatt scale – this is anticipated to be true for all thin film PV manufacturing. While much lower than the water intensity of thermoelectric (e.g., coal) energy generation, significant issues and gaps with PV manufacturing data quality in academic studies are identified which cause estimates to vary by over 1000x (0.04 – 49 trillion liters/terawatt). Data issues are discussed and the need for accurate accounting of water resources (e.g., via continuous, updated information during PV manufacturing) is highlighted. The opportunity to reconfigure decommissioned thermoelectric sites to PV manufacturing is also explored. Finally, factors that influence PV manufacturing water intensity, from individual manufacturing steps to trends across the PV industry, are examined and water conservation opportunities are presented.

14 SOLAR ENERGY↗

Strong Lensing by Galaxies

Strong gravitational lensing at the galaxy scale is a valuable tool for various applications in astrophysics and cosmology. Some of the primary uses of galaxy-scale lensing are to study elliptical galaxies’ mass structure and evolution, constrain the stellar initial mass function, and measure cosmological parameters. Since the discovery of the first galaxy-scale lens in the 1980s, this field has made significant advancements in data quality and modeling techniques. In this review, we describe the most common methods for modeling lensing observables, especially imaging data, as they are the most accessible and informative source of lensing observables. We then summarize the primary findings from the literature on the astrophysical and cosmological applications of galaxy-scale lenses. We also discuss the current limitations of the data and methodologies and provide an outlook on the expected improvements in both areas in the near future.

79 ASTRONOMY AND ASTROPHYSICS↗

Daily, 30 m Resolution NDSI Data for the East River Watershed, CO for 2000-2020

This dataset contains daily Normalized Difference Snow Index (NDSI) values at 30 m spatial resolution for the East River watershed in Colorado, USA. The temporal range of these data includes water years 2001-2020. These data were created using the Spatial and Temporal Adaptive Reflectance Fusion Model (STARFM). This model fuses low spatial and high temporal resolution data from MODIS (500 m, daily) with high spatial and low temporal resolution data from Landsat (30 m, 16 days) to create a 30m synthetic daily snow product. This product allows for the analysis of historical snow covered area trends in the East River Watershed at fine spatiotemporal resolutions where it was not available previously. This research was performed as a part of the Department of Energy’s Subsurface Biogeochemical Research Program with the primary intent of better understanding the timing and spatial patterns of water delivery to the Critical Zone in mountain watersheds. Each .zip file contains one "water year" of data (October 1 - September 30; i.e., water year 2010 starts October 1, 2010 and ends September 30, 2011). Each zip file contains the following: STARFM daily Normalized Difference Snow Index (NDSI) fusion data files in GeoTiff format with one layer for each day between Landsat data acquisition dates (i.e., for dates of Landsat acquisition, the Landsat image is included for that date). The study area is located in an area of Landsat path overlap, so Landsat dates acquisitions are every 7-9 days. Landsat NDSI files containing the high spatial (30m), low temporal (7-9 days due to Landsat path overlap) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Dates for which no Landsat data were obtained are included as NoData layers. MODIS NDSI files containing the high temporal (daily), low spatial (500m) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Please note the MODIS data were resampled to 30m pixels for input into the STARFM model. The data have a scale factor of 10,000 and a no data value of -32767. The projection of all datasets is WGS 84 (EPSG: 4326), which has a latitude/longitude based degree resolution of 0.0002694946 X 0.0002694946, and approximates to the 30 m spatial resolution mentioned above. The Layer Index files in .csv format. They contain information for each layer in the above GeoTiff files regarding the corresponding date for each layer, the fraction of pixels in the image that contain valid data (missing data is due to either cloud cover or poor data quality; these values are not percent snow cover). Dates of Landsat overpass are indicated in these files. If no Landsat data were able to be obtained due to cloud cover or lack of Landsat Tier 1 data available on Google Earth Engine, this is also noted.

EARTH SCIENCE > CRYOSPHERE > SNOW/ICE↗

Data about data – when, why and how metadata can support the digital plant

A structured approach for recording data quality and contextual information about how and why a signal exists – i.e. metadata – is central to interpret and use sensor data correctly. This is becoming increasingly important with the global trend with data-driven applications such as digital twins and AI-models. But a structured metadata collection and organization of sensor data is not routine in most plants, which can result in lost information and missed opportunities to make use of the investments made in the data collection. Therefore, the IWA task group on Metadata Collection and Organization in wastewater resource recovery systems (MetaCO) was initiated in 2020 and recently delivered the IWA scientific and technical report number 31. The report gives and in-depth description about metadata in water resources recovery facilities (WRRFs) and is available as open access at IWA publishing. The report is the outcome of the collaboration between more than 80 water professionals with the intention to serve WRRF data users with a guide on how to structure and make use of metadata throughout the data pipeline in order to maximize the value of sensor data.

Alferes, Janelcy [VITO, Belgium]↗

A Markov chain Monte Carlo (MCMC) Bayesian inference approach to analyze apparent activation barriers and reaction orders from microreactor data

Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.

42 ENGINEERING↗

Searches for CE ν NS and physics beyond the standard model using Skipper-CCDs at CONNIE

The Coherent Neutrino-Nucleus Interaction Experiment (CONNIE) aims to detect the coherent scattering ( CE ν NS ) of reactor antineutrinos off silicon nuclei using thick fully depleted high-resistivity silicon CCDs. Two Skipper-CCD sensors with subelectron readout noise capability were installed at the experiment next to the Angra-2 reactor in 2021, making CONNIE the first experiment to employ Skipper-CCDs for reactor neutrino detection. We report on the performance of the Skipper-CCDs, the new data processing, data quality, and event selection for CE ν NS interactions, which enable CONNIE to reach a record low detection threshold of 15 eV. The data were collected over 300 days in 2021–2022 and correspond to exposures of 14.9 g-days with the reactor-on and 3.5 g-days with the reactor-off. The difference between the reactor-on and off event rates shows no excess and yields upper limits for the neutrino interaction rates, comparable with previous CONNIE limits from standard CCDs and higher exposures. Searches for new neutrino interactions beyond the Standard Model improve the previous CONNIE limit on a simplified model with light vector mediators. A first dark matter (DM) search by diurnal modulation by CONNIE obtains the best limits on the DM-electron scattering cross section by a surface-level experiment. These promising results, obtained using a very small-mass sensor, illustrate the potential of Skipper-CCDs to probe rare neutrino interactions and motivate the plans to increase the detector mass in the near future.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Indoor Air Quality (IAQ) Monitoring for Space Farming Institute [Slides]

Through the U.S. Department of Energy's Energy to Communities (E2C) program, NREL, other national laboratory experts, and select organizations provide Expert Match - free, short-term technical assistance to address near-term energy challenges and questions. Expert Match is for community stakeholders who have decision-making power or influence in their community but need access to additional energy expertise to inform key upcoming decisions. This Expert Match request supported the Space Farming Institute, a nonprofit organization located in Anchorage, AK, with an indoor air quality analysis. The NREL team analyzed indoor air quality data provided by the Space Farming Institute, which experts at PNNL used to design an indoor bioreactor to grow Ulva algae for indoor air quality mitigation purposes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

On High-Temperature Dynamometer Test Stand Development

RePED is a novel prototype alternator that has been designed and built to operate at 35 kW under 250 degrees C ambient temperature conditions. The operating efficiency will be measured using a 225-kW motor to drive the machine in a back-to-back configuration. The goal is to determine if the RePED alternator can deliver 35 kW at a rotational speed of 1,000 rpm while operating in 250 degrees C ambient temperature and have a minimum efficiency of 60%. Dynamometer testing of a high-temperature alternator for geothermal drilling presents several unique challenges, especially around maintaining and managing the various sources and types of energy passing through the system without impacting data quality. The objectives of this report are to (1) describe the test stand in detail, the relative orientation of each major component, and the data acquisition system and (2) discuss the behavior of the test stand components such that this report can serve as a reference document for future projects.

15 GEOTHERMAL ENERGY↗