Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]

Data Quality Assessment Process for Real-Time Data-Driven Traffic Microsimulation of Smart Corridor

Smart corridor digital twins are often created for the development and evaluation of emerging intelligent transportation systems and Connected and Autonomous Vehicle (CAV) technologies. However, limited guidance exists for data quality assessment for digital twin development. To address this, this paper discusses the data quality assessment utilized to develop data-driven real-time microscopic simulation models, i.e., digital twins, for two separate smart corridors: the North Avenue Smart Corridor in Atlanta, GA, and the Martin Luther King Smart Corridor in Chattanooga, Tennessee. This paper provides a summary of the author’s investigations of data requirements and data characteristics for the given smart corridor digital twin development efforts. With a focus on data, this summary includes a description of the data investigation process, key data issues observed, and strategies to address observed issues. Discussion is provided to help expand the lessons from these studies to other digital twin development efforts.

Saroj, Abhilasha [ORNL] (ORCID:0000000191178063)

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS

An Open-Access Repository of Synchrophasor Data Quality Examples: Curation and Example Applications

Synchrophasor measurements are critical in providing wide-area situational awareness to power system operators. However, data artifacts may be introduced due to various issues such as loss of communication, loss of GPS signal, internal clock error, and vendor-specific implementation of phasor estimation algorithms. Tools designed to provide actionable insights from synchrophasor data, hence, must be designed to be robust to these data quality issues. In this work, two years of synchrophasor data sourced from multiple electric utilities in the United States were analyzed to identify examples of data quality problems. These examples were then labeled and published in the Grid Event Signature Library, a publicly available repository of power system measurements hosted by the Oak Ridge National Laboratory. This paper describes the data curation process, and illustrates two application use cases where the dataset can be valuable to the research community. In the first use case, a random forest classifier is trained to distinguish power system disturbance signatures from data anomalies introduced in synchrophasor measurements due to clock errors. The second use case studies the impact of data quality issues on an example synchrophasor application (specifically, event start time determination). The choice of data quality problems investigated is informed by the examples in the repository curated in this work.

24 POWER TRANSMISSION AND DISTRIBUTION

Data Quality Assessment of Optiwatt Vehicle Telematics Data

In October 2024, the Idaho National Laboratory (INL) received data from Optiwatt (Compass Global, Inc.) describing the driving and charging behavior of electric vehicle (EV) drivers. The data shared had been collected from approximately 10,000 vehicles and included vehicle specifications, driving information like odometer readings at the beginning and end of origin-destination pairs (i.e., trips with identification of home for trip start and end for Tesla vehicles), and charging information such as charging energy consumed per charge session and if the charge occurred at home. The vehicle data were provided from 9 EV makes and 18 EV models, with production years ranging from 2012–2024, but more than 9,500 of the vehicles were Tesla EVs. The data includes more than six million trips and more than three million charging events that occurred between June 2023 to Aug 2024 and collected from California and the Eastern United States. The purpose of this report is to review the quality of the data received from Optiwatt and the feedback INL received from Optiwatt after data concerns were shared with them.

33 - ADVANCED PROPULSION SYSTEMS

Review of Technical Photovoltaic Key Performance Indicators and the Importance of Data Quality Routines

Technical key performance indicators (KPIs) are important metrics used to assess and quantitatively summarize various aspects of photovoltaic (PV) systems, including long-term performance, economic viability, and carbon footprint. Herein, a group of experts of the International Energy Agency's Photovoltaic Power Systems Programme Task 13 collect and describ the most important technical KPIs used in the industry. Thereby, a set of best practices for reliably handling PV system data is presented and the impact of data quality and climatic variability on KPI calculation is investigated. Further, the effective use of technical KPIs allows triggering data-driven and informed decisions to optimize PV systems and providing a comprehensive overview of how PV systems operate across different conditions and climates. With the worldwide growth of the PV industry, more companies operate/own PV systems in different regions, where the climatic and seasonal profiles differ. This requires context-aware evaluation of KPIs, or the judicious application of multiple KPIs, to ensure that each asset is evaluated correctly. Beyond that, there is untapped potential in the utilization of KPIs through geospatial mapping and extrapolation of fleet KPIs. This study demonstrates that the uncertainty in KPI estimation is not well understood and depends on data quality, climatic variability, and system configuration.

14 SOLAR ENERGY

Anomaly Detection Based on Machine Learning for the CMS Electromagnetic Calorimeter Online Data Quality Monitoring

Using a semi-supervised machine learning approach we present a real-time anomaly detection system based on an autoencoder used for online data quality monitoring of the CMS electromagnetic calorimeter operating at the CERN LHC. We introduce a novel method that maximizes the anomaly detection performance making use of the time-dependence of anomalies and the spatial variations in the detector response. The autoencoder-based system efficiently detects anomalies in real time and maintains a very low false discovery rate. We validate the performance of this novel system with anomalies from LHC collision data taken in 2018 and 2022. In addition, results are presented after deploying the autoencoder-based system in the CMS online Data Quality Monitoring workflow at the beginning of LHC Run 3 resulting in the system to detect issues that were missed by the existing system.

Harilal, Abhirami [Carnegie Mellon University, Pit

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Online and Offline Data Quality Monitoring for the Mu2e Calorimeter

This thesis presents the design, implementation, and validation of a calorimeter Data Quality Monitoring (DQM) toolchain for the Mu2e experiment at Fermilab. Mu2e searches for charged lepton flavor violation via coherent muon-to-electron conversion in the field of an aluminum nucleus, $\mu^- Al \rightarrow e^-Al$, a process whose observation would constitute clear evidence of physics beyond the Standard Model. Achieving target sensitivity requires stringent control of detector performance and data integrity during acquisition, as subtle issues in readout configuration, data formatting, or electronics behavior can compromise reconstruction and bias downstream analyzes. To address these challenges, this work develops a multi-layer DQM approach spanning both raw data validation and reconstructed digi-level diagnostics. At the low level, a fragment analysis component performs word- and bit-field decoding of calorimeter readout blocks, enabling sanity checks of the expected structure and producing detailed error and integrity statistics useful for commissioning and troubleshooting. At the digi level, the CaloDigiDQM analyzer is implemented within the art framework and transforms each CaloDigiCollection into a structured hierarchy of ROOT histograms designed for fast drill-down diagnostics. The module generates coherent monitoring views at global, disk, board, and channel granularity, including occupancy, waveform-derived features (baseline, RMS, peak amplitude and position), and left-right sensor consistency metrics. Detector-aware channel-to-electronics mapping is performed through the conditions system (CaloDAQMap), ensuring that diagnostics remain aligned with hardware identifiers used in operations. For end-to-end testing without reliance on live DAQ data, a synthetic CaloDigi producer is developed to generate realistic waveforms with controlled noise and pulse shapes. The resulting system supports both offline ROOT-file production and online operation, including optional histogram streaming through otsdaq via ots::HistoSender. This toolchain provides a practical and scalable foundation for calorimeter commissioning and stable data collection, enabling early detection of anomalies and reducing operational risk for Mu2e.

Vakulenko, Mark [Drew U.] (ORCID:0009000276197818)

Hydra: An AI-Based Framework for Interpretable and Portable Data Quality Monitoring

Hydra is an advanced framework designed for training and managing AI models for near real time data quality monitoring at Jefferson Lab. Deployed in all four experimental halls, Hydra has analyzed over 2 million images and has extended its capabilities to offline monitoring and validation. Hydra utilizes computer vision to continually analyze sets of images of monitoring plots generated 24/7 during experiments. Generally, these sets of images are produced at a rate and quantity that is exceedingly difficult for shift crews to effectively monitor. Significant effort has been devoted to enhancing Hydra’s user interface, to ensure that it provides clear, actionable insights for shift workers and other users. Gradient Weighted Class Activation Maps (GradCAM) provide added interpretability, allowing users to visualize important regions of the image for classification. Hydra has been containerized to enable the creation of portable demos and seamless integration with container-based technologies such as Kubernetes and Docker. With the user interface enhancements and containerization, Hydra can be rapidly deployed for new use cases and experiments. This talk will describe the Hydra framework, its user interface and experience, and the challenges inherent in its design and deployment.

Britton, Thomas [Thomas Jefferson National Acceler

Hydra: computer vision for data quality monitoring

Hydra, initially developed for Hall-D in 2019, is a system that utilizes computer vision to perform near real time data quality monitoring. Since then, it has been deployed across all experimental halls at Jefferson Lab, with the CLAS12 collaboration in Hall-B being the first outside of GlueX to fully utilize Hydra. The system comprises back end processes that manage the models, their inferences, and the data flow. Finally, the front-end components, accessible via web pages, allow detector experts and shift crews to view and interact with the system.

47 OTHER INSTRUMENTATION

PMU Data Quality and Sensor Health Monitoring

Phasor Measurement Units (PMUs) play a critical role in the evolution of the electric power industry by providing high-precision, real-time monitoring of essential power system metrics. However, effectively detecting abnormalities and critical events from PMU data is a complex task, complicated by intricate temporal patterns, a scarcity of labeled data for training algo- rithms, and constraints on online computational power. In this study, we apply TranAD, an innovative algorithm that combines transformer architectures with the refinement of adversarial learning, to both synthetic and real-world PMU datasets for developing a data quality and sensor online health monitoring platform for utilities. Our findings reveal that TranAD not only provides efficient detection and localization but also enhances the detail with which abnormalities are detected, marking a a significant step forward in the field of clean data acquisition processes for power system monitoring

deep neural network, machine learning (ML)

Data Quality Monitoring for the Hadron Calorimeters Using Transfer Learning for Anomaly Detection

The proliferation of sensors brings an immense volume of spatio-temporal (ST) data in many domains, including monitoring, diagnostics, and prognostics applications. Data curation is a time-consuming process for a large volume of data, making it challenging and expensive to deploy data analytics platforms in new environments. Transfer learning (TL) mechanisms promise to mitigate data sparsity and model complexity by utilizing pre-trained models for a new task. Despite the triumph of TL in fields like computer vision and natural language processing, efforts on complex ST models for anomaly detection (AD) applications are limited. In this study, we present the potential of TL within the context of high-dimensional ST AD with a hybrid autoencoder architecture, incorporating convolutional, graph, and recurrent neural networks. Motivated by the need for improved model accuracy and robustness, particularly in scenarios with limited training data on systems with thousands of sensors, this research investigates the transferability of models trained on different sections of the Hadron Calorimeter of the Compact Muon Solenoid experiment at CERN. The key contributions of the study include exploring TL’s potential and limitations within the context of encoder and decoder networks, revealing insights into model initialization and training configurations that enhance performance while substantially reducing trainable parameters and mitigating data contamination effects.

47 OTHER INSTRUMENTATION

Surface Water Quality Data from Beaver-Impacted Streams; Trail Creek and East River, Colorado 2025

This data package contains surface water chemistry measurements collected in 2025 to evaluate how beaver damming and low-tech process-based stream restoration influence water quality and metal mobility in mountainous headwater systems of the Upper Colorado River Basin. Sampling was conducted at Trail Creek (Taylor Park watershed, Colorado), a tributary undergoing restoration through installation of low-tech process-based structures (i.e., beaver dam analogs), and at off-channel beaver ponds within the East River floodplain (East River watershed, Colorado). Samples were collected along longitudinal transects spanning upstream control reaches, beaver-influenced ponded reaches, and downstream segments. Additional samples were collected from near-surface pore waters within a beaver dam seepage face. The dataset includes concentrations of major and trace elements measured by inductively coupled plasma–mass spectrometry (ICP-MS) and inductively coupled plasma–optical emission spectrometry (ICP-OES), major anions measured by ion chromatography (IC), and dissolved organic carbon (DOC; reported as non-purgeable organic carbon, NPOC). Samples were size-fractionated at 0.45 micrometers (µm), 0.22 µm, and 0.02 µm to distinguish particulate (>0.45 µm), colloidal (0.22–0.02 µm), and dissolved (<0.02 µm) fractions. The data package consists of comma-separated value (.csv) files containing tabulated chemical concentration data, sample metadata (site identifiers, geographic coordinates, sampling dates, fraction type), and quality control flags. All files are provided in open, non-proprietary formats that can be accessed using standard data analysis software such as Microsoft Excel, R, Python, MATLAB, or other programs capable of reading .csv files. Units, detection limits, and analytical methods are documented in accompanying metadata files. The dataset is designed to support analyses of (1) how beaver impoundment and restoration structures alter elemental partitioning and transport, (2) the role of iron and organic carbon in mediating trace metal mobility, and (3) reach-scale changes in water quality across restoration gradients. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

Anions

Water quality data collected from the Muskegon River (MI, USA) using AquaBOT June-July 2024.

This dataset contains processed output from AquaBOT, which combines 30-second interval measurements of GPS location and YSI water quality parameters (temperature (degrees Celsius), dissolved oxygen (mg/l), specific conductance(microSiemens/cm at 25 degrees Celsius), and turbidity (NTU)), for 3 sections of the Muskegon River in Michigan, USA. The file aquabot2024_MuskegonRiver.csv has information on locations, dates and times, and observations. The data were processed by removing observations recorded before and after AquaBOT was in the water and observations where no data values were recorded. No other QA/QC was done. This research was performed as part of the DOE Research Development and Partnership Pilot award “Expanding Collaborative Capacity to Address Climate Resiliency in the Great Lakes Region”, which aims to expand collaborations between researchers at Central Michigan University and U.S. Department of Energy labs and projects focused on enhancing climate resilience in Great Lakes communities and ecosystems.

54 ENVIRONMENTAL SCIENCES

CROCUS Air Quality Data at Argonne National Laboratory Prairie Site

The AQT (Vaisala AQT530) instrument provides observations on meteorological conditions, including particulate matter (PM2.5, PM10), gas species concentrations (NO, NO2, O3, CO), and environment temperature and moisture. These measurements are critical for understanding air quality. These measurements are useful for understanding changes in aerosol properties, air quality research, and comparing to model experiments especially in urban environments. These measurements are collected at the Argonne Testbed for Multiscale Observational Science (ATMOS), a prairie field site at Argonne National Laboratory in Lemont, Illinois. Data is available in the netCDF data format, we encourage data users review documentation through Project Pythia to understand how to work with netCDF data https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html. Each file contains one day's worth of data (24 hours, starting at 0000 UTC). The data is aggregated into daily frequency to make it easier to process multiple days, and compress the higher-resolution fields. File naming convention includes the project (CROCUS), location (atmos), data level (raw, a1), date (year, month, day), and hour (0000).

54 ENVIRONMENTAL SCIENCES

CROCUS Air Quality Data at University of Illinois - Chicago Tower

The AQT (Vaisala AQT530) instrument provides observations on meteorological conditions, including particulate matter (PM2.5, PM10), gas species concentrations (NO, NO2, O3, CO), and environment temperature and moisture. These measurements are critical for understanding air quality. These measurements are useful for understanding changes in aerosol properties, air quality research, and comparing to model experiments especially in urban environments. These measurements are collected at the University of Illinois in Chicago, Illinois, on the meteorological tower near the greenhouse on campus. Data is available in the netCDF data format, we encourage data users review documentation through Project Pythia to understand how to work with netCDF data https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html. Each file contains one day's worth of data (24 hours, starting at 0000 UTC). File naming convention includes the project (CROCUS), location (UIC), data level (raw, a1), date (year, month, day), and hour (0000).

54 ENVIRONMENTAL SCIENCES

Sensor Data Analytics and Data Quality Assessment Software

The proposed framework derives a set of quality metrics to provide critical insights into and tracking of grid operations, sensor performance, sensor longevity, and event statistics. Power grid engineers can utilize this information to identify problems with existing sensor locations and problematic power grid assets including generators, transmission lines, load centers, and substations. This information can also be used to identify unexpected/abnormal behavior of power grid components, improve power grid observability, and operational monitoring, and thus enhance real-time decision-making support system. Power grid planners can utilize this information to augment existing sensing architecture with new sensors and improve the observability of the network.

Mahapatra, Kaveri