Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

MODIS and VIIRS Calibration and Characterization in Support of Producing Long-Term High-Quality Data Products

Terra and Aqua Moderate Resolution Imaging Spectroradiometer (MODIS) have successfully operated since their launches in 1999 and 2002, respectively, and generated various data products to support the Earth remote sensing disciplines and users worldwide for their research activities and applications, including studies of the Earth system, and its changes over time and geographic regions. The MODIS data have also significantly contributed to the continuity of multi-decadal satellite data records and led to major advances in the Earth remote sensing field. The long-term data records from MODIS observations have been and will continue to be extended by the Visible Infrared Imaging Radiometer Suite (VIIRS) instruments, currently operated aboard the Suomi-National Polar-Orbiting Partnership (NPP) and NOAA-20 satellites. The data quality of satellite instruments strongly depends on their calibration accuracy and stability. In order to help scientists and users gain a better understanding of MODIS and VIIRS data quality, this paper provides an overview of their on-orbit calibration methodologies, approaches, and results derived from instrument on-board calibrators and lunar observations, as well as select Earth view targets. What is also discussed is the calibration consistency between MODIS and VIIRS and its potential impact on producing multi-sensor long-term data records. As illustrated, the overall performance of both MODIS and VIIRS continues to meet their design requirements.

MODIS

Data Quality: 1979 Southeastern Virginia Urban Plume Study (SEV-UPS): Surface and Airborne studies

The SEV-UPS field study conducted during August 1979 involved the coordinated air quality data collection efforts of three different organizations operating a total of four aircraft and 12 ground stations. Incorporated into plans for that program were numerous quality assurance activities designed to identify potential hazards for data quality in time for their correction and to take data to allow assessment of the consistency of the data from different stations. The purpose of this study was to document QA procedures that were performed by participating agencies and revise data from all platforms for mutual compatibility. For the latter effort, cases were identified on 21 occasions on seven different days where two or more monitoring systems were operating in the same location at nearly the same time. Data from the systems involved were tabulated in 200 m altitude segments, averaged and compared to detect any bias between measurements made from different systems. The results of these comparisons along with the results of independent performance audits are presented.

White, J. H.

A Signal Detection Theory Approach to Evaluating Oculometer Data Quality

Currently, data quality is described in terms of spatial and temporal accuracy and precision [Holmqvist et al. in press]. While this approach provides precise errors in pixels, or visual angle, often experiments are more concerned with whether subjects'points of gaze can be said to be reliable with respect to experimentally-relevant areas of interest. This paper proposes a method to characterize oculometer data quality using Signal Detection Theory (SDT) [Marcum 1947]. SDT classification results in four cases: Hit (correct report of a signal), Miss (failure to report a ), False Alarm (a signal falsely reported), Correct Reject (absence of a signal correctly reported). A technique is proposed where subjects' are directed to look at points in and outside of an AOI, and the resulting Points of Gaze (POG) are classified as Hits (points known to be internal to an AOI are classified as such), Misses (AOI points are not indicated as such), False Alarms (points external to AOIs are indicated as in the AOI), or Correct Rejects (points external to the AOI are indicated as such). SDT metrics describe performance in terms of discriminability, sensitivity, and specificity. This paper presentation will provide the procedure for conducting this assessment and an example of data collected for AOIs in a simulated flightdeck environment.

Latorella, Kara

Process air quality data

Air quality sampling was conducted. Data for air quality parameters, recorded on written forms, punched cards or magnetic tape, are available for 1972 through 1975. Computer software was developed to (1) calculate several daily statistical measures of location, (2) plot time histories of data or the calculated daily statistics, (3) calculate simple correlation coefficients, and (4) plot scatter diagrams. Computer software was developed for processing air quality data to include time series analysis and goodness of fit tests. Computer software was developed to (1) calculate a larger number of daily statistical measures of location, and a number of daily monthly and yearly measures of location, dispersion, skewness and kurtosis, (2) decompose the extended time series model and (3) perform some goodness of fit tests. The computer program is described, documented and illustrated by examples. Recommendations are made for continuation of the development of research on processing air quality data.

Butler, C. M.

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]

International Metadata Standards and Enterprise Data Quality Metadata Systems

Well-documented data quality is critical in situations where scientists and decision-makers need to combine multiple datasets from different disciplines and collection systems to address scientific questions or difficult decisions. Standardized data quality metadata could be very helpful in these situations. Many efforts at developing data quality standards falter because of the diversity of approaches to measuring and reporting data quality. The one size fits all paradigm does not generally work well in this situation. I will describe these and other capabilities of ISO 19157 with examples of how they are being used to describe data quality across the NASA EOS Enterprise and also compare these approaches with other standards.

data quality

Data Quality Assessment Process for Real-Time Data-Driven Traffic Microsimulation of Smart Corridor

Smart corridor digital twins are often created for the development and evaluation of emerging intelligent transportation systems and Connected and Autonomous Vehicle (CAV) technologies. However, limited guidance exists for data quality assessment for digital twin development. To address this, this paper discusses the data quality assessment utilized to develop data-driven real-time microscopic simulation models, i.e., digital twins, for two separate smart corridors: the North Avenue Smart Corridor in Atlanta, GA, and the Martin Luther King Smart Corridor in Chattanooga, Tennessee. This paper provides a summary of the author’s investigations of data requirements and data characteristics for the given smart corridor digital twin development efforts. With a focus on data, this summary includes a description of the data investigation process, key data issues observed, and strategies to address observed issues. Discussion is provided to help expand the lessons from these studies to other digital twin development efforts.

Saroj, Abhilasha [ORNL] (ORCID:0000000191178063)

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS

Systematic monitoring and evaluation of M7 scanner performance and data quality

An investigation was conducted to provide the information required to maintain data quality of the Michigan M7 Multispectral scanner by systematic checks on specific system performance characteristics. Data processing techniques which use calibration data gathered routinely every mission have been developed to assess current data quality. Significant changes from past data quality are thus identified and attempts made to discover their causes. Procedures for systematic monitoring of scanner data quality are discussed. In the solar reflective region, calculations of Noise Equivalent Change in Radiance on a permission basis are compared to theoretical tape-recorder limits to provide an estimate of overall scanner performance. M7 signal/noise characteristics are examined.

Stewart, S.

Preliminary Evaluation of Thematic Mapper Image Data Quality

Thematic Mapper (TM) data from Mississippi County, Arkansas, and Webster County, Iowa, were examined for the purpose of evaluating the image data quality of the TM which was launched on board the LANDSAT-4 spacecraft. Preliminary clustering and principal component analysis indicates that the middle infrared and thermal infrared data of TM appear to add significant information over that of the near IR and visible bands of the multispectral scanner data. Moreover, the higher spatial resolution of TM appears to provide better definition of the edges and the within variability of agricultural fields. The geometric performance of TM data, without ground control correction, was found to exceed expectations. The modulation transfer function for the 1.65 m band was found to agree with prelaunch specifications when the effects of the GSFC cubic convolution and the atmosphere were removed. The band to band registration for the bands within the noncooled focal plane was found to be better than specified. However, the middle infrared and thermal infrared, which are on a separate cooled focal plane were found to be misregistered and were significantly worse than prelaunch specifications.

Macdonald, R. B.

Evaluation of Landsat-4 Thematic Mapper and multispectral scanner data quality

Landsat-4 image data quality was evaluated for test sites in Iowa and Illinois. Radiometric and geometric quality was tested and an applications evaluation was carried out using a cooling-pond thermal-mapping example. Geometric quality was found to be generally very good. Small errors were found in registration of the middle IR bands of the TM and the thermal IR band was found to be misregistered by one 120-meter pixel. Radiometric quality of the TM is excellent with only minor striping effects.

Bartolucci, L. A.

NASA GLOBE CLOUD GAZE: Creating Data Quality Flags for Citizen Science Cloud Observations Matched to NASA Satellite Data

The GLOBE Program, NASA’s largest and longest lasting citizen science program about the Earth, has been collecting cloud observations matched to multiple satellite data daily. The program’s cloud protocol is historically the most popular protocol as your eyes are the only instruments you need to collect observations of the sky. This dataset includes over 3,300,000 cloud observations with variables like total cloud cover, cloud type and opacity that are collocated to the nearest overpass times of geostationary satellites (GOES-15, GOES-16, GOES-17, Meteosat-8, Meteosat-11, or Himawari-8), or to Clouds and the Earth’s Radiant Energy System (CERES) instruments onboard Aqua and Terra, or the Cloud–Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) satellite. In order to increase the usability of this dataset, the Community science project Leveraging Online and User Data through GLOBE And Zooniverse Engagement (CLOUD GAZE) has been developed to generate data quality flags of these ground-up and top-down perspectives of sky and clouds. Recently funded through NASA’s Citizen Science for Earth Systems Program, CLOUD GAZE has partnered with the Zooniverse online platform to obtain reference data and image tagging of sky photographs collected through The GLOBE Program’s clouds protocol. This paper will present the GLOBE Clouds dataset matched to NASA satellite data, integration of CLOUD GAZE to develop data quality flags, and research applications of the dataset (includes ground-up and top-down perspective comparisons, ground observations of dust storms and smoke plumes, and cloud observations in polar regions). The paper will also present on techniques and recommendations for classroom use and for community engagement, particularly for those looking to online resources.

Marilé Colón Robles

An Open-Access Repository of Synchrophasor Data Quality Examples: Curation and Example Applications

Synchrophasor measurements are critical in providing wide-area situational awareness to power system operators. However, data artifacts may be introduced due to various issues such as loss of communication, loss of GPS signal, internal clock error, and vendor-specific implementation of phasor estimation algorithms. Tools designed to provide actionable insights from synchrophasor data, hence, must be designed to be robust to these data quality issues. In this work, two years of synchrophasor data sourced from multiple electric utilities in the United States were analyzed to identify examples of data quality problems. These examples were then labeled and published in the Grid Event Signature Library, a publicly available repository of power system measurements hosted by the Oak Ridge National Laboratory. This paper describes the data curation process, and illustrates two application use cases where the dataset can be valuable to the research community. In the first use case, a random forest classifier is trained to distinguish power system disturbance signatures from data anomalies introduced in synchrophasor measurements due to clock errors. The second use case studies the impact of data quality issues on an example synchrophasor application (specifically, event start time determination). The choice of data quality problems investigated is informed by the examples in the repository curated in this work.

24 POWER TRANSMISSION AND DISTRIBUTION

Evaluation of Various Radar Data Quality Control Algorithms Based on Accumulated Radar Rainfall Statistics

The primary function of the TRMM Ground Validation (GV) Program is to create GV rainfall products that provide basic validation of satellite-derived precipitation measurements for select primary sites. A fundamental and extremely important step in creating high-quality GV products is radar data quality control. Quality control (QC) processing of TRMM GV radar data is based on some automated procedures, but the current QC algorithm is not fully operational and requires significant human interaction to assure satisfactory results. Moreover, the TRMM GV QC algorithm, even with continuous manual tuning, still can not completely remove all types of spurious echoes. In an attempt to improve the current operational radar data QC procedures of the TRMM GV effort, an intercomparison of several QC algorithms has been conducted. This presentation will demonstrate how various radar data QC algorithms affect accumulated radar rainfall products. In all, six different QC algorithms will be applied to two months of WSR-88D radar data from Melbourne, Florida. Daily, five-day, and monthly accumulated radar rainfall maps will be produced for each quality-controlled data set. The QC algorithms will be evaluated and compared based on their ability to remove spurious echoes without removing significant precipitation. Strengths and weaknesses of each algorithm will be assessed based on, their abilit to mitigate both erroneous additions and reductions in rainfall accumulation from spurious echo contamination and true precipitation removal, respectively. Contamination from individual spurious echo categories will be quantified to further diagnose the abilities of each radar QC algorithm. Finally, a cost-benefit analysis will be conducted to determine if a more automated QC algorithm is a viable alternative to the current, labor-intensive QC algorithm employed by TRMM GV.

Robinson, Michael

LANDAT-4/5 image data quality analysis

LANDSAT-4/5 data quality analysis was covered. Focus was on estimation of two-dimensional point-spread function estimation. A brief description is included.

Anuta, P. E.