Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Information Quality Cluster”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ESIP Information Quality Cluster (IQC)

The Information Quality Cluster (IQC) within the Federation of Earth Science Information Partners (ESIP) was initially formed in 2011 and has evolved significantly over time. The current objectives of the IQC are to: 1. Actively evaluate community data quality best practices and standards; 2. Improve capture, description, discovery, and usability of information about data quality in Earth science data products; 3. Ensure producers of data products are aware of standards and best practices for conveying data quality, and data providers distributors intermediaries establish, improve and evolve mechanisms to assist users in discovering and understanding data quality information; and 4. Consistently provide guidance to data managers and stewards on how best to implement data quality standards and best practices to ensure and improve maturity of their data products. The activities of the IQC include: 1. Identification of additional needs for consistently capturing, describing, and conveying quality information through use case studies with broad and diverse applications; 2. Establishing and providing community-wide guidance on roles and responsibilities of key players and stakeholders including users and management; 3. Prototyping of conveying quality information to users in a more consistent, transparent, and digestible manner; 4. Establishing a baseline of standards and best practices for data quality; 5. Evaluating recommendations from NASA's DQWG in a broader context and proposing possible implementations; and 6. Engaging data providers, data managers, and data user communities as resources to improve our standards and best practices. Following the principles of openness of the ESIP Federation, IQC invites all individuals interested in improving capture, description, discovery, and usability of information about data quality in Earth science data products to participate in its activities.

data products↗

Information Quality Cluster and Usability

The Information Quality Cluster (IQC) of the Federation of Earth Science Information Partners (ESIP) has been active since 2014 with membership from multiple organizations including NASA and NOAA. The purpose of this presentation is to foster collaboration between the IQC and the ESIP Usability Cluster. The IQC's activities are motivated partly by the guidelines on information quality from several federal agencies. The agencies developed the guidelines complying with a request in 2002 from the Office of Management and Budget (OMB). The OMB request resulted from a congressional mandate, namely, Section 515 of the Treasury and General Government Appropriations Act for Fiscal Year 2001 (Public Law 106-554; H.R. 5658). NASA's guidelines, for example, emphasize the need for high information quality indicating the various types of public users of information from NASA's missions and programs. The IQC's vision is to become an authoritative and responsive resource of information and guidance to data providers on how best to implement data quality standards and best practices, so that the implementations comply with the various agencies' guidelines, as well as provide users with the best quality of information possible. The IQC interacts with various national and international organizations and encourages collaboration for exchange of information. The IQC considers four aspects of information quality: Scientific Quality, Product Quality, Stewardship Quality and Service Quality. The IQC has considered several use cases to identify issues in capturing, describing, providing access to, and enabling use of information on quality. Several of these use cases point to issues about the usability of information. Collaboration between the IQC and Usability Cluster will be beneficial for arriving at solutions to such issues.

Remote Sensing; Data Systems; Information Quality;↗

Ensuring and Improving Information Quality for Earth Science Data and Products: Role of the ESIP Information Quality Cluster

Quality of products is always of concern to users regardless of the type of products. The focus of this paper is on the quality of Earth science data products. There are four different aspects of quality - scientific, product, stewardship and service. All these aspects taken together constitute Information Quality. With increasing requirement on ensuring and improving information quality, there has been considerable work related to information quality during the last several years. Given this rich background of prior work, the Information Quality Cluster (IQC), established within the Federation of Earth Science Information Partners (ESIP) has been active with membership from multiple organizations. Its objectives and activities, aimed at ensuring and improving information quality for Earth science data and products, are discussed briefly.

Earth Science↗

Ensuring and Improving Information Quality for Earth Science Data and Products Role of the ESIP Information Quality Cluster

Quality of products is always of concern to users regardless of the type of products. The focus of this paper is on the quality of Earth science data products. There are four different aspects of quality scientific, product, stewardship and service. All these aspects taken together constitute Information Quality. With increasing requirement on ensuring and improving information quality, there has been considerable work related to information quality during the last several years. Given this rich background of prior work, the Information Quality Cluster (IQC), established within the Federation of Earth Science Information Partners (ESIP) has been active with membership from multiple organizations. Its objectives and activities, aimed at ensuring and improving information quality for Earth science data and products, are discussed briefly.

Information Quality↗

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality↗

FAIR-ness Assessment of NASA’s Earth Observation System Data and Information System (EOSDIS)

This presentation addresses the challenge of evaluating a multi-disciplinary institutional network of data repositories in operation since 1994 against the relatively recent criteria that constitute FAIR (Findable, Accessible, Interoperable, Reusable) data. NASA’s Earth Observation System Data and Information System (EOSDIS), with its 12 discipline-based Distributed Active Archive Centers (DAACs), preceded the definition and popularization of FAIR by over two decades. An assessment is very useful to describe how well the FAIR principles are met and to identify any improvements needed. In 2020, A “self-assessment” of EOSDIS and DAACs was performed by the ESDIS Project staff and the DAACs from the points of view of human actionability and machine actionability. More recently, a draft of a Science Mission Directorate (SMP) Program Directive (SPD-41a) has been released by NASA Headquarters for comment, where it is recommended that all SMD-funded data should follow the FAIR principles. This presentation is timely to initiate community discussion within the Information Quality Cluster (IQC) of the Earth Science Information Partners (ESIP) and help strategize and develop implementation guidelines for EOSDIS and DAACs to conform to FAIR principles.

Remote sensing↗

Image Information Mining Utilizing Hierarchical Segmentation

The Hierarchical Segmentation (HSEG) algorithm is an approach for producing high quality, hierarchically related image segmentations. The VisiMine image information mining system utilizes clustering and segmentation algorithms for reducing visual information in multispectral images to a manageable size. The project discussed herein seeks to enhance the VisiMine system through incorporating hierarchical segmentations from HSEG into the VisiMine system.

Tilton, James C.↗

LANDSAT-4 MSS and Thematic Mapper data quality and information content analysis

LANDSAT-4 thematic mapper (TM) and multispectral scanner (MSS) data were analyzed to obtain information on data quality and information content. Geometric evaluations were performed to test band-to-band registration accuracy. Thematic mapper overall system resolution was evaluated using scene objects which demonstrated sharp high contrast edge responses. Radiometric evaluation included detector relative calibration, effects of resampling, and coherent noise effects. Information content evaluation was carried out using clustering, principal components, transformed divergence separability measure, and supervised classifiers on test data. A detailed spectral class analysis (multispectral classification) was carried out to compare the information content of the MSS and TM for a large number of scene classes. A temperature-mapping experiment was carried out for a cooling pond to test the quality of thermal-band calibration. Overall TM data quality is very good. The MSS data are noisier than previous LANDSAT results.

Anuta, P.↗

Documentation Resources on the ESIP Wiki

The ESIP community includes data providers and users that communicate with one another through datasets and metadata that describe them. Improving this communication depends on consistent high-quality metadata. The ESIP Documentation Cluster and the wiki play an important central role in facilitating this communication. We will describe and demonstrate sections of the wiki that provide information about metadata concept definitions, metadata recommendation, metadata dialects, and guidance pages. We will also describe and demonstrate the ISO Explorer, a tool that the community is developing to help metadata creators.

ESIP Documentation Cluster↗

SIMBAD quality-control

The astronomical database SIMBAD developed at the Centre de donnees astronomiques de Strasbourg presently contains 760,000 objects (stellar and non-stellar). It has the unique characteristic of being structured specifically for astronomical objects. All types of heterogeneous data (bibliographic references, measurements, and sets of identification) are connected with each object. The attributes that define quality of the database include the following. Reliability: cross-identification should not rely upon just exact values object coordinates. It also means that information attached to one simple object should be consistent. The existing data must be controlled in order to start with a reliable base and to cross-identify new data assuring the quality as data grows. Exhaustivity: delays between publication of new informations and their inclusion in the database should be as short as possible. The integrity of the database has to be maintained as data accumulates. Taking the amount of data into consideration and the rate of new data production, it is necessary to use automatic methods. One of the possibilities is to use multivariate data analysis. The factor-space is a n-dimensional relevancy space which is described by the n-axes representing a set of n subject matter headings; the words and phrases can be used to scale the axes and the documents are then a vector average of the terms within them. The application reported herein is based on the NASA-STI bibliographical database. The selected data concern astronomy, astrophysics, and space radiation (102,963 references from 1975 to 1991 included 8070 keywords). The F-space is built from this bibliographical data. By comparing the F-space position obtained from the NASA-STI keywords with the F-space position obtained from the SIMBAD references, the authors will be able to show whether it is possible to retrieve information with a restricted set of words only. If the comparison is valid, this will be a way to enter bibliographic information in the SIMBAD quality control process. Furthermore, it is possible to connect the physical measurements of stars from SIMBAD to literature concerning these stars from the NASA-STI abstracts. The physical properties of stars (e.g. UBV colors) are not randomly distributed. Stars are distributed among different clusters in a physical parameter space. The authors will show that there are some relations between this classification and the literature concerning these objects clusters in a factor space. They will investigate the nature of the relationship between the SIMBAD measurements and the bibliography. These would be new relationships that are not pre-established by an astronomer. In addition, the bibliography could be neutral information that can be used in combination with the measured parameters.

Lesteven, Soizick↗

Preliminary Evaluation of Thematic Mapper Image Data Quality

Thematic Mapper (TM) data from Mississippi County, Arkansas, and Webster County, Iowa, were examined for the purpose of evaluating the image data quality of the TM which was launched on board the LANDSAT-4 spacecraft. Preliminary clustering and principal component analysis indicates that the middle infrared and thermal infrared data of TM appear to add significant information over that of the near IR and visible bands of the multispectral scanner data. Moreover, the higher spatial resolution of TM appears to provide better definition of the edges and the within variability of agricultural fields. The geometric performance of TM data, without ground control correction, was found to exceed expectations. The modulation transfer function for the 1.65 m band was found to agree with prelaunch specifications when the effects of the GSFC cubic convolution and the atmosphere were removed. The band to band registration for the bands within the noncooled focal plane was found to be better than specified. However, the middle infrared and thermal infrared, which are on a separate cooled focal plane were found to be misregistered and were significantly worse than prelaunch specifications.

Macdonald, R. B.↗

Leveraging ­ CPF Spectral Information for Effective Angular Corrections in Sensor Inter-calibration Studies

The Climate Absolute Radiance and Refractivity Observatory Pathfinder (CPF) mission will provide a high-accuracy (0.3% radiometric uncertainty at k=1) SI-traceable on-orbit calibration reference for sensor intercalibration in reflective solar spectral region. CPF inter-calibration algorithms have been developed to address various sampling differences between CPF and target sensors, particularly in spectral, angular, spatial and temporal domains. CPF provides unique hyper-spectral measurements that allow the utilization of rich spectral information for critical inter-calibration procedures including angular correction and spectral-gap filling. The angular correction scheme of CPF mission has been constructed based on the spectral correlation relationship between CPF observed spectra and the spectra to be measured at the viewing/sun geometry angles of target sensors. The angular correction relationship is both angular and scene dependent. The scene variability issue can impose undesired uncertainties that complicates the correction for those errors. We have demonstrated that hyper-spectral radiance information can be well utilized to identify different scene types associated with CPF observations. Using spectral radiances-based scene clustering analysis can greatly reduce the uncertainty associated with the angular correction algorithm. Additionally, a reliable quality control scheme has been established, utilizing spectral similarity and regression-prediction error analyses, to ensure the successful and efficient implementation of the angular correction method.

Wan Wu↗

Landsat-4 data quality analysis

Landsat-4 satellite Thematic Mapper (TM) and multispectral scanner (MSS) data have been analyzed in order to ascertain data quality and information content. Geometric evaluations have tested band-to-band registration accuracy, and the TM's overall system resolution was evaluated for the case of image objects with high contrast, sharp edge responses. The information content evaluation employed clustering, principal components, and the transformed divergence separability measured on data from Iowa and Chicago, Illinois. The MSS classification analysis compared MSS and TM information contents for a large number of science classes.

Anuta, P.↗

Failure Mode Identification Through Clustering Analysis

Research has shown that nearly 80% of the costs and problems are created in product development and that cost and quality are essentially designed into products in the conceptual stage. Currently, failure identification procedures (such as FMEA (Failure Modes and Effects Analysis), FMECA (Failure Modes, Effects and Criticality Analysis) and FTA (Fault Tree Analysis)) and design of experiments are being used for quality control and for the detection of potential failure modes during the detail design stage or post-product launch. Though all of these methods have their own advantages, they do not give information as to what are the predominant failures that a designer should focus on while designing a product. This work uses a functional approach to identify failure modes, which hypothesizes that similarities exist between different failure modes based on the functionality of the product/component. In this paper, a statistical clustering procedure is proposed to retrieve information on the set of predominant failures that a function experiences. The various stages of the methodology are illustrated using a hypothetical design example.

Arunajadai, Srikesh G.↗

Evaluation of the radiometric quality of the TM data using clustering and multispectral distance measures

Radiometrically and geometrically corrected TM data from three different geographic locations were examined. Histograms were inspected for each band to determine the dynamic range of the data, the shape of the distributions, and to verify whether empty bins were introduced by the radiometric correction process. The effect of geometric correction on the radiometry of the resampled pixels was determined. The information content between TM and MSS data sets were compared and the TM data were used to map the thermal effluent discharge into a river ecosystem from a nuclear thermal power plant, and application only possible previously only possible through the acquisition of thermal infrared scanner data from aircraft altitudes.

Bartolucci, L. A.↗

Landsat-4 MSS and Thematic Mapper data quality and information content analysis

Landsat-4 Thematic Mapper and Multispectral Scanner data were analyzed to obtain information on data quality and information content. Geometric evaluations were performed to test band-to-band registration accuracy. Thematic Mapper overall system resolution was evaluated using scene objects which demonstrated sharp high contrast edge responses. Radiometric evaluation included detector relative calibration, effects of resampling, and coherent noise effects. Information content evaluation was carried out using clustering, principal components, transformed divergence separability measure, and numerous supervised classifiers on data from Iowa and Illinois. A detailed spectral class analysis (multispectral classification) was carried out on data from the Des Moines, IA area to compare the information content of the MSS and TM for a large number of scene classes.

Anuta, P. E.↗

A study of the feasibility of using sea and wind information from the ERS-1 satellite. Part 1: Wind scatterometer data

The use of scatterometer and altimeter data in wind and wave assimilation, and the benefits this offers for quality assurance and validation of ERS-1 data were examined. Real time use of ERS-1 data was simulated through assimilation of Seasat scatterometer data. The potential for quality assurance and validation is demonstrated by documenting a series of substantial problems with the scatterometer data, which are known but took years to establish, or are new. A data impact study, and an analysis of the performance of ambiguity removal algorithms on real and simulated data were conducted. The impact of the data on analyses and forecasts is large in the Southern Hemisphere, generally small in the Northern Hemisphere, and occasionally large in the Tropics. Tests with simulated data give more optimistic results than tests with real data. Errors in ambiguity removal results occur in clusters. The probabilities which can be calculated for the ambiguous wind directions on ERS-1 contain more information than is given by a simple ranking of the directions.

Anderson, D.↗