Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Deliverable 6.7-Final Technical Report: Development Summary and Evaluation of the Solar Uncertainty Integrator (SUNI) Software

The Data Quality and Uncertainty Integration Project was a three-year effort to address stakeholder needs for assessing solar radiation resource data quality based on existing tools for estimating radiometer measurement uncertainties and assessing post-measurement data quality. The annual research objectives for the project addressed a logical progression of effort needed to achieve the ultimate project goal of developing the Solar Uncertainty Integrator (SUNI) software. This final technical report summarizes the development process for achieving these key research objectives and addresses the outreach and code development efforts in the final year of the project to develop a new solar irradiance data uncertainty integration software package.

14 SOLAR ENERGY↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

FC Site 4.0 - NLR Scanning Lidar (Halo XR+ 235) / Standardized and Quality-Controlled Data

This dataset contains lidar data that have been standardized and quality-controlled through NatLabRockies/FIEXTA/LiDARGO (https://github.com/NatLabRockies/FIEXTA/tree/main/lidargo). Standardization rearranges the lidar data into convenient range vs beamID vs scanID coordinates that facilitate data analysis. The scan geometry (i.e., azimuth, elevation) is shifted on a regular grid based on the most likely angles within the scan file. Quality control of radial wind speed is performed through a generalized version of the dynamic lidar filter (Beck and Kuhn, 2017).

17 WIND ENERGY↗

FC Site 4.2 - NLR Scanning Lidar (Halo XR+ 199) / Standardized and Quality-Controlled Data

This dataset contains lidar data that have been standardized and quality-controlled through NatLabRockies/FIEXTA/LiDARGO (https://github.com/NatLabRockies/FIEXTA/tree/main/lidargo). Standardization rearranges the lidar data into convenient range vs beamID vs scanID coordinates that facilitate data analysis. The scan geometry (i.e., azimuth, elevation) is shifted on a regular grid based on the most likely angles within the scan file. Quality control of radial wind speed is performed through a generalized version of the dynamic lidar filter (Beck and Kuhn, 2017).

17 WIND ENERGY↗

FC Site 1.9 - NLR Scanning Lidar (Halo XR+ 200) / Standardized and Quality-Controlled Data

This dataset contains lidar data that have been standardized and quality-controlled through NatLabRockies/FIEXTA/LiDARGO (https://github.com/NatLabRockies/FIEXTA/tree/main/lidargo). Standardization rearranges the lidar data into convenient range vs beamID vs scanID coordinates that facilitate data analysis. The scan geometry (i.e., azimuth, elevation) is shifted on a regular grid based on the most likely angles within the scan file. Quality control of radial wind speed is performed through a generalized version of the dynamic lidar filter (Beck and Kuhn, 2017).

17 WIND ENERGY↗

Anomaly Detection in Seismic Data with Deep Learning: Application for Instrument Failure Detection and Forecasting

Seismic data quality assessment (QA) is the first and one of the most important steps before conducting any further data analysis. Traditional methods involve checking various metrics, such as spike detection and power spectral density, by setting strict thresholds or comparing data against synthetic benchmarks. However, these approaches often rely on pre-existing knowledge and assumptions about data anomalies, leading to potential misclassification of unusual cases. Here, in this study, we propose a deep autoencoder model, an unsupervised learning approach that evaluates data quality without making assumptions about normal and anomalous data, which can be used to identify deviations in recorded data that may indicate nascent instrument failure. We test the model with the U.S. International Monitoring System (IMS) seismic stations and demonstrate the capability of detecting anomalies on a monthly scale. This could prompt station operators to examine potential problems early, allowing sufficient time for instrument maintenance to prevent data outages. In addition, we use a new manually selected testing dataset to compare our model performance against two supervised machine learning (ML) approaches and a standard QA package, as baseline models. When applied to the dataset containing known data anomalies, performance of the supervised and unsupervised ML approaches is similar, with an accuracy of 88.1% for our model compared to ∼90% for the supervised ML approach and 78.2% for the standard QA package. Our model outperforms the baseline models when applied to new stations, where new types of data anomalies can be station-specific and not included in the training dataset. Finally, we show model transferability by training the model with data from the Global Seismograph Network only and applying it to the IMS network data. The results suggest that our model is generalizable and can be applied to new stations with good accuracy.

Lin, Jiun-Ting [Lawrence Livermore National Labora↗

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES↗

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES↗

Livewire: Automatic Annotations

Diogenes processes datasets to provide data quality metrics for the Livewire platform and creates standardized data dictionaries from data annotations. Diogenes needs data annotations that clearly outline thenformat and organization of the data. It also relies on the type, class, and unit of each data piece for comprehensive analysis, which it cannot determine independently. The Annotation Tool significantly reduces the time needed to create annotations for Diogenes by generating data annotations with the correct formatting and content. It also employs machine learning and hard-coded models to automatically annotate data class, quality type, and data units.

33 - ADVANCED PROPULSION SYSTEMS↗

Practical procedures for sensor quality assessment

Sensors are increasingly deployed for process monitoring and control. These produce on-line measurements at a high frequency, in parallel with low-frequency laboratory measurements. Compared to laboratory practices, sensor data quality assessment and control practices are far less structured at most utilities. This leads to inaccurate sensor data with unknown uncertainty factors.This chapter shows how to establish standard operating procedures (SOPs) to support sensor data quality assessment and control and subsequent maintenance actions by producing relevant sensor metadata. Furthermore, SOPs are provided for the most commonly used wastewater quality sensors, inspired by utility and academic best practices. This chapter builds on definitions provided in Chapter 3 and provides additional definitions specifically related to sensors maintenance. Chapter 6 complements the methods in this chapter, which are based on reference measurements, with data-analytical techniques.

Alferes, Janelcy↗

Towards Autonomous Experiments by Connecting High Performance Microscopy with High Performance Computing

The digitization of controls, data, and analysis in microscopy is bringing the idea of autonomous microscopes closer to reality than ever before. Automated transmission electron microscopy (TEM) is already fairly routine for some experiments the only require simple repetitive tasks such as imaging biological macromolecules for single particle cryoEM [1], tilt series for electron tomography [2], and movies for crystallography [3]. The vast majority of TEM experiments are conducted completely by human operators who choose the regions of interest, optimize experimental parameters, and make decisions about data quality visually during an experiment. The field is still a long way from having completely autonomous TEMs that can adapt to sample difficulties and tune experimental parameters based on data quality and desired experimental outcomes. Part of the issue is the lack of capability for feeding information learned from on-line, live data analysis back into the on-going experiment [4]. Furthermore, this presentation will discuss current capabilities for large scale data reduction and analysis using high performance computing (i.e. supercomputing) and progress towards developing a true feed-back loop that places data analysis and theory in the experimental loop.

97 MATHEMATICS AND COMPUTING↗

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi↗

High Yield Xray Imager Final Design Review

The High Yield Xray Imager (HYXI) is a new NIF target diagnostic system currently under development. The goal of HYXI is to provide high-fidelity, high temporal resolution x-ray imaging capability on high yield NIF implosions at 10MJ and above. The HYXI instrument design concept is based on the combination of two technologies that have been successfully utilized at the NIF on previous instruments, electron pulse-dilation and hybrid-CMOS sensor imaging. The combination of these two techniques will give HYXI sufficient data quality to ascertain differences in hot spot formation dynamics between high and low yield implosions. This information will highlight the critical hot spot conditions needed for ignition and burn. The HYXI design leverages the successful operation of the PDIXI x-ray imager at the NIF on multi MJ yield shots. A new radiation tolerant CMOS imaging array (HYPERION) is being developed to eliminate the significant background noise which limits the data quality of PDIXI. We successfully placed the contract with Advanced hCMOS Systems (AHS) to develop the HYPERION sensor, which fulfils our criteria to place long lead time item procurements by end of FY24. The HYXI Final Design Review was completed at the end of Q4 FY24 (Sep 24 th and Sep 30 th ). The HYXI project is a multi-year effort with a phased approach to be bring up system functionality over time in parallel with the development and fabrication effort of the HYPERION CMOS imaging array. In Phase 1, time-integrated x-ray images on NIF DT experiments will be collected starting in Q3 FY25. In Phase 2 of the project, time-resolved imaging with HYXI utilizing a spare microchannel plate detector back-end will begin in Q3 FY26. Phase 3 concludes the project with the installation of the HYPERION sensor array and the final performance qualification of the HYXI instrument which is scheduled for Q3 FY27 as discussed in the PDR and MRT report on this project in FY23.

42 ENGINEERING↗

Reward based optimization of resonance-enhanced piezoresponse spectroscopy

Dynamic spectroscopies in scanning probe microscopy (SPM) are critical for probing material properties, such as force interactions, mechanical properties, polarization switching, electrochemical reactions, and ionic dynamics. However, the practical implementation of these measurements is constrained by the need to balance imaging time and data quality. Signal to noise requirements favor long acquisition times and high frequencies to improve signal fidelity. However, these are limited on the low end by contact resonant frequency and photodiode sensitivity and on the high end by the time needed to acquire high-resolution spectra or the propensity for sample degradation under high field excitation over long times. The interdependence of key parameters such as instrument settings, acquisition times, and sampling rates makes manual tuning labor-intensive and highly dependent on user expertise, often yielding operator-dependent results. These limitations are prominent in techniques like dual amplitude resonance tracking in piezoresponse force microscopy that utilize multiple concurrent feedback loops for topography and resonance frequency tracking. Here, a reward-driven workflow is proposed that automates the tuning process, adapting experimental conditions in real time to optimize data quality. Furthermore, this approach significantly reduces the complexity and time required for manual adjustments and can be extended to other SPM spectroscopic methods, enhancing overall efficiency and reproducibility.

47 OTHER INSTRUMENTATION↗

Sampling and Analysis Plan for LSL2 Underground Storage Tank Evaluation of PFAS in Soil: Protection of Groundwater

This SAP gathers all pertinent information relevant to the sampling and analysis of soil at the removal site for a former 10,000-gallon UST and reference sites in the vicinity of the excavation site. Data Quality Objectives establishes the process and steps for the acquisition of data. Data Flow Diagram establishes the decisions and next steps. Quality Assurance Project Plan establishes the quality requirements for data collection, including planning, implementation, and assessment of sampling, field measurements, and laboratory analysis. The work is being completed for DOE-SC Pacific Northwest Site Office.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

The Dark Energy Bedrock All-sky Supernova Program: Motivation, Design, Implementation, and Preliminary Data Release

Precise measurements of Type Ia supernovae (SNe Ia) at low redshifts (z) serve as one of the most viable keys to unlocking our understanding of cosmic expansion, isotropy, and growth of structure. The Dark Energy Bedrock All-Sky Supernovae (DEBASS) program will deliver a uniformly calibrated low-z dataset of more than 400 spectroscopically confirmed SNe Ia in the Southern Hemisphere. DEBASS utilizes the Dark Energy Camera to image supernovae in conjunction with the Wide-Field Spectrograph to gather comprehensive host-galaxy information. By using the same photometric instrument as both the Dark Energy Survey (DES) and the DECam Local Volume Exploration Survey, DEBASS not only benefits from a robust photometric pipeline and well-calibrated images across the Southern sky, but can replace the historic and external low-z samples that were used in the final DES supernova analysis. In this paper, along with a companion paper, we present an early data release of 77 DEBASS SNe within the DES footprint. We introduce the DEBASS program, discuss its scientific goals and the advantages it offers for supernova cosmology, and present our initial results demonstrating data quality. With this early data release, we find a robust median absolute standard deviation of Hubble diagram residuals of ∼0.10 mag and an initial measurement of the host-galaxy mass step of 0.06 ± 0.04 mag, both before performing bias corrections. This low scatter shows the promise of a low-z SN Ia program with a well-calibrated telescope and high signal-to-noise ratio across multiple bands.

Sherman, Nora F. [Boston U.] (ORCID:00000001539901↗

ARM Lead Mentor Selection Process

The Atmospheric Radiation Measurement (ARM) Program was created in 1989 with funding from the U.S. Department of Energy (DOE) to develop several highly instrumented ground stations to study cloud-formation processes and their influence on radiative transfer. This scientific infrastructure provides for fixed sites, mobile facilities, an aerial facility, and a data archive available for use by scientists worldwide through the ARM Climate Research Facility—a scientific user facility. The ARM Climate Research Facility currently operates more than 300 instrument systems that provide ground-based observations of the atmospheric column. To keep ARM at the forefront of climate observations, the ARM infrastructure depends heavily on instrument scientists and engineers, known as Mentors. Mentors must have an excellent understanding of instrumentation theory and operation for their instrument areas and have comprehensive knowledge of critical scale-dependent atmospheric processes. They must also possess the technical and analytical skills to develop new data retrievals that provide innovative approaches for creating research-quality data sets. The ARM Facility seeks the best overall qualified candidate, or team when appropriate, that can fulfill Mentor requirements in a timely manner. The roles and responsibilities of the ARM Instrument Operations Manager are provided in Appendix A. The key role and responsibilities and detailed responsibilities of ARM Lead Mentors are provided in Appendix B and Appendix C, respectively.

47 OTHER INSTRUMENTATION↗