Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Hydropower Infrastructure - LAkes, Reservoirs, and RIvers (HILARRI), v4

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2025) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2025) – Power plants that are listed in the 2025 U.S. Hydropower Development Pipeline Data or were listed in previous versions of the dataset These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) – EPA SuRGE sampling locations Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

Hansen, Carly [ORNL] (ORCID:0000000193280838)↗

Accelerating Control Systems with GitOps: A Path to Automation and Reliability

GitOps is a foundational approach for modernizing infrastructure by leveraging Git as the single source of truth for declarative configurations. The poster explores how GitOps transforms traditional control system infrastructure, services and applications by enabling fully automated, auditable, and version-controlled infrastructure management. Cloud-native and containerized environments are shifting the ecosystem not only in the IT industry but also within the computational science field, as is the case of CERN and Diamond Light Source among other Accelerator/Science facilities which are slowly shifting towards modern software and infrastructure paradigms. The ACORN project, which aims to modernize Fermilab’s control system infrastructure and software is implementing proven best-practices and cutting-edge technology standards including GitOps, containerization, infrastructure as code and modern data pipelines for control system data acquisition and the inclusion of AI/ML in our accelerator complex.

Gonzalez, M. [Fermilab]↗

TESS Science Processing Operations Center Pipeline and Data Products

TESS launched 18-4-2018 to conduct a two-year, near all-sky survey for at least 50 nearby exoplanets for which masses can be. TESS just completed surveying the southern hemisphere, identifying hundreds of candidate exoplanet systems and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) at NASA Ames Research Center processes the image data downlinked from TESS every two weeks to generate a variety of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each ~1 month sector, the SPOC calibrates the image data for both 30-min Full Frame Images (FFIs) and up to 20,000 pre-selected 2-min target star postage stamps. Simple aperture photometry and systematic error-corrected flux time series are generated for the 2-min data. The data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions (PRF). The archival files are modeled after Kepler's for ease of use, and include Target Pixel Files (TPFs) containing original and calibrated 2-min image data, Light Curve files (LCs) containing the photometric time series for each 2-min target, as well as the Data Validation products. New products derived from the FFIs include light curves for the 2-min targets and CBVs. The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Jenkins, Jon M.↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

Dynamic Black-Level Correction and Artifact Flagging for Kepler Pixel Time Series

Methods applied to the calibration stage of Kepler pipeline data processing [1] (CAL) do not currently use all of the information available to identify and correct several instrument-induced artifacts. These include time-varying crosstalk from the fine guidance sensor (FGS) clock signals, and manifestations of drifting moire pattern as locally correlated nonstationary noise, and rolling bands in the images which find their way into the time series [2], [3]. As the Kepler Mission continues to improve the fidelity of its science data products, we are evaluating the benefits of adding pipeline steps to more completely model and dynamically correct the FGS crosstalk, then use the residuals from these model fits to detect and flag spatial regions and time intervals of strong time-varying black-level which may complicate later processing or lead to misinterpretation of instrument behavior as stellar activity.

Kolodziejczak, J. J.↗

Data about data – when, why and how metadata can support the digital plant

A structured approach for recording data quality and contextual information about how and why a signal exists – i.e. metadata – is central to interpret and use sensor data correctly. This is becoming increasingly important with the global trend with data-driven applications such as digital twins and AI-models. But a structured metadata collection and organization of sensor data is not routine in most plants, which can result in lost information and missed opportunities to make use of the investments made in the data collection. Therefore, the IWA task group on Metadata Collection and Organization in wastewater resource recovery systems (MetaCO) was initiated in 2020 and recently delivered the IWA scientific and technical report number 31. The report gives and in-depth description about metadata in water resources recovery facilities (WRRFs) and is available as open access at IWA publishing. The report is the outcome of the collaboration between more than 80 water professionals with the intention to serve WRRF data users with a guide on how to structure and make use of metadata throughout the data pipeline in order to maximize the value of sensor data.

Alferes, Janelcy [VITO, Belgium]↗

Principal Components as a Data Reduction and Noise Reduction Technique

The potential of principal components as a pipeline data reduction technique for thematic mapper data was assessed and principal components analysis and its transformation as a noise reduction technique was examined. Two primary factors were considered: (1) how might data reduction and noise reduction using the principal components transformation affect the extraction of accurate spectral classifications; and (2) what are the real savings in terms of computer processing and storage costs of using reduced data over the full 7-band TM complement. An area in central Pennsylvania was chosen for a study area. The image data for the project were collected using the Earth Resources Laboratory's thematic mapper simulator (TMS) instrument.

Imhoff, M. L.↗

Product Operations Status Summary Metrics

The Product Operations Status Summary Metrics (POSSUM) computer program provides a readable view into the state of the Phoenix Operations Product Generation Subsystem (OPGS) data pipeline. POSSUM provides a user interface that can search the data store, collect product metadata, and display the results in an easily-readable layout. It was designed with flexibility in mind for support in future missions. Flexibility over various data store hierarchies is provided through the disk-searching facilities of Marsviewer. This is a proven program that has been in operational use since the first day of the Phoenix mission.

Takagi, Atsuya↗

Development of the Science Data System for the International Space Station Cold Atom Lab

Cold Atom Laboratory (CAL) is a facility that will enable scientists to study ultra-cold quantum gases in a microgravity environment on the International Space Station (ISS) beginning in 2016. The primary science data for each experiment consists of two images taken in quick succession. The first image is of the trapped cold atoms and the second image is of the background. The two images are subtracted to obtain optical density. These raw Level 0 atom and background images are processed into the Level 1 optical density data product, and then into the Level 2 data products: atom number, Magneto-Optical Trap (MOT) lifetime, magnetic chip-trap atom lifetime, and condensate fraction. These products can also be used as diagnostics of the instrument health. With experiments being conducted for 8 hours every day, the amount of data being generated poses many technical challenges, such as downlinking and managing the required data volume. A parallel processing design is described, implemented, and benchmarked. In addition to optimizing the data pipeline, accuracy and speed in producing the Level 1 and 2 data products is key. Algorithms for feature recognition are explored, facilitating image cropping and accurate atom number calculations.

bose einstein condensate↗

Advanced Infrastructure Integrity Modeling (AIIM) Onshore Pipeline Database

The Advanced Infrastructure Integrity Modeling (AIIM) Onshore Pipeline Database is an interoperable spatial resource containing critical environmental, operational, and reported stressors tied to publicly available oil and gas pipeline locations across the contiguous U.S. and Alaska. This database contains two layers: 1. Pipeline point locations (‘pipeline_points’) – More than 500,000 points (at every kilometer along pipelines, and end points) to which more than 350 stress-related variables have been appended. 2. Merged pipelines (‘merged_pipelines’) – The original, publicly available pipeline data (see table below) merged together into one feature class.

Carbon Transport↗