Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

U.S. Hydropower Development Pipeline Data, 2026

The U.S. Hydropower Development Pipeline dataset provides a comprehensive, regularly updated view of proposed and potential hydropower projects across the United States. This resource compiles information from federal agencies and other public sources to track non-powered dams considered for electrification, proposed hydropower facilities at stream reaches with no existing dams, conduit exemptions, and emerging pumped storage hydropower proposals. The dataset includes project characteristics such as location, development status, technology type, ownership category, and other attributes that support analysis of future hydropower trends. It is designed to help researchers, planners, policymakers, and stakeholders assess national‑scale development patterns, understand the evolving hydropower landscape, and explore opportunities and challenges associated with new hydropower deployment. The dataset is updated annually to reflect changes in project status, new proposals entering the pipeline, and projects that are cancelled, completed, or otherwise removed from active consideration. Note: Capacity additions to existing hydropower plants are not included in this database due to reliance on a proprietary data source.

Johnson, Megan [ORNL] (ORCID:0000000290141741)

RectifHydPlus Data Pipeline

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab

RectifHydPlus Data Pipeline v1.1.0

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab

POWER DATA PIPELINE

SF-25-081 Utility software for creating high-performance data pipelines to extract, load, and transform raw electric power systems measurements. For use with anomaly detection models training workflows. The software supports the project: Adaptive Cybersecurity for DER: A Game-Theoretic and Machine Learning approach for Real-Time Threat Detection and Mitigation

Plathottam, Silby Jose [Argonne National Laborator

Integrase-On-Demand-Pipeline Data Set

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

McClain, Hannah Marie [Sandia National Laboratorie

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL),

NLR Data Processing Pipeline for MADIS [SWR-26-050]

The NLR Data Processing Pipeline for MADIS software package is for downloading, processing, and performing QA/QC on MADIS data. Designed to handle the following steps: 1) Download all MADIS data as compressed netcdf files for a given time period. 2) Unpack netcdf files into timeseries csvs for each coordinate within the given bounding box. 3) Process the csvs to filter according to quality control checks and convert variables to correct units. 4) Write processed csvs to a single nc file.

Benton, Brandon [National Laboratory of the Rockie

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)

Selection of a Pair of Experiments to Optimally Reduce Uncertainty in Targeted Nuclear Data

We propose a novel process to select a pair of differential and integral experiments that best reduce uncertainties in targeted 239 ⁢Pu nuclear data while compressing the current nuclear data pipeline from 20 to 3 years. 239⁢ Pu nuclear data are poorly understood for neutrons in the intermediate energy range due to sparsity and uncertainty in historical experiments. New experiments targeting this range will enable better understanding of these nuclear data, but choosing the ideal experiments to conduct is challenging. Beginning with a prior distribution represented by samples of nuclear data generated from theory, generalized least squares adjustments are made to incorporate data from historical experiments. To quantify potential uncertainty reduction obtainable from a pair of candidate experiments, we compute the D-optimality criterion of the posterior covariance of intermediate energy range nuclear data compared to the equivalent covariance after additional adjustment to the pair of candidate experiments. Repeating the process for each of many candidate pairs facilitates the final selection. Results support 63⁢ Cu total cross section measurements for differential experiments and alumina and alumina/graphite configurations for integral experiments. This analysis enables choosing differential and integral experiments to be executed concurrently while shortening decision times relative to the current nuclear data pipeline.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Establishing Data Analysis Pipeline for Bulk ATAC-Seq Datasets

We developed an analysis pipeline for transposase-accessible chromatin sequencing (ATAC-Seq) data derived from bulk samples, which brings together publicly available R packages in addition to command-line tools designed for analysis of bulk ATAC-Seq data and can be run on any computer running a Linux-like operating system such as Ubuntu or Apple OSX.

97 MATHEMATICS AND COMPUTING

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING

Hydropower Infrastructure - LAkes, Reservoirs, and RIvers (HILARRI), v4

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2025) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2025) – Power plants that are listed in the 2025 U.S. Hydropower Development Pipeline Data or were listed in previous versions of the dataset These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) – EPA SuRGE sampling locations Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

Hansen, Carly [ORNL] (ORCID:0000000193280838)

Accelerating Control Systems with GitOps: A Path to Automation and Reliability

GitOps is a foundational approach for modernizing infrastructure by leveraging Git as the single source of truth for declarative configurations. The poster explores how GitOps transforms traditional control system infrastructure, services and applications by enabling fully automated, auditable, and version-controlled infrastructure management. Cloud-native and containerized environments are shifting the ecosystem not only in the IT industry but also within the computational science field, as is the case of CERN and Diamond Light Source among other Accelerator/Science facilities which are slowly shifting towards modern software and infrastructure paradigms. The ACORN project, which aims to modernize Fermilab’s control system infrastructure and software is implementing proven best-practices and cutting-edge technology standards including GitOps, containerization, infrastructure as code and modern data pipelines for control system data acquisition and the inclusion of AI/ML in our accelerator complex.

Gonzalez, M. [Fermilab]