Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

Pipeline Hydrogen Decarbonization and Repurposing Analyzers (P-HyDRAs)

The Pipeline Hydrogen Decarbonization and Repurposing Analyzers (P-HyDRAs) are a set of prototype computational tools for simulating and optimizing midstream natural gas pipeline system operations subject to location and time-dependent hydrogen blending. The models can accurately resolve dynamic gas flows through large-scale pipeline networks using non-ideal gas equations of state. The codes can be used as decision support for planning and design decisions involving intra-day energy flow schedules as well as spatiotemporal economic values of natural gas, hydrogen, and net energy delivered to consumers while ensuring that pipeline hydraulic limitations, gas compressor station constraints, operational factors, and pre-existing shipping contracts are satisfied. The inputs to the codes are a model of the pipeline system as well as time-series data that specify boundary conditions on the network. For optimization, the code module requires price and quantity offers for natural gas and hydrogen and price and quantity bids for energy, which are used as time-dependent constraints in an optimal control problem. The outputs are time-series data that provide a predictive simulation of gas flows, mass fractions, and pressures, or with additional degrees of freedom give an approximately optimal solution for gas injections/withdrawals, compressor settings, and sensitivities to the objective function that provide locational values of energy.

Zlotnik, Anatoly↗

A Data Processing Pipeline To Extract A Knowledge Graph From Sec Documents For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest (disk) expressed as a latitude/longitude point and distance, and a set of SEC form types from which to extract entities and relations. There are three main components to this pipeline as currently implemented: Social Network Extraction, Critical Infrastructure Network Extraction, and Inference and Fusion. First, Social Network Extraction, implemented as the `organizations_sec` component of the workflow graph queries the SEC EDGAR webservice using the list of initial companies from the configuration file. Given this, it extracts metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Second, the Critical Network Extraction component extracts entities and relations for a critical infrastructure sector. Currently, we focus on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Third, the Inference and Fusion component relates the social network graph to the critical infrastructure graph in order to understand the impact of a company within a geographic region. Relations include ownership of the EV Charging Station asset as well as maintenance/ownership of the EV payment networks. The fused network can be represented in many ways and currently we emit a knowledge graph.

Weaver, GabrielA.↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

Abstract for CRADA between National Energy Technology Laboratory and Colonial Pipeline Company

The National Energy Technology Laboratory (NETL) and Colonial Pipeline Company (Participant) will collaborate in the field demonstration of optical fiber sensor systems developed at NETL on Participant’s fuel pipeline. The optical fiber sensor technologies are capable of distributed temperature and strain sensing, distributed acoustic sensing, and ultrasensitive acoustic sensing. Real-time monitoring of these parameters enables pipeline integrity monitoring, security monitoring, flow rate monitoring, etc. Successful demonstration on a real fuel pipeline at Participant’s facilities will validate the sensor technologies and installation methods at a real scale. This effort aligns with NETL’s mission in reliable and sustainable energy and reducing environmental effects due to pipeline failures.

42 ENGINEERING↗

Onshore U.S. Carbon Pipeline Deployment: Siting, Safety, and Regulation

Carbon capture, utilization, and storage technology has significant potential to reduce greenhouse gas emissions and mitigate the impact of climate change, particularly in hard-to-decarbonize industrial and commercial sectors. Unless utilization or storage occurs at the same location as carbon capture, which is rare, carbon must be transported from a point source to a utilization or storage site. Pipelines offer significant advantages for large-scale transportation of carbon dioxide over other methods. While studies show that reaching net-zero carbon emissions in the United States by 2050 will require between 29,000 and 66,000 miles of carbon pipelines, the U.S. had deployed fewer than 6,000 miles of carbon pipelines by 2022. Although closing this gap is important to achieving low-carbon goals, carbon pipelines operate in a complex and uncertain local, state and federal regulatory landscape and face public concerns about safety and siting. This report covers numerous regulatory issues surrounding carbon pipeline development, including the current narrow federal definition of carbon dioxide and the considerable variation in state and local governments’ laws and regulations. The report serves as a primer for regulators and stakeholders who seek to better understand the regulatory challenges and opportunities facing this critical infrastructure

01 COAL, LIGNITE, AND PEAT↗

Transport Affordable Clean Hydrogen Energy via Existing Pipeline Infrastructure (CRADA 573)

Our Nation’s vast network of oil pipelines span over 230,000 miles connecting remote energy producing regions to distant markets. The COVID-19 pandemic’s influence on oil prices intensified technical and economic vulnerabilities afflicting the Trans-Alaska Pipeline System (TAPS, operating below 25% capacity) that may soon resurface as energy demand shifts to low carbon sources. Through the opportunity afforded by the Arctic Advanced Manufacturing Program, I’ve toured the great State of Alaska learning about its culture, unique energy challenges, business ecosystem, research institutions, and regulatory agencies. With encouragement and support from dedicated Alaskans and my professional mentors, I’ve founded Mighty Pipeline for the purpose of developing and commercializing proprietary technology and hardware to convert oil pipelines into clean hydrogen energy transmission systems. Leveraging the advanced technical capabilities of Pacific Northwest National Laboratory, we’ve demonstrated proof-of-concept (TRL = 3) and are seeking pre-pilot development project opportunities to conduct hardware and system level tests. If we successfully achieve technical and regulatory milestones, Mighty Pipeline’s technology could be available to help facilitate bulk clean hydrogen energy export this decade.

08 HYDROGEN↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN↗

PhytoOracle: Scalable, modular phenomics data processing pipelines

As phenomics data volume and dimensionality increase due to advancements in sensor technology, there is an urgent need to develop and implement scalable data processing pipelines. Current phenomics data processing pipelines lack modularity, extensibility, and processing distribution across sensor modalities and phenotyping platforms. To address these challenges, we developed PhytoOracle (PO), a suite of modular, scalable pipelines for processing large volumes of field phenomics RGB, thermal, PSII chlorophyll fluorescence 2D images, and 3D point clouds. PhytoOracle aims to ( i ) improve data processing efficiency; ( ii ) provide an extensible, reproducible computing framework; and ( iii ) enable data fusion of multi-modal phenomics data. PhytoOracle integrates open-source distributed computing frameworks for parallel processing on high-performance computing, cloud, and local computing environments. Each pipeline component is available as a standalone container, providing transferability, extensibility, and reproducibility. The PO pipeline extracts and associates individual plant traits across sensor modalities and collection time points, representing a unique multi-system approach to addressing the genotype-phenotype gap. To date, PO supports lettuce and sorghum phenotypic trait extraction, with a goal of widening the range of supported species in the future. At the maximum number of cores tested in this study (1,024 cores), PO processing times were: 235 minutes for 9,270 RGB images (140.7 GB), 235 minutes for 9,270 thermal images (5.4 GB), and 13 minutes for 39,678 PSII images (86.2 GB). These processing times represent end-to-end processing, from raw data to fully processed numerical phenotypic trait data. Repeatability values of 0.39-0.95 (bounding area), 0.81-0.95 (axis-aligned bounding volume), 0.79-0.94 (oriented bounding volume), 0.83-0.95 (plant height), and 0.81-0.95 (number of points) were observed in Field Scanalyzer data. We also show the ability of PO to process drone data with a repeatability of 0.55-0.95 (bounding area).

59 BASIC BIOLOGICAL SCIENCES↗

Investigation of the Hydrogen Embrittlement of API 5L Natural Gas Pipeline Steels

Hydrogen ions produced during corrosion or through hydrogen blending into natural gas pipelines can lead to the degradation of ductility of metals used in these transmission pipelines. The effect of hydrogen on the mechanical properties of pipeline steels was studied for three grades of API 5L steels: X56, X65, and X100. These steels are either commonly used in or considered for use in natural gas transmission pipelines which are being considered for use in hydrogen blending. These steels were subjected to constant strain rate tensile testing after charging with electrochemically generated hydrogen. This presentation reports on work to extend the life of the natural gas pipeline network via understanding the effect of hydrogen embrittlement induced by corrosion processes.

Teeter, Lucas↗

Pilot-Scale Validation of Distributed Optical Fiber Sensors for Underground Pipeline Monitoring

Distributed fiber optic sensing is a cutting-edge technology that has found extensive applications in the monitoring of Ensuring the safety, integrity, and operational efficiency of underground product pipelines is vital for maintaining the nation’s critical infrastructure. Monitoring parameters such as hoop strain, pressure, and acoustic vibrations is key to detecting potential leaks, intrusions, or structural issues. Distributed optical fiber sensor (DOFS) systems provide a compelling solution for continuous, real-time monitoring over long distances. This paper details the development and pilot-scale implementation of DOFS systems for underground pipeline monitoring, evolving from a proof-of-concept stage. Multiple custom-designed DOFS interrogator units—such as optical frequency-domain reflectometry (OFDR), Brillouin optical time-domain analysis (BOTDA), and multimodal interferometer-based fiber acoustic sensors—were employed to measure key parameters like hoop strain, pressure, and acoustic vibrations. The underground product pipeline's outer diameter is 30 inches, the wall thickness is 1.28 inches, and the 3-foot depth. The fiber deployment strategies, and sensing data acquisition methods for these systems are discussed. The results demonstrate the effectiveness of DOFS in detecting hoop strain, temperature changes, and acoustic vibrations, showcasing their potential for real-time monitoring and enhancing pipeline safety.

distributed fiber sensing↗

Automated Image Segmentation and Processing Pipeline Applied to X–Ray Computed Tomography Studies of Pitting Corrosion in Aluminum Wires

Understanding pitting corrosion is critical, yet its kinetics and morphology remain challenging to study from X-ray computed tomography (XCT) due to manual segmentation barriers. To address this, an automated pipeline leveraging deep learning for efficient large-scale XCT analysis is developed, revealing new corrosion insights. The pipeline enables pit segmentation, 3D reconstruction, statistical characterization, and a topological transformation for visualization. Here, the pipeline is applied to 87 648 XCT images capturing commercial purity aluminum (1100 Al) wire exposed to sodium chloride (NaCl) salt particles over a period of 122 h. The pipeline achieves complete feature extraction and statistical quantification across the entire XCT dataset, leveraging distributed computing environment for high efficiency. Global growth kinetics such as high-level stepwise sigmoidal volume loss patterns and granular individual pit developments are both captured for 36 detected pits. By combining automation, computer vision, and extensive XCT datasets, this research accelerates precise corrosion assessment to enable materials science discoveries at scale.

36 MATERIALS SCIENCE↗

Alaska Liquid Natural Gas Pipeline Front-End Engineering & Design (Final Technical Report)

The Alaska Gasline Development Corporation (AGDC) is Alaska’s natural gas infrastructure development corporation established in 2013. AGDC’s mission is to maximize the benefit of Alaska’s vast North Slope natural gas resources for Alaskans through the development of infrastructure necessary to move the gas into local and international markets. AGDC was identified for a Congressionally Directed Spending (CDS) project for funding in the Energy and Water Development and Related Agencies Appropriations Act, 2023 under the heading: “Congressionally Directed Energy Efficiency and Renewable Energy Projects.” The CDS included $\$$4,000,000 of direct funding, with required match funds, to move the project forward. Alaska’s North Slope holds America’s largest proven and conventional natural gas supply. The integrated Alaska LNG Project will deliver 3.5 billion cubic feet of natural gas per day from Alaska’s North Slope gas fields to Alaskans as well as to a marine terminal located at tidewater in Cook Inlet. Alaska LNG is an integrated gas infrastructure project with three major components: a gas treatment plant (GTP) located at Prudhoe Bay, an 807-mile (1,287 km) gas pipeline (Mainline Pipeline) to Southcentral Alaska with interconnections for in-state gas use, and a natural gas liquefaction facility (LNG Facility) in Nikiski, Alaska. The integrated Alaska LNG Project has several strategic advantages including proven gas resources, existing upstream infrastructure, an advantageous arctic climate for LNG production, proximity to LNG markets, a track record of reliability from a state that first began exporting LNG to Japan in 1969, and broad support from Alaskans. North Slope natural gas is a conventional resource and can be produced with minimal drilling at a fraction of the carbon dioxide emissions of shale gas from the Lower 48 states. Through the development of the Alaska LNG Project, Alaska can provide energy security to Alaskans and a stable source of LNG to the Asia-Pacific region for generations. The Alaska LNG Project has been progressed through Pre-Front-End Engineering Design (Pre-FEED) and has obtained all major federal and State of Alaska permits and authorizations to construct the project, including the Federal Energy Regulatory Commission (FERC) Order Granting Authorization Under Section 3 of the Natural Gas Act. On September 5, 2024, the U.S. Department of Energy (DOE), National Energy Technology Laboratory (NETL) awarded Project No. DE-FE0032307 to AGDC with the objective to progress the project to Front-End Engineering Design (FEED) entry for the Alaska LNG Project Phase 1 Pipeline. The award Start Date was made effective July 1, 2023, with a Period of Performance through June 30, 2025. On March 27, 2025, AGDC announced the execution of definitive commercial agreements with Glenfarne Alaska LNG, LLC, an affiliate of Glenfarne Group, LLC, (together as “Glenfarne”), to lead the development of the Alaska LNG Project and enter FEED for the Phase 1 Pipeline. Project activities are now funded and directed by this private sector partner who holds a 75% interest in 8 Star Alaska, LLC (8 Star). 8 Star holds the assets of the Alaska LNG Project. As planned, AGDC continues to hold 25% minority interest in 8 Star and will play a governance role moving forward with Alaska LNG. This definitive commercial agreement milestone led to the successful completion of AGDC’s Statement of Project Objectives (SOPO) for FEED entry and led to the completion of DOE Project No. DE-FE0032307. At conclusion of the SOPO, AGDC also reached the award’s maximum federal cost share of $\$$4,000,000. AGDC is, therefore, providing Final Technical Report to close out DOE Project No. DE-FE0032307.

02 PETROLEUM↗

Quasi-Distributed Fiber Sensor-Based Approach for Pipeline Health Monitoring: Generating and Analyzing Physics-Based Simulation Datasets for Classification

This study presents a framework for detecting mechanical damage in pipelines, focusing on generating simulated data and sampling to emulate distributed acoustic sensing (DAS) system responses. The workflow transforms simulated ultrasonic guided wave (UGW) responses into DAS or quasi-DAS system responses to create a physically robust dataset for pipeline event classification, including welds, clips, and corrosion defects. This investigation examines the effects of sensing systems and noise on classification performance, emphasizing the importance of selecting the appropriate sensing system for a specific application. The framework shows the robustness of different sensor number deployments to experimentally relevant noise levels, demonstrating its applicability in real-world scenarios where noise is present. Overall, this study contributes to the development of a more reliable and effective method for detecting mechanical damage to pipelines by emphasizing the generation and utilization of simulated DAS system responses for pipeline classification efforts. The results on the effects of sensing systems and noise on classification performance further enhance the robustness and reliability of the framework.

36 MATERIALS SCIENCE↗

An Update on the Geothermal Data Repository's Data Standards and Pipelines: Geospatial Data and Distributed Acoustic Sensing Data: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.

cloud-optimized↗

An Update on the Geothermal Data Repository's Data Standards and Pipelines: Geospatial Data and Distributed Acoustic Sensing Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.

cloud-optimized↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗