Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

HopBox: An image analysis pipeline to characterize hop cone morphology

Abstract Hop cone morphology can influence picking and drying ability, and color can impact consumer preference and may be indicative of quality. However, these characteristics are not generally evaluated in hop breeding programs due to the tedious nature of trait quantification and the extensive variation among cones within a genotype. We developed the HopBox, which is a simply constructed light box with a camera mount, and a publicly available image processing pipeline that identifies hop cones within color‐corrected images, reads a QR code within the image, and outputs data on hop cone length, width, area, perimeter, openness, weight, color, and density. The trained model was applied to images of 500 cones each from 15 replicated advanced hop genotypes from the USDA‐ARS breeding program in Prosser, Washington. Analysis of variance revealed significant ( p < 0.001) differences between genotypes for all traits measured, enabling breeders to discriminate between genotypes for selection purposes. Broad sense heritability for all traits ranged from 0.23 to 0.59. A random sampling of hop cones from the complete dataset revealed that imaging only 5–10 cones adequately captured genotypic variation and provided acceptable rank correlations ( r s > 0.75); however, increasing the sample size to 30 provided optimal precision. Instructions for constructing a HopBox and the code for the analysis pipeline are publicly available online and have wide applicability for hop breeding and research.

Altendorf, Kayla R.↗

Influence of Rare Earth Ce Additions on Microstructure and Mechanical Properties of Experimental Pipeline Steels

Herein, the effect of Ce additions ranging from 57 to 263 ppm is evaluated for an experimental pipeline steel. Compared to the Ce-free steel, progressive Ce additions result in a slightly refined microstructure, significantly improve transverse impact properties, and slightly increase strength. All these observations can be attributed to the gradual transformation of Mn sulfide and Mn–Si–Al oxide inclusions to Ce-containing oxide/sulfides. In particular, the inclusions consist exclusively of sub-5 μm spherical Ce 2 O 2 S particles upon near-stoichiometric additions of Ce, considering the oxygen and sulfur impurity level of the steel. In conclusion, the results suggest that Ce is a potentially promising alloying addition for next-generation pipeline steels by replacing overtly deleterious inclusions with potentially beneficial ones.

36 MATERIALS SCIENCE↗

Data processing pipeline for Tianlai experiment

The Tianlai project is a 21cm intensity mapping experiment for detecting dark energy by measuring the baryon acoustic oscillation (BAO) features in the large scale structure power spectrum. This experiment provides an opportunity to test the data processing methods for cosmological 21cm signal extraction, which is still a great challenge in current radio astronomy research. The 21cm signal is much weaker than the foregrounds and easily aected by the imperfections in the instrumental responses. Furthermore, processing the large volumes of interferometer data poses a practical challenge. We have developed a data processing pipeline called tlpipe to process the drift scan survey data from the Tianlai experiment. It performs oine data processing tasks such as radio frequency interference (RFI) agging, array calibration, binning, and map-making, etc. It also includes utility functions needed for the data analysis, such as data selection, transformation, visualization and others. A number of new algorithms are implemented, for example the eigenvector decomposition method for array calibration and the Tikhnov regularization for m-mode analysis. In this paper we describe the design and implementation of the pipeline and illustrate its functions with some analysis of real data. Finally, we outline directions for future development of this publicly code.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A machine learning pipeline for membrane segmentation of cryo-electron tomograms

We describe how to use several machine learning techniques organized in a learning pipeline to segment and identify cell membrane structures from cryo electron tomograms. These tomograms are difficult to analyze with traditional segmentation tools. The learning pipeline in our approach starts from supervised learning via a special convolutional neural network trained with simulated data. It continues with semi-supervised reinforcement learning and/or a region merging technique that tries to piece together disconnected components belonging to the same membrane structure. A parametric or non-parametric fitting procedure is then used to enhance the segmentation results and quantify uncertainties in the fitting. Domain knowledge is used in generating the training data for the neural network and in guiding the fitting procedure through the use of appropriately chosen priors and constraints. We demonstrate that the approach proposed here works well for extracting membrane surfaces in two real tomogram datasets.

97 MATHEMATICS AND COMPUTING↗

Machine learning pipeline for denoising low signal-to-noise ratio and out-of-distribution transmission electron microscopy datasets

High-resolution transmission electron microscopy (HRTEM) is crucial for observing material’s structural and morphological evolution at Angstrom scales, but the electron beam can alter these processes. Devices such as CMOS-based direct-electron detectors operating in electron-counting mode can be utilized to substantially reduce the electron dosage. However, the resulting images often lead to a low signal-to-noise ratio, which requires frame integration that sacrifices temporal resolution. Several machine learning (ML) models have been recently developed to successfully denoise HRTEM images. Yet, these models are often computationally expensive, and their inference speeds on GPUs are outpaced by the imaging speed of advanced detectors, precluding in situ analysis. Furthermore, the performance of these denoising models on datasets with imaging conditions that deviate from the training datasets has not been evaluated. To mitigate these gaps, we propose a new self-supervised ML denoising pipeline specifically designed for time-series HRTEM images. This pipeline integrates a blind-spot convolution neural network with pre-processing and post-processing steps, including drift correction and low-pass filtering. Results demonstrate that our model outperforms various other ML and non-ML denoising methods in noise reduction and contrast enhancement, leading to improved visual clarity of atomic features. Additionally, the model is drastically faster than U-Net-based ML models and demonstrates excellent out-of-distribution generalization. The model’s computational inference speed is in the order of milliseconds per image, rendering it suitable for application in in-situ HRTEM experiments.

36 MATERIALS SCIENCE↗

A reusable neural network pipeline for unidirectional fiber segmentation

Abstract Fiber-reinforced ceramic-matrix composites are advanced, temperature resistant materials with applications in aerospace engineering. Their analysis involves the detection and separation of fibers, embedded in a fiber bed, from an imaged sample. Currently, this is mostly done using semi-supervised techniques. Here, we present an open, automated computational pipeline to detect fibers from a tomographically reconstructed X-ray volume. We apply our pipeline to a non-trivial dataset by Larson et al . To separate the fibers in these samples, we tested four different architectures of convolutional neural networks. When comparing our neural network approach to a semi-supervised one, we obtained Dice and Matthews coefficients reaching up to 98%, showing that these automated approaches can match human-supervised methods, in some cases separating fibers that human-curated algorithms could not find. The software written for this project is open source, released under a permissive license, and can be freely adapted and re-used in other domains.

79 ASTRONOMY AND ASTROPHYSICS↗

Large-scale real-time signal processing in physics experiments: the ALICE TPC FPGA pipeline

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB s -1 of raw detector data. This requirement is met by a custom FPGA-based processing pipeline that performs the complete front-end data treatment fully in-stream, including common-mode correction, pedestal subtraction, ion-tail filtering, zero suppression, and dense data packing. A central element of the design is a highly parallel common-mode correction algorithm operating directly on the streaming data. It robustly identifies signal-free readout channels on a time-bin basis and applies pad-dependent scaling to compensate for local variations in capacitive coupling in the GEM readout. In combination with pedestal subtraction and ion-tail filtering, this enables accurate baseline restoration under extreme high-occupancy conditions, preventing signal loss while efficiently suppressing noise prior to zero suppression. The pipeline operates continuously at the full detector bandwidth and reduces the raw input rate of approximately 3 TB s -1 to about 900 GBps for Pb-Pb collisions at the target interaction rate. Overall, it represents a large-scale FPGA-based real-time signal-processing implementation for high-energy physics detector readout.

Digital signal processing (DSP)↗

New high-throughput endstation to accelerate the experimental optimization pipeline for synchrotron X-ray footprinting

Synchrotron X-ray footprinting (XF) is a growing structural biology technique that leverages radiation-induced chemical modifications via X-ray radiolysis of water to produce hydroxyl radicals that probe changes in macromolecular structure and dynamics in solution states of interest. The X-ray Footprinting of Biological Materials (XFP) beamline at the National Synchrotron Light Source II provides the structural biology community with access to instrumentation and expert support in the XF method, and is also a platform for development of new technological capabilities in this field. Hee, the design and implementation of a new high-throughput endstation device based around use of a 96-well PCR plate form factor and supporting diagnostic instrumentation for synchrotron XF is described. This development enables a pipeline for rapid comprehensive screening of the influence of sample chemistry on hydroxyl radical dose using a convenient fluorescent assay, illustrated here with a study of 26 organic compounds. The new high-throughput endstation device and sample evaluation pipeline now available at the XFP beamline provide the worldwide structural biology community with a robust resource for carrying out well optimized synchrotron XF studies of challenging biological systems with complex sample compositions.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

A Deep Learning Pipeline for Optimizing Large-scale Phase Field Simulations

Phase field (PF) simulations are computationally expensive but remain a key analysis tool to understand the complex mechanisms of additive manufacturing (AM) processes. Each PF simulation-aided analysis requires thousands of node hours on leadership-class supercomputers. One of the main goals of these analyses is the study of microstructure evolution during the build process which begins with the onset of nucleation. Nucleation occurs under certain thermomechanical conditions which are not known a priori and many PF simulations are required to identify ranges of input thermo-mechanical parameters that can result in the onset of nucleation. Since many of the simulations do not result in nucleation, an analysis campaign often ends up wasting tremendous amounts of precious computing resources executing nucleation-absent simulations. The goal of this work is to design and train deep learning models to inform a PF simulation about the likelihood of the occurrence of nucleation in a future simulation time-step based on the state summary over a finite number of past time-steps of a running simulation. If the prediction determines that the running simulation is unlikely to reach nucleation in the allotted time, then its execution is stopped immediately ultimately resulting in vast reduction in wasted computations when accrued over all the PF simulations typically performed in a single or multiple analysis campaign(s). The paper presents the performance of a machine learning pipeline that uses a convolutional neural network (CNN) model to learn an embedding which is then used with a self-attention network to build a multi-task deep learning model to predict the likelihood of nucleation. The model also predicts the input parameters used in a simulation. Performance is compared with a baseline pipeline that uses an off-the-shelf LeNet-5 model to learn the initial embedding. Despite their smaller size, performance results indicate significant improvement in accuracy of the proposed models compared to the larger baseline models.

Kannan, Ramakrishnan {ramki}↗

Online and Scalable Data Compression Pipeline with Guarantees on Quantities of Interest

Data compression is becoming critical for data-intensive scientific applications. Scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Prior work has shown that a pipeline can be built to guarantee error on the primary data (PD) within user-defined bounds and achieve near-floating point QoI errors. In this paper, we present novel computational approaches for accelerating the pipeline and demonstrate results that enable concurrent execution of compression in parallel with the simulation nodes. This allows compression, including the writing of the required compression data, for the previous time step to be completed while the simulation proceeds with the current time step. Overall, the approach presented in this paper results in a 6–8 times improvement in computational overhead compared to previous work. These results were obtained using data generated by a large-scale fusion code called XGC, which produces hundreds of terabytes of data in a single day.

Banerjee, Tania↗

IMG Annotation Pipeline (IMGAP) v5.1.13

The IMG Annotation Pipeline is a collection of Bash and Python scripts to control a workflow for structural and functional annotation of prokaryotic genomes, metagenomes, and metatranscriptomes. The bash scripts in general control the overall workflow and are wrappers around 3rd party executables (not included in repo) that predict features or functions. Whereas the Python scripts do post-processing of raw output in terms of filtering or format transformation and in some cases contain some logic for picking the correct predictions or resolving overlaps. The pipeline is tailored to produce results required by IMG (https://img.jgi.doe.gov/) and is executed on every dataset submitted to IMG via https://img.jgi.doe.gov/submit. These consist of internal genomes, metagenomes and metatranscriptomes sequenced and assembled at the JGI, as well as datasets submitted by external users (non-lab/JGI affiliates).

Huntemann, Marcel↗

Biosynth Pipeline v1.0

BioPKS Pipeline is a computational pipeline for retrosynthetic design of small molecule biosynthesis pathways (e.g. retrobiosynthesis). It combines capabilities by interfacing with existing retrobiosynthesis tools- RetroTide (developed at LBNL) and DORAnet to create pathways that combine multiple biosynthesis approaches- both megasynthase assembly line enzymes and single step enzymes.

Backman, Tyler [Lawrence Berkeley National Laborat↗

A pipeline for targeted metagenomics of environmental bacteria

Background:Metagenomics and single cell genomics provide a window into the genetic repertoire of yet uncultivated microorganisms, but both methods are usually taxonomically untargeted. The combination of fluorescence in situ hybridization (FISH) and fluorescence activated cell sorting (FACS) has the potential to enrich taxonomically well-defined clades for genomic analyses. Methods:Cells hybridized with a taxon-specific FISH probe are enriched based on their fluorescence signal via flow cytometric cell sorting. A recently developed FISH procedure, the hybridization chain reaction (HCR)-FISH, provides the high signal intensities required for flow cytometric sorting while maintaining the integrity of the cellular DNA for subsequent genome sequencing. Sorted cells are subjected to shotgun sequencing, resulting in targeted metagenomes of low diversity. Results: Pure cultures of different taxonomic groups were used to (1) adapt and optimize the HCR-FISH protocol and (2) assess the effects of various cell fixation methods on both the signal intensity for cell sorting and the quality of subsequent genome amplification and sequencing. Best results were obtained for ethanol-fixed cells in terms of both HCR-FISH signal intensity and genome assembly quality. Our newly developed pipeline was successfully applied to a marine plankton sample from the North Sea yielding good quality metagenome assembled genomes from a yet uncultivated flavobacterial clade. Conclusions: With the developed pipeline, targeted metagenomes at various taxonomic levels can be efficiently retrieved from environmental samples. The resulting metagenome assembled genomes allow for the description of yet uncharacterized microbial clades.

59 BASIC BIOLOGICAL SCIENCES↗

MetaboDirect: an analytical pipeline for the processing of FT-ICR MS-based metabolomic data

Background: Microbiomes are now recognized as the main drivers of ecosystem function ranging from the oceans and soils to humans and bioreactors. However, a grand challenge in microbiome science is to characterize and quantify the chemical currencies of organic matter (i.e., metabolites) that microbes respond to and alter. Critical to this has been the development of Fourier transform ion cyclotron resonance mass spectrometry (FT-ICR MS), which has drastically increased molecular characterization of complex organic matter samples, but challenges users with hundreds of millions of data points where readily available, user-friendly, and customizable software tools are lacking. Results: Here, we build on years of analytical experience with diverse sample types to develop MetaboDirect, an open-source, command-line-based pipeline for the analysis (e.g., chemodiversity analysis, multivariate statistics), visualization (e.g., Van Krevelen diagrams, elemental and molecular class composition plots), and presentation of direct injection high-resolution FT-ICR MS data sets after molecular formula assignment has been performed. When compared to other available FT-ICR MS software, MetaboDirect is superior in that it requires a single line of code to launch a fully automated framework for the generation and visualization of a wide range of plots, with minimal coding experience required. Among the tools evaluated, MetaboDirect is also uniquely able to automatically generate biochemical transformation networks (ab initio) based on mass differences (mass difference network-based approach) that provide an experimental assessment of metabolite connections within a given sample or a complex metabolic system, thereby providing important information about the nature of the samples and the set of microbial reactions or pathways that gave rise to them. Finally, for more experienced users, MetaboDirect allows users to customize plots, outputs, and analyses. Conclusion: Application of MetaboDirect to FT-ICR MS-based metabolomic data sets from a marine phage-bacterial infection experiment and a Sphagnum leachate microbiome incubation experiment showcase the exploration capabilities of the pipeline that will enable the research community to evaluate and interpret their data in greater depth and in less time. It will further advance our knowledge of how microbial communities influence and are influenced by the chemical makeup of the surrounding system. The source code and User’s guide of MetaboDirect are freely available through (https://github.com/Coayala/MetaboDirect) and (https://metabodirect.readthedocs.io/en/latest/), respectively.

54 ENVIRONMENTAL SCIENCES↗

Distributed fiber sensor and machine learning data analytics for pipeline protection against extrinsic intrusions and intrinsic corrosions

This paper presents an integrated technical framework to protect pipelines against both malicious intrusions and piping degradation using a distributed fiber sensing technology and artificial intelligence. A distributed acoustic sensing (DAS) system based on phase-sensitive optical time-domain reflectometry (φ-OTDR) was used to detect acoustic wave propagation and scattering along pipeline structures consisting of straight piping and sharp bend elbow. Signal to noise ratio of the DAS system was enhanced by femtosecond induced artificial Rayleigh scattering centers. Data harnessed by the DAS system were analyzed by neural network-based machine learning algorithms. The system identified with over 85% accuracy in various external impact events, and over 94% accuracy for defect identification through supervised learning and 71% accuracy through unsupervised learning.

Peng, Zhaoqiang↗

End-to-End Pipeline for Trigger Detection on Hit and Track Graphs

There has been a surge of interest in applying deep learning in particle and nuclear physics to replace labor-intensive offline data analysis with automated online machine learning tasks. This paper details a novel AI-enabled triggering solution for physics experiments in Relativistic Heavy Ion Collider and future Electron-Ion Collider. The triggering system consists of a comprehensive end-to-end pipeline based on Graph Neural Networks that classifies trigger events versus background events, makes online decisions to retain signal data, and enables efficient data acquisition. Here, the triggering system first starts with the coordinates of pixel hits lit up by passing particles in the detector, applies three stages of event processing (hits clustering, track reconstruction, and trigger detection), and labels all processed events with the binary tag of trigger versus background events. By switching among different objective functions, we train the Graph Neural Networks in the pipeline to solve multiple tasks: the edge-level track reconstruction problem, the edge-level track adjacency matrix prediction, and the graph-level trigger detection problem. We propose a novel method to treat the events as track-graphs instead of hit-graphs. This method focuses on intertrack relations and is driven by underlying physics processing. As a result, it attains a solid performance (around 72% accuracy) for trigger detection and outperforms the baseline method using hit-graphs by 2% higher accuracy.

97 MATHEMATICS AND COMPUTING↗

Fatigue Performance of High-Strength Pipeline Steels and Their Welds in Hydrogen Gas Service

Objectives of the project include: Enable the use of high strength steel hydrogen pipelines, as significant cost savings can result by implementing high strength steels as compared to lower strength pipes. Demonstrate that girth welds in high-strength steel pipe exhibit fatigue performance similar to lower-strength steels in high-pressure hydrogen gas. Identify pathways for developing high-strength pipeline steels by establishing the relationship between microstructure constituents and hydrogen-accelerated fatigue crack growth (HA-FCG)

08 HYDROGEN↗

Remote methane sensor for emissions from pipelines and compressor stations using chirped-laser dispersion spectroscopy

Leak rates of methane (CH 4 ) from the natural gas supply chain result in lost profit from unsold product, public safety and property concerns due to potential explosion hazards, and a potentially large source of economic damages from legal liabilities. Yet large measurement challenges exist in identifying and quantifying CH 4 leak rates along the vast number and type of components in the natural gas supply chain. This is particularly true of the “midstream” components involved in the gathering, processing, compression, transmission, and storage of natural gas. This project developed and deployed new advances in chirped laser dispersion spectroscopy (CLaDS) to detect methane leaks from pipelines, compressor stations, and other midstream infrastructure from a remote position (standoff detection). The system was deployed from a van to measure fugitive methane leaks from a local compressor station as a proof-of-concept. The technique was validated through mobile laboratory measurements with in-situ sensors as well as controlled releases of methane. The system also mapped a plume from a controlled release of methane by tracking a small unmanned aerial system (sUAS) that carried a corner cube retroreflector which reflected the beam back to the instrument. The drone-based system quantified a leak rate to within 30% of the actual rate and localized the emission location within 5 m of the actual release location at standoff distances of 25-45 m. This sUAS-reflector tracking approach has benefits for mapping leak locations remotely using small drones flying around a facility. Benefits of a commercial sensor with these capabilities include reductions of leaks for pipeline operators (more profit), earlier detection of leaks to avoid catastrophic explosion hazards for public health and to mitigate property damage, and reduced methane emissions to the atmosphere (improving air quality).

03 NATURAL GAS↗