Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

L-PBF High-Throughput Data Pipeline Approach for Multi-modal Integration

Abstract Metal-based additive manufacturing requires active monitoring solutions for assessing part quality. Multiple sensors and data streams, however, generate large heterogeneous data sets that are impractical for manual assessment and characterization. In this work, an automated pipeline is developed that enables feature extraction from high-speed camera video and multi-modal data analysis. The framework removes the need for manual assessment through the utilization of deep learning techniques and training models in a weakly supervised paradigm. We demonstrate this pipeline’s capability over 700,000 high-speed camera frames. The pipeline successfully extracts melt pool and spatter geometries and links them to corresponding pyrometry, radiography, and processparameter information. 715 individual prints are examined to reveal melt pool areas that exceeds 0.07 mm 2 and pyrometry signal over a threshold (375 pyrometry units) were more likely to have defects. These automated processes enable massive throughput of characterization techniques.

36 MATERIALS SCIENCE↗

Optimal high-throughput virtual screening pipeline for efficient selection of redox-active organic materials

As global interest in renewable energy continues to increase, there has been a pressing need for developing novel energy storage devices based on organic electrode materials that can overcome the shortcomings of the current lithium-ion batteries. One critical challenge for this quest is to find materials whose redox potential (RP) meets specific design targets. In this study, we propose a computational framework for addressing this challenge through the effective design and optimal operation of a high-throughput virtual screening (HTVS) pipeline that enables rapid screening of organic materials that satisfy the desired criteria. Starting from a high-fidelity model for estimating the RP of a given material, we show how a set of surrogate models with different accuracy and complexity may be designed to construct a highly accurate and efficient HTVS pipeline. We demonstrate that the proposed HTVS pipeline construction and operation strategies substantially enhance the overall screening throughput.

36 MATERIALS SCIENCE↗

A simulation pipeline for fast neutron imaging and spectroscopy using quantified detector attributes

Radiation imaging capabilities, essential in the nuclear nonproliferation regime, facilitate source localization and, in certain cases, spectroscopy. Scatter-based neutron cameras, which can measure the neutron signatures from special nuclear material, hold particular interest. Systems incorporating organic scintillators can extract neutron energy spectra, potentially distinguishing fission neutron sources from others, such as alpha-neutron sources. The development and testing of a scatter-based neutron imager, however, can be challenging without having an accurate simulation model or first constructing a prototype. This work describes a simulation pipeline that takes output from MCNPX-PoliMi simulations and creates the expected back-projection neutron images and neutron energy spectra. This pipeline was developed to improve the modeling of fast neutron imagers and bridge the current gap in literature, which predominantly focuses on gamma-ray Compton imager models. This work also reports on the significance of various real-world system considerations and their effects on the simulated detector responses. The pipeline was verified and validated with experimental data collected using a 252 Cf spontaneous fission source using a fast neutron scattering imager developed at the University of Michigan.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

AmeriFlux BASE data pipeline to support network growth and data sharing

Abstract AmeriFlux is a network of research sites that measure carbon, water, and energy fluxes between ecosystems and the atmosphere using the eddy covariance technique to study a variety of Earth science questions. AmeriFlux’s diversity of ecosystems, instruments, and data-processing routines create challenges for data standardization, quality assurance, and sharing across the network. To address these challenges, the AmeriFlux Management Project (AMP) designed and implemented the BASE data-processing pipeline. The pipeline begins with data uploaded by the site teams, followed by the AMP team’s quality assurance and quality control (QA/QC), ingestion of site metadata, and publication of the BASE data product. The semi-automated pipeline enables us to keep pace with the rapid growth of the network. As of 2022, the AmeriFlux BASE data product contains 3,130 site years of data from 444 sites, with standardized units and variable names of more than 60 common variables, representing the largest long-term data repository for flux-met data in the world. The standardized, quality-ensured data product facilitates multisite comparisons, model evaluations, and data syntheses.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Image processing pipeline for AI-driven nanoparticle megalibrary characterization

Recent innovations have made it possible to produce megalibraries, millions of structurally and compositionally distinct nanoparticles on a chip. These megalibraries yield vast volumes of data that are impossible to analyze manually, necessitating the development of automated tools. In previous work, we created a binary classification machine learning model to select quality nanoparticle images for downstream analysis. In this work, we show that adding a custom image processing step before training can produce significantly higher-performing models in a fraction of the time and make them more robust to different image noise levels and microscope acquisition settings. The image processing pipeline proposed here effectively cleans raw nanoparticle images, enhances key features, and allows us to use much lower resolution images and simpler neural network model architectures. These features result in higher performance and significant cost savings. Experiments demonstrate superior performance relative to baseline, including an 18.2% improvement in recall and a 13.1% increase in accuracy. Given the high cost of downstream analysis, it is critical to minimize false positives, and our best-performing model reaches a precision of 95.9% and a weighted F-score of 95.1% on an unseen test set. Additionally, model training time is reduced from hours to less than a minute. We also show that, using this custom image processing pipeline, model performance is significantly improved at lower pixel resolutions compared to downsizing alone. We expect that adopting this pipeline for AI-driven automated nanoparticle characterization will allow researchers to rapidly and accurately analyze much greater volumes of data, thereby accelerating materials discovery.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Assessing Compatibility of Natural Gas Pipeline Materials with Hydrogen, CO 2 , and Ammonia

Here, in this study, we examine the efficacy of repurposing natural gas (NG) pipelines for transporting hydrogen blends (with NG), ammonia, and CO 2 (gaseous and supercritical) from the standpoint of materials compatibility. Some information pertaining to component performance is also included, especially those components critical for pressurization and monitoring of flow. A listing of critical pipeline components and materials was developed, and their compatibilities was assessed for each fluid or gas type based on known compatibilities. Results indicate that pipeline materials should be suitable for gaseous CO 2 and anhydrous ammonia, but hydrogen blends greater than 12% may be problematic. Current compressor/regulator stations will not be suitable for use with either supercritical CO 2 or ammonia. Important knowledge gaps were identified, including (1) polymer performance with hydrogen/NG blends at low pressures, (2) compressor/regulator station polymers and epoxy coating materials with supercritical CO 2 , and (3) metal performances of hydrogen/NG blends at low pressures.

03 NATURAL GAS↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

The FRB-searching Pipeline of the Tianlai Cylinder Pathfinder Array

This paper presents the design, calibration, and survey strategy of the Fast Radio Burst (FRB) digital backend and its real-time data processing pipeline employed in the Tianlai Cylinder Pathfinder Array. The array, consisting of three parallel cylindrical reflectors and equipped with 96 dual-polarization feeds, is a radio interferometer array designed for conducting drift scans of the northern celestial semi-sphere. The FRB digital backend enables the formation of 96 digital beams, effectively covering an area of approximately 40 square degrees with the 3 dB beam. Our pipeline demonstrates the capability to conduct an automatic search of FRBs, detecting at quasi-real-time and classifying FRB candidates automatically. The current FRB searching pipeline has an overall recall rate of 88%. During the commissioning phase, we successfully detected signals emitted by four well-known pulsars: PSR B0329+54, B2021+51, B0823+26, and B2020+28. We report the first discovery of an FRB by our array, designated as FRB 20220414A. We also investigate the optimal arrangement for the digitally formed beams to achieve maximum detection rate by numerical simulation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

Feature Engineering and Ensemble Methods for Imbalanced ICS Intrusion Detection: Pipeline Audit and Constrained Evaluation

Industries are becoming increasingly connected and are more vulnerable to cyberattacks due to the widened attack surface. Industrial Control Systems (ICS) are among the most critical sectors that malicious actors can target, as such attacks can cause significant operational disruption and physical damage. It is imperative to detect such attacks as early as possible. This paper evaluates constraint-conditioned optimistic performance estimates for traditional ML models in ICS intrusion detection (i.e., estimates obtained under contiguous, non-shuffled temporal evaluation without test-set alteration, but with pre-split feature engineering that may introduce temporal leakage, due to dataset constraints). Our findings are threefold. First, we quantify how iterative feature engineering affects tree-based ensemble performance and examine how pipeline decisions (split strategy, sampling scope, and cleaning policy) can inflate or reduce reported IDS results under constraint-bound evaluation. Second, we compare intrinsic class-imbalance handling across ensemble models. Third, under our current pipeline constraints (including pre-split feature engineering), CatBoost achieves the best performance on Water Storage Tank (accuracy: 0.9831, class-1 F1: 0.9682), while Light- GBM achieves the best performance on Gas Pipeline (accuracy: 0.9618, class-1 F1: 0.9086).

97 MATHEMATICS AND COMPUTING↗

In-Situ and Ex-Situ Studies on the Morphology Changes of Polymer Pipeline Materials for Use in Hydrogen Gas Environments

The US natural gas infrastructure is a national asset that could be used to deliver hydrogen and hydrogen blends of natural gas as a pathway to reduce carbon emissions. The distribution system comprises nearly 50% plastic pipe composed of medium- and high-density polyethylene materials (MDPE and HDPE). While these materials perform adequately for natural gas, research on their hydrogen compatibility is essential to understand if any immediate and long-term risks are associated with hydrogen addition. The Blended Gas CRADA, a HyBlend project, has established a comprehensive test method for evaluating MDPE and HDPE of various plastic resin compositions of pipeline material in pure hydrogen and 20% hydrogen/80% methane blends. Both in-situ and ex-situ measurements were performed to capture hydrogen-induced changes in the polyethylene material's crystalline, amorphous and their interphase regions. We investigated MDPE and HDPE pipeline materials made from different polymer resin systems to evaluate the effects of hydrogen gas. The materials were characterized by their density, diffusion coefficient, free volume ratio, and degree of crystallinity. Various advanced characterization methods, including in situ NMR, ex situ XRD, ex situ DSC, and ex situ TDA, were used to analyze the effects of changes in crystalline, amorphous, and interphase regions due to gas exposure. Time-dependent post-decompression quasi-static tensile tests were conducted to explore the effects of gas exposure time on the mechanical behavior of the pipe materials. This work will highlight the time sensitivities during and after gas exposure. The correlation between gas-induced polyethylene morphology changes and the associated material performance will be addressed for the intended applications. These studies will show that polyethylene resin composition and material exposure are important factors when considering whether hydrogen gas affects pipeline materials positively or negatively.

Simmons, Kevin L.↗

Multi-choice Viromics Pipeline (MVP) v1

MVP stands for Multi-choice Viromics Pipeline. It is a pipeline that utilizes a suite of state-of-art tools: geNomad to identify viruses, proviruses, and plasmids in sequencing data, CheckV to assess the quality, and completeness of identified viral genomes, including identification of host contamination for integrated proviruses, A custom code for a rapid genome clustering based on pairwise ANI, Bowtie2, Samtools, and CoverM to calculate coverage of individual viral genomes by read mapping, A custom code to create a vOTU table of abundance, MMseqs2 to compare viral proteins to multiple databases. It provides a quick, and intuitive pipeline to get viral sequences and corresponding properties that can be used for downstream analyses.

Roux, Simon↗

Data Agnostic Feature-Target Analysis & Ranking Machine Learning Pipeline (DAFTAR-ML) v0.1.0

DAFTAR-ML is a specialized machine-learning pipeline that identifies relevant features based on their relationship to a target variable. Many ML pipelines focus solely on prediction, and feature ranking is often absent or lacks robust statistical methods. DAFTAR-ML performs its tasks with this outcome in mind. Model training is robust, using nested cross-validation and hyperparameter tuning. Instead of relying on native feature-importance scores, it employs SHAP (SHapley Additive exPlanations) to quantify feature importance. The pipeline also produces comprehensive results, including publication-quality visualizations.

Melie, Tina [Lawrence Berkeley National Laboratory↗

HT-SIP: a semi-automated stable isotope probing pipeline identifies cross-kingdom interactions in the hyphosphere of arbuscular mycorrhizal fungi

Abstract Background Linking the identity of wild microbes with their ecophysiological traits and environmental functions is a key ambition for microbial ecologists. Of many techniques that strive for this goal, Stable-isotope probing—SIP—remains among the most comprehensive for studying whole microbial communities in situ. In DNA-SIP, actively growing microorganisms that take up an isotopically heavy substrate build heavier DNA, which can be partitioned by density into multiple fractions and sequenced. However, SIP is relatively low throughput and requires significant hands-on labor. We designed and tested a semi-automated, high-throughput SIP (HT-SIP) pipeline to support well-replicated, temporally resolved amplicon and metagenomics experiments. We applied this pipeline to a soil microhabitat with significant ecological importance—the hyphosphere zone surrounding arbuscular mycorrhizal fungal (AMF) hyphae. AMF form symbiotic relationships with most plant species and play key roles in terrestrial nutrient and carbon cycling. Results Our HT-SIP pipeline for fractionation, cleanup, and nucleic acid quantification of density gradients requires one-sixth of the hands-on labor compared to manual SIP and allows 16 samples to be processed simultaneously. Automated density fractionation increased the reproducibility of SIP gradients compared to manual fractionation, and we show adding a non-ionic detergent to the gradient buffer improved SIP DNA recovery. We applied HT-SIP to 13 C-AMF hyphosphere DNA from a 13 CO 2 plant labeling study and created metagenome-assembled genomes (MAGs) using high-resolution SIP metagenomics (14 metagenomes per gradient). SIP confirmed the AMF Rhizophagus intraradices and associated MAGs were highly enriched (10–33 atom% 13 C), even though the soils’ overall enrichment was low (1.8 atom% 13 C). We assembled 212 13 C-hyphosphere MAGs; the hyphosphere taxa that assimilated the most AMF-derived 13 C were from the phyla Myxococcota, Fibrobacterota, Verrucomicrobiota, and the ammonia-oxidizing archaeon genus Nitrososphaera . Conclusions Our semi-automated HT-SIP approach decreases operator time and improves reproducibility by targeting the most labor-intensive steps of SIP—fraction collection and cleanup. We illustrate this approach in a unique and understudied soil microhabitat—generating MAGs of actively growing microbes living in the AMF hyphosphere (without plant roots). The MAGs’ phylogenetic composition and gene content suggest predation, decomposition, and ammonia oxidation may be key processes in hyphosphere nutrient cycling.

59 BASIC BIOLOGICAL SCIENCES↗

A miniaturized feedstocks-to-fuels pipeline for screening the efficiency of deconstruction and microbial conversion of lignocellulosic biomass

Sustainably grown biomass is a promising alternative to produce fuels and chemicals and reduce the dependency on fossil energy sources. However, the efficient conversion of lignocellulosic biomass into biofuels and bioproducts often requires extensive testing of components and reaction conditions used in the pretreatment, saccharification, and bioconversion steps. This restriction can result in a significant and unwieldy number of combinations of biomass types, solvents, microbial strains, and operational parameters that need to be characterized, turning these efforts into a daunting and time-consuming task. Here we developed a high-throughput feedstocks-to-fuels screening platform to address these challenges. The result is a miniaturized semi-automated platform that leverages the capabilities of a solid handling robot, a liquid handling robot, analytical instruments, and a centralized data repository, adapted to operate as an ionic-liquid-based biomass conversion pipeline. The pipeline was tested by using sorghum as feedstock, the biocompatible ionic liquid cholinium phosphate as pretreatment solvent, a “one-pot” process configuration that does not require ionic liquid removal after pretreatment, and an engineered strain of the yeast Rhodosporidium toruloides that produces the jet-fuel precursor bisabolene as a conversion microbe. By the simultaneous processing of 48 samples, we show that this configuration and reaction conditions result in sugar yields (~70%) and bisabolene titers (~1500 mg/L) that are comparable to the efficiencies observed at larger scales but require only a fraction of the time. We expect that this Feedstocks-to-Fuels pipeline will become an effective tool to screen thousands of bioenergy crop and feedstock samples and assist process optimization efforts and the development of predictive deconstruction approaches.

09 BIOMASS FUELS↗

Literature Review of Electromagnetic Pulse (EMP) and Geomagnetic Disturbance (GMD) Effects on Oil and Gas Pipeline Systems

This document summarizes the findings of a review of published literature regarding the potential impacts of electromagnetic pulse (EMP) and geomagnetic disturbance (GMD) phenomena on oil and gas pipeline systems. The impacts of telluric currents on pipelines and their associated cathodic protection systems has been well studied. The existing literature describes implications for corrosion protection system design and monitoring to mitigate these impacts. Effects of an EMP on pipelines is not a thoroughly explored subject. Most directly related articles only present theoretical models and approaches rather than specific analyses and in-field testing. Literature on SCADA components and EMP is similarly sparse and the existing articles show a variety of impacts to control system components that range from upset and damage to no effect. The limited research and the range of observed impacts for the research that has been published suggests the need for additional work on GMD and EMP and natural gas SCADA components.

02 PETROLEUM↗

Machine Learning Modeling Pipeline for Extracting Nuclear Proliferation Events of Interest from Open Data Sources (U)

In FY2020, the Savannah River National Laboratory (SRNL) and the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Polytechnic Institute and State University entered a collaboration funded by Department of Energy’s (DOE) Office of Defense Nuclear Nonproliferation Research and Development. The project’s mission was to take the first steps toward developing a demonstration prototype system that uses multiple machine learning and data analytics methods on largescale open data sources to identify new, developing, and/or undeclared nuclear programs. Given the SRNL team’s on-site perspective of events culminating in the DOE’s decision to pursue the Savannah River Plutonium Processing Facility (SRPPF), the team targeted the identification of events and indicators in retrospective datasets that pointed to the activity of “fissile core fabrication at the Savannah River Site” prior to the official announcement in May of 2018. A preliminary modeling pipeline was developed in FY20 that showed the datasets contained adequate signal for continuation of efforts. In FY21, a modular demonstration prototype modeling pipeline has continued in development for two text-based data sources: a broad internet archive (Webhose Ltd.) and a decahose Twitter database (i.e., a global sampling of one in every ten Tweets). The techniques that have been developed rely on graph theory and anomaly detection to identify contextual shifts in key words and phrases at various points in time such that indicators of events of interest could be identified and subsequently, events could be extracted from the corpuses. The foundational concept behind the approaches is that contextual shifts in key words and phrases can act as indicators of events of interest. Both datasets have proven successful in extracting events of interest related to pit production at the Savannah River Site prior to the official announcement. In addition, the pipelines have generated a wide range of events broadly summarized as: the awarding of DOE contracts at major sites, DOE investments in various programs, accidents at DOE national laboratories, speculations about the fate of pit production in the DOE complex, domestic and international shipments and receipts of nuclear materials at DOE sites, termination of non-proliferation agreements with Russia, termination of MOX, new weapons development approvals/testing, nuclear posture reviews, major DOE cleanup/production milestones, political opinions, and nuclear watch groups’ opinions, among many others.

97 MATHEMATICS AND COMPUTING↗