Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Protocol for applying a network-enabled gene discovery pipeline to non-model plant species

Identifying upstream regulators of key genes is essential for understanding gene regulatory mechanisms and translating these insights into functional targets. Here, we present a protocol for applying the network-enabled gene discovery pipeline (NEEDLE) to non-model plant species. We describe steps for environment setup, data preparation, computational analysis, expected outputs, and parameter considerations. NEEDLE integrates RNA sequencing (RNA-seq) processing, weighted gene co-expression analysis (WGCNA), Gene Network Inference with Ensemble of trees (GENIE3), and promoter conservation analysis to prioritize candidate transcriptional regulators.

Plant Sciences↗

Geospatial Data Workflow Orchestration and Architecture

In an era characterized by explosive growth in geospatial data, the selection of appropriate technologies for data storage, processing, and orchestration is critical for organizations aiming to maintain competitive advantages. This white paper provides a comprehensive analysis of how Oak Ridge National Laboratory (ORNL) has effectively employed various cloud technologies, including containerized applications, container orchestrators, and workflow orchestrators, to develop robust geospatial data processing solutions. We explore the fundamental concepts behind these technologies and compare multiple deployment models tailored to diverse use cases. Our findings conclude that while Kubernetes has emerged as the preferred platform for truly scalable and fault-tolerant production workflows, the choice of workflow orchestration tool requires careful consideration of team needs, pipeline complexity, and deployment environments. This paper aims to serve as a strategic guide for organizations leveraging geospatial data, articulating the balance between technology choices and practical implementation to enhance workflow efficacy and scalability.

97 MATHEMATICS AND COMPUTING↗

Event Classifications on DNE2 Main Experiment Data using a Convolutional Neural Network Ensemble

The Dynamic Networks (DN) Experiment for FY24 (DNE2) is an experiment within DN with the goal of quantitatively evaluating the effectiveness of solutions developed so far by various researchers under the Low Yield Nuclear Monitoring (LYNM) program using a shared set of metrics and datasets. A key component of this experiment is the mimicking of a signature processing pipeline, and comparing currently accepted and standard-use processing methods to more state-of-the-art processes developed under DN. In this work, we focus specifically on the Event Characterization (EC) Focus Area (FA) of the pipeline, where a seismic event’s magnitude, yield and class are identified. We use Deep Learning (DL) to classify the type of events being processed as either earthquakes (EQs) or explosions (EXs) for three iterations of experiment datasets. The model is noticeably more confident and accurate in classifying explosions than earthquakes, reflecting a known shortcoming of the model, that being of a bias towards predicting explosions over earthquakes in the west coast due to training data biases.

97 MATHEMATICS AND COMPUTING↗

Segmentation Model Distillation [Poster]

The process of training object detection (OD) or image segmentation model requires both a substantial amount of data and technical knowledge, which often creates challenges in applying these types of models to their full potential. In order to streamline the process of developing these models, we propose a new pipeline where a foundation model assists in the dataset generation. Then this resulting dataset is used to fine-tune a fast light-weight model to perform the custom segmentation or OD. This resulting model is also fit for real-time image segmentation, such as in a video stream.

97 MATHEMATICS AND COMPUTING↗

Exploring DAOS as a Burst Buffer for a 100 Gbps DAQ Real-Time Streaming System

We present an experimental evaluation of a burst buffer for a real-time DAQ streaming system designed to transmit instrument data to remote data centers. The system is based on EJ-FAT, a load balancing system capable of Nx 100Gbps streams, distributing data from event sources to processing nodes. We explore applying the DAOS system as a burst buffer to serve a number of purposes: improve resiliency, elasticity and add new functions into the processing pipeline. In the evaluation a sender transmits events over a 100Gbps network to a receiver integrated with DAOS to store the reassembled events using DAOS APIs. We evaluate the system for possible bottlenecks and provide end-to-end evaluation with a burst buffer using DAOS storage abstractions. We show that a receiver node can support 38.1 Gbps. This proves the viability of our approach and allows us to extend this work to investigate scale-out properties and new streaming optimizations.

Mei, Xinxin↗

Advanced Distributed Optical Fiber Sensor Systems for Pipeline Integrity Monitoring

Distributed fiber optic sensors allow the measurement of structural parameters such as static/dynamic strain, temperature, pressure, and vibrations at thousands of locations along a single fiber cable. Deep neural network (DNN) algorithms were developed for rapid data processing speed and vibration event classification.

Lalam, Nageswara↗

A Radiation-Hard 8-Channel 15-Bit 40-MSPS ADC for the ATLAS Liquid Argon Calorimeter Readout

The custom design of a radiation-hardened, 8-channel, 40-MSPS, 15-bit resolution, 14.2-bit dynamic range, 11.4-ENOB ADC data acquisition ASIC fabricated in a commercial 65-nm triple-well CMOS technology is presented. The ADC is developed for and integrates seamlessly into the readout system for the ATLAS liquid argon (LAr) calorimeter in the high-luminosity large hadron collider (HLLHC) upgrade at CERN, which will require a total of 364 936 ADC channels. A three-stage MDAC+SAR pipelined ADC architecture was designed to meet the physics requirements and scientific goals of the ATLAS experiment. The ADC is a fully self-contained data acquisition system that includes foreground calibration, digital data processing, digital control, and supporting circuitry. The measured performance shows the ADC achieves a competitive dynamic range and SNDR, and it meets or exceeds the ATLAS analog requirements. Radiation tolerance and scalability design considerations were implemented at the device-, circuit-, and system-level. Radiation-hardening-by-design techniques used include redundancy for digital circuits, the use of MiM capacitors, and a hybrid RC-DAC for the ADC core. The ADC ASIC was demonstrated to be robust against the effects of the intense radiation expected in the HL-LHC experimental environment.

DAQ↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Active learning enables generation of molecules that advance the known Pareto front

Although generative models hold promise for discovering molecules with optimized desired properties, they often fail to suggest synthesizable molecules that improve upon the properties of the structures represented in the training distribution. We find that this limitation arises not only from the molecule generation process itself, but also from the poor generalization capabilities of molecular property predictors. We address this challenge by creating a closed-loop molecule generation pipeline with iterative retraining on new quantum chemical simulation data. Compared against static, single-pass generative modeling approaches, only our closed-loop iterative workflow generates molecules with properties extending beyond the training distribution (up to 0.44 standard deviations beyond the original range) and achieves a 79% improvement in out-of-distribution molecule classification accuracy. Furthermore, by conditioning molecular generation on thermodynamic stability data obtained during the iterative loop, the proportion of stable and hence potentially synthesizable molecules generated is 3.5x higher than the next-best model.

Chemistry↗

Fiducial-cosmology-dependent systematics for the DESI 2024 BAO analysis

When measuring the Baryon Acoustic Oscillations (BAO) scale from galaxy surveys, one typically assumes a fiducial cosmology when converting redshift measurements into comoving distances and also when defining input parameters for the reconstruction algorithm. A parameterised template for the model to be fitted is also created based on a (possibly different) fiducial cosmology. This model reliance can be considered a form of data compression, and the data is then analysed allowing that the true answer is different from the fiducial cosmology assumed. In this study, we evaluate the impact of the fiducial cosmology assumed in the BAO analysis of the Dark Energy Spectroscopic Instrument (DESI) survey Data Release 1 (DR1) on the final measurements in DESI 2024 III. We utilise a suite of mock galaxy catalogues with survey realism that mirrors the DESI DR1 tracers: the bright galaxy sample (BGS), the luminous red galaxies (LRG), the emission line galaxies (ELG) and the quasars (QSO), spanning a redshift range from 0.1 to 2.1. We compare the four secondary AbacusSummit cosmologies against DESI's fiducial cosmology (Planck 2018). The secondary cosmologies explored include a lower cold dark matter density, a thawing dark energy universe, a higher number of effective species, and a lower amplitude of matter clustering. The mocks are processed through the BAO pipeline by consistently iterating the grid, template, and reconstruction reference cosmologies. We determine a conservative systematic contribution to the error of 0.1% for both the isotropic and anisotropic dilation parameters αiso and αAP. We then directly test the impact of the fiducial cosmology on DESI DR1 data.

79 ASTRONOMY AND ASTROPHYSICS↗

A miniaturized feedstocks-to-fuels pipeline for screening the efficiency of deconstruction and microbial conversion of lignocellulosic biomass

Sustainably grown biomass is a promising alternative to produce fuels and chemicals and reduce the dependency on fossil energy sources. However, the efficient conversion of lignocellulosic biomass into biofuels and bioproducts often requires extensive testing of components and reaction conditions used in the pretreatment, saccharification, and bioconversion steps. This restriction can result in a significant and unwieldy number of combinations of biomass types, solvents, microbial strains, and operational parameters that need to be characterized, turning these efforts into a daunting and time-consuming task. Here we developed a high-throughput feedstocks-to-fuels screening platform to address these challenges. The result is a miniaturized semi-automated platform that leverages the capabilities of a solid handling robot, a liquid handling robot, analytical instruments, and a centralized data repository, adapted to operate as an ionic-liquid-based biomass conversion pipeline. The pipeline was tested by using sorghum as feedstock, the biocompatible ionic liquid cholinium phosphate as pretreatment solvent, a “one-pot” process configuration that does not require ionic liquid removal after pretreatment, and an engineered strain of the yeast Rhodosporidium toruloides that produces the jet-fuel precursor bisabolene as a conversion microbe. By the simultaneous processing of 48 samples, we show that this configuration and reaction conditions result in sugar yields (~70%) and bisabolene titers (~1500 mg/L) that are comparable to the efficiencies observed at larger scales but require only a fraction of the time. We expect that this Feedstocks-to-Fuels pipeline will become an effective tool to screen thousands of bioenergy crop and feedstock samples and assist process optimization efforts and the development of predictive deconstruction approaches.

09 BIOMASS FUELS↗

BLADE: An Automated Framework for Classifying Light Curves from the Center for Near-Earth Object Studies Fireball Database

Fireballs (bolides) are high-energy luminous phenomena produced when meteoroids and small asteroids enter Earth’s atmosphere at hypersonic speeds, often resulting in fragmentation or complete disintegration accompanied by significant energy release. The resulting bolide light curves capture temporal brightness variations as these objects traverse increasingly dense atmospheric layers, providing essential information on meteoroid entry dynamics, fragmentation behavior, and atmospheric energy deposition processes. The Center for Near-Earth Object Studies’ (CNEOS) continuously expanding fireball database offers a globally comprehensive archive of bolide events, including light curves and associated metadata. Events associated with infrasound detections allow direct correlations between acoustic signatures and light curve features, therefore enabling detailed analyses of fragmentation dynamics and energy deposition. Here, we introduce Bolide Light-curve Analysis and Discrimination Explorer (BLADE), a robust and high-fidelity framework specifically designed to analyze bolide light curves for objects detected from space. BLADE incorporates a processing pipeline integrating Savitzky–Golay filtering, prominence-based peak detection, and gradient analysis, enabling systematic identification and classification of fragmentation events and their associated energy release characteristics. Preliminary results demonstrate that BLADE reliably distinguishes distinct bolide behaviors, providing an objective, scalable methodology for characterization and analysis of large bolide light curve data sets. This foundational work establishes a novel pathway for advanced bolide research, with promising applications in planetary defense and global atmospheric monitoring. Future research should adopt an integrative approach combining CNEOS optical data with complementary infrasound measurements, further clarifying relationships between bolide energy deposition and acoustic signatures, thus refining our understanding of meteoroid and asteroid atmospheric entry processes.

Asteroids↗

NanoPSD: A software for automatic detection of Nano-Particle Shape Distribution in electron microscopy images

Accurate quantification of the size and morphology of nanoparticles from electron microscopy (EM) images is essential to understand growth mechanisms, surface reactivity, and functional behavior in nanoscale materials. Manual analysis remains slow, subjective, and difficult to reproduce in large datasets. We introduce NanoPSD (Nano-Particle Shape Distribution), an open-source and fully automated framework for quantitative particle detection and morphology analysis from EM images. NanoPSD integrates adaptive contrast enhancement, polarity-agnostic scale-bar detection, Optical Character Recognition (OCR)-based calibration, and classical segmentation via Otsu thresholding with morphological refinement. Particle contours are used to extract geometric descriptors, including equivalent circular diameter, aspect ratio, circularity, and solidity, enabling automated classification into spherical, rod-like, and aggregate morphologies. The framework supports both single-image and batch processing, generating publication-quality visualizations, LaTeX-ready tables, and structured comma-separated values (CSV) datasets. As a demonstration, we applied NanoPSD to plasma-synthesized nanoparticle samples diagnosed via transmission electron microscopy (TEM). The code produced statistically robust size and morphology distributions spanning a few to tens of nanometers with minimal user supervision. The pipeline demonstrates high reproducibility and scalability, processing large image collections with consistent calibration and output formatting. Its modular design enables seamless integration of future deep-learning-based segmentation models, providing a pathway toward intelligent, data-driven electron microscopy analysis.

36 MATERIALS SCIENCE↗

OpenSAMPL: An Open Source Library for Timing and Synchronization Measurements and Analytics

Today's power grid operators are implementing timing and synchronization solutions that provide resilience to Global Navigation Satellite System (GNSS) vulnerabilities. These vendor-specific solutions often come with additional software applications that are designed to monitor that vendor's synchronization performance data. However, resilient timing architectures often resulting in multi-vendor solutions, including approaches that blend terrestrial clocks with space-based subscription services. In such an environment, collecting, analyzing, and visualizing data from a variety of sources within a single platform was heretofore not possible. To address this need, the US Department of Energy's Center for Alternative Synchronization and Timing (CAST) developed OpenSAMPL, the Open Synchronized Analytics and Monitoring Platform, an open-source Python framework for processing, loading, and observing clock measurement data from distributed devices. OpenSAMPL enables the ingestion of diverse clock-probe sources into a scalable time-series database and applies robust analytics. OpenSAMPL currently supports two vendor data pipelines, and will be extended to more in the near future, enabling seamless monitoring of a variety of timing and synchronization devices in a common environment.

Grant, Josh [ORNL] (ORCID:0000000163475060)↗

Data for "Enhancing Lipid Production in Plant Cells through Automated High-Throughput Genome Engineering and Phenotyping"

Plant bioengineering is a time-consuming and labor-intensive process with no guarantee of achieving desired traits. Here, we present a fast, automated, scalable, high-throughput pipeline for plant bioengineering (FAST-PB) in maize (Zea mays) and Nicotiana benthamiana. FAST-PB enables genome editing and product characterization by integrating automated biofoundry engineering of callus and protoplast cells with single-cell matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). We first demonstrated that FAST-PB could streamline Golden Gate cloning, with the capacity to construct 96 vectors in parallel. Using FAST-PB in protoplasts, we found that PEG2050 increased transfection efficiency by over 45%. For proof-of-concept, we established a reporter-gene-free method for CRISPR editing and phenotyping via mutation of high chlorophyll fluorescence 136. We show that diverse lipids were enhanced up to 6-fold using CRISPR activation of lipid controlling genes. In callus cells, an automated transformation platform was employed to regenerate plants with enhanced lipid traits through introducing multigene cassettes. Lastly, FAST-PB enabled high-throughput single-cell lipid profiling by integrating MALDI-MS with the biofoundry, protoplast, and callus cells, differentiating engineered and unengineered cells using single-cell lipidomics. These innovations massively increase the throughput of synthetic biology, genome editing, and metabolic engineering and change what is possible using single-cell metabolomics in plants.

AI/ML↗

Dark Energy Survey Year 6 Results: Synthetic-source Injection Across the Full Survey Using Balrog

Synthetic source injection (SSI), the insertion of sources into pixel-level on-sky images, is a powerful method for characterizing object detection and measurement in wide-field, astronomical imaging surveys. Within the Dark Energy Survey (DES), SSI plays a critical role in characterizing all necessary algorithms used in converting images to catalogs, and in deriving quantities needed for the cosmology analysis, such as object detection rates, galaxy redshift estimation, galaxy magnification, star-galaxy classification, and photometric performance. We present here a source injection catalog of 146 million injections spanning the entire 5000 deg 2 DES footprint, generated using the Balrog SSI pipeline. Through this SSI sample, we demonstrate that the DES Year 6 (Y6) image processing pipeline provides accurate estimates of the object properties, for both galaxies and stars, at the percent-level, and we highlight specific regimes where the accuracy is reduced. We then show the consistency between SSI and data catalogs, for all galaxy samples developed within the weak lensing and galaxy clustering analyses of DES Y6. The consistency between the two catalogs also extends to their correlations with survey observing properties (seeing, airmass, depth, extinction, etc.). Lastly, we highlight a number of applications of this catalog to the DES Y6 cosmology analysis, such as estimates of the redshift distribution and lens magnification. This dataset is the largest SSI catalog produced at this fidelity and will serve as a key testing ground for exploring the utility of SSI catalogs in upcoming surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time.

79 ASTRONOMY AND ASTROPHYSICS↗