Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Accelerating the identification of novel secondary metabolites in bioenergy plant root exudates using MicroED

Small molecule metabolites drive inter- and intraspecies communication and dependencies in diverse biological systems, yet a large proportion of these important chemical compounds remain uncharacterized in plants and microbes. Approximately 90% of the metabolites in root exudate profiles are unknown compounds, despite the importance of root exudate composition in plant-microbe interactions. We need advanced analytical capabilities that will support rapid discovery and structural elucidation of metabolites from biological samples that may be limited in quantity and high in complexity. To fill this gap, this project aimed to develop an integrated workflow involving metabolite extraction, separation, and crystallization from plant root exudates followed by characterization using nuclear magnetic resonance (NMR) spectroscopy, mass spectrometry, and microcrystal electron diffraction (MicroED). Using crude root exudates from sorghum, this project successfully developed higher throughput exudate fractionation strategies to obtain pure compounds for crystallization and identified crystals in multiple fractions that diffracted. Additional efforts to increase the throughput of high-quality crystal generation for MicroED, such as crystallization screening and crystallization chaperone exploration, will be needed to further advance root exudate metabolite identification. The overall optimized sample preparation process can then be integrated with the existing data collection and data analysis pipelines for MicroED at PNNL to facilitate more rapid natural product discovery.

59 BASIC BIOLOGICAL SCIENCES↗

Chlamydomonas reinhardtii responses to Fe-excess, Fe-deficiency, and Fe-limitation in either photoautotrophic or mixotrophic growth

A systems level analysis of Chlamydomonas reinhardtii grown photoautotrophically or mixotrophically with a reduced carbon source, acetate, under four different defined Fe stages of Fe-replete, Fe-deficient, Fe-limited, or Fe-excess. Samples were digested with trypsin, labeled with TMT 10-Plex, then analyzed by LC-MS/MS. Data was searched with MS-GF+ using PNNL's DMS Processing pipeline. [doi:10.25345/C5707X12X] [dataset license: CC0 1.0 Universal (CC0 1.0)]

59 BASIC BIOLOGICAL SCIENCES↗

Iron-starvation induces photosystem I antenna remodeling in green algae

Dunaliella salina and Dunaliella tertiolecta are extremophile, marine algae that can survive in very low Fe conditions. In this study, we used TMT-proteomics to compare the Fe starvation responses to the Fe replete responses. Samples were digested with trypsin, labeled with TMT 10-Plex, then analyzed by LC-MS/MS. Data was searched with MS-GF+ using PNNL's DMS Processing pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Adapt: A Weather Radar Data Analysis and Nowcasting Platform for Informed Adaptive Scanning

SF-26-021 Adapt is a data processing platform for real-time data analysis, short term prediction of targets convective cells and tracking for archived data. It provides tools for downloading, processing, segmenting, projecting, analyzing, and visualizing storm cell data from weather radar. The pipeline includes cell detection, motion estimation using optical flow, cell property extraction, and persistence to NetCDF and SQLite/Parquet for guiding adaptive scanning.

Raut, Bhupendra Ashokrao [Argonne National Laborat↗

Long-term measurements of ice nucleating particles at Atmospheric Radiation Measurement (ARM) sites worldwide

Ice nucleating particles (INPs) play a critical role in cloud microphysics and precipitation formation, yet long-term, spatially extensive observational datasets remain limited. Here, we present one of the most comprehensive publicly available datasets of immersion-mode INP concentrations using a single analytical method, generated through the U.S. Department of Energy's (DOE) Atmospheric Radiation Measurement (ARM) user facility. INP filter samples have been collected across a broad range of environments – including agricultural plains, Arctic coastlines, high-elevation mountain sites, marine regions, and urban areas – via fixed observatories, mobile facility deployments, and vertically-resolved tethered balloon system operations. We describe the standardized processing and quality assurance pipeline, from filter collection and processing using the Ice Nucleation Spectrometer to final data products archived on the ARM Data Discovery portal. The dataset includes both total INP concentrations and selectively treated samples, allowing for classification of biological, organic, and inorganic INP types. It features a continuous 5-year record of INP measurements from a central U.S. site, with data collection still ongoing. Seasonal and site-specific differences in INP concentrations are illustrated through intercomparisons at −10 and −20 °C, revealing distinct regional sources and atmospheric drivers. We also outline mechanisms for researchers to access existing data, request additional sample analyses, and propose future field campaigns involving ARM INP measurements. This dataset supports a wide range of scientific applications, from observational and mechanistic studies to model development, and provides critical constraints on aerosol-cloud interactions across diverse atmospheric regimes (Creamean et al., 2024, 2020b; https://doi.org/10.5439/1770816).

Creamean, Jessie M. [Colorado State Univ., Fort Co↗

Selection of a Pair of Experiments to Optimally Reduce Uncertainty in Targeted Nuclear Data

We propose a novel process to select a pair of differential and integral experiments that best reduce uncertainties in targeted 239 ⁢Pu nuclear data while compressing the current nuclear data pipeline from 20 to 3 years. 239⁢ Pu nuclear data are poorly understood for neutrons in the intermediate energy range due to sparsity and uncertainty in historical experiments. New experiments targeting this range will enable better understanding of these nuclear data, but choosing the ideal experiments to conduct is challenging. Beginning with a prior distribution represented by samples of nuclear data generated from theory, generalized least squares adjustments are made to incorporate data from historical experiments. To quantify potential uncertainty reduction obtainable from a pair of candidate experiments, we compute the D-optimality criterion of the posterior covariance of intermediate energy range nuclear data compared to the equivalent covariance after additional adjustment to the pair of candidate experiments. Repeating the process for each of many candidate pairs facilitates the final selection. Results support 63⁢ Cu total cross section measurements for differential experiments and alumina and alumina/graphite configurations for integral experiments. This analysis enables choosing differential and integral experiments to be executed concurrently while shortening decision times relative to the current nuclear data pipeline.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Fusion of Experiments and Simulations for Real-Time Identification of Pipeline Defects

In this study, we explored fusion of experiments and simulations for real time identification of pipeline defects across physical and non-physical domains. The challenges associated to data processing were addressed and a combined classification models was presented via CNN models. In addition, regression model based on XGBOOST is built to determine the defect location and defect dimension from data-driven features of guided wave signals captured by SMS fiber optic sensor.

deep learning↗

Virtual Inspection of Advanced Manufacturing via Process-Scale Digital Twins (Abbreviated Report)

Inspection and certification comprise the most significant bottlenecks in advanced manufacturing for NNSA applications, often requiring far more time and resources than the fabrication of the parts themselves. Traditional methods, such as manual review and X-ray computed tomography, are not only slow and costly, but also struggle to provide a clear connection between manufacturing instructions and the final performance of critical components. This gap limits both the agility and assurance needed to support the modernization and safety of the United States nuclear stockpile. In response, our Strategic Initiative established a digital twin framework that integrates realtime process monitoring, automated data analysis, and immersive virtual reality collaboration into a unified inspection pipeline. By leveraging data from sensors, machine instructions, and imaging, we created high-fidelity virtual models of manufactured parts that could be rapidly analyzed and certified. This approach was first demonstrated with Direct Ink Write, and then extended to other manufacturing settings, including conventional (or “subtractive”) manufacturing and to predict the end of life performance of parts per the aging and lifetimes programs. The result is a transformational capability: inspection times have been reduced by a factor of 120,000 without loss of accuracy and while simultaneously improving traceability and confidence in part quality. This framework not only streamlines certification for critical applications, but also positions the national security enterprise to respond more flexibly to emerging challenges, supporting agile manufacturing and digital engineering practices across a broad range of mission-relevant domains.

42 ENGINEERING↗

Fusion of Experiments and Simulations for Real-Time Identification of Pipeline Defects

In this study, we explored fusion of experiments and simulations for real time identification of pipeline defects across physical and non-physical domains. The challenges associated to data processing were addressed and a combined classification models was presented via CNN models. In addition, regression model based on XGBOOST is built to determine the defect location and defect dimension from data-driven features of guided wave signals captured by SMS fiber optic sensor.

deep learning↗

Laminography as a tool for imaging large-size samples with high resolution

Despite the increased brilliance of the new generation synchrotron sources, there is still a challenge with high-resolution scanning of very thick and absorbing samples, such as a whole mouse brain stained with heavy elements, and, extending further, brains of primates. Samples are typically cut into smaller parts, to ensure a sufficient X-ray transmission, and scanned separately. Compared with the standard tomography setup where the sample would be cut into many pillars, the laminographic geometry operates with slab-shaped sections significantly reducing the number of sample parts to be prepared, the cutting damage and data stitching problems. In this work, a laminography pipeline for imaging large samples (>1 cm) at micrometre resolution is presented. The implementation includes a low-cost instrument setup installed at the 2-BM micro-CT beamline of the Advanced Photon Source. Additionally, sample mounting, scanning techniques, data stitching procedures, a fast reconstruction algorithm with low computational complexity, and accelerated reconstruction on multi-GPU systems for processing large-scale datasets are presented. The applicability of the whole laminography pipeline was demonstrated by imaging four sequential slabs throughout an entire mouse brain sample stained with osmium, in total generating approximately 12 TB of raw data for reconstruction.

47 OTHER INSTRUMENTATION↗

The Future of a Myriad of Accelerated Biodiscoveries Lies in AI‐Powered Mass Spectrometry and Multiomics Integration

The intersection of modern artificial intelligence (AI) and mass spectrometry (MS) is set to transform the MS‐based “omics” research fields, particularly proteomics, metabolomics, lipidomics, and glycomics, enabling advancements across a wide range of domains, from health to environment and industrial biotechnology. Beginning with an overview of key challenges inherent in MS software pipelines, this personal perspective explores how AI‐driven solutions can address them to enhance data processing, integration and interpretation. It proposes a paradigm shift in molecular identification and quantitation algorithms, leveraging AI to enable holistic interpretation of MS‐based multiomics data. While centered on MS‐based omics, this holistic AI‐driven paradigm is also critical for connecting dynamic biochemical changes to genomics and transcriptomics contexts, reinforcing the integrative value of MS in multiomics research. Ultimately, this AI‐driven approach could enhance efficiency, accuracy, and molecular breadth of coverage, deepening our systems‐level understanding of biological processes and accelerating a myriad of biodiscoveries.

47 OTHER INSTRUMENTATION↗

Integrative SP3 Workflow for Multi-PTM Proteomics Profiling (TZ-DP0)

The goal of the experiment was to demonstrate that the optimized multiplexed multi-PTM profiling workflow can comprehensively and quantitatively capture dynamic changes in protein abundance, cysteine oxidation, phosphorylation, and acetylation in cytokine-induced inflammatory stress in mouse pancreatic ß-cells. Global proteomic, redox proteomic, phosphoproteomic, and acetylomic were data collected from mouse Beta-TC-6 pancreatic Beta-cells, untreated (mock) and cytokine-treated Beta-cells at 4, 8, and 24 hours with 4 biological replicates. Samples were digested with trypsin and Lys-C, then analyzed by LC-MS/MS. Data were searched with MS-GF+, MASIC, and MaxQuant using PNNL's DMS processing pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

DMTN-277: The Monster: A reference catalog with synthetic ugrizy-band fluxes for the Vera C. Rubin observatory

In order to facilitate bootstrap photometric calibrations of early Rubin Observatory data we have created an all sky reference catalog called The Monster. This reference catalog uses a rank-ordered set of other reference catalogs to generate synthetic ugrizy-band fluxes that can be used calibrate images processed with the LSST science pipelines. This document describes the methodology used to create The Monster, documents the input external reference catalogs, and performs basic data validation of the first version of The Monster.

79 ASTRONOMY AND ASTROPHYSICS↗