Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Redis-Based Streaming Architecture for Accelerator Beam Instrumentation DAQ Systems

The Fermilab Acceleraor Division, Beam Instrumentation Department, is always adopting modern and current software methodologies for complex DAQ architectures. This paper highlights the Redis Adapter (RA) as the key software component enabling high performance, modular communication between digitizers and distributed control systems by leveraging Redis and containerization. The RA provides a unified, efficient interface between Redis based data streams and consumer systems. In the legacy architecture, digitized data flowed through the custom, UDP based Distributed Data Communication Protocol in the middle layer. In the current system, DDCP remains the ingestion path, while the RA serves as the decoupling layer. The proposed system replaces old VME digitizers with a SOM-based digitizer that communicates with Redis using the RA. The RA acts as both a performance-critical bridge and a protocol-agnostic adapter, ensuring compatibility with legacy control frameworks while enabling future scalability and modularity. This restructuring of the middle layer also helps the system achieve high throughput, reduce latency, and simplify the data path. Finally, we will demonstrate how RA is utilized in our two core products to deliver both legacy compatibility and future flexibility.

Joshi, S. [Fermilab]↗

Bridging microscopy with molecular dynamics and quantum simulations: an atomAI based pipeline

Recent advances in (scanning) transmission electron microscopy have enabled a routine generation of large volumes of high-veracity structural data on 2D and 3D materials, naturally offering the challenge of using these as starting inputs for atomistic simulations. In this fashion, the theory will address experimentally emerging structures, as opposed to the full range of theoretically possible atomic configurations. However, this challenge is highly nontrivial due to the extreme disparity between intrinsic timescales accessible to modern simulations and microscopy, as well as latencies of microscopy and simulations per se. Addressing this issue requires as a first step bridging the instrumental data flow and physics-based simulation environment, to enable the selection of regions of interest and exploring them using physical simulations. Here we report the development of the machine learning workflow that directly bridges the instrument data stream into Python-based molecular dynamics and density functional theory environments using pre-trained neural networks to convert imaging data to physical descriptors. Additionally, the pathways to ensure structural stability and compensate for the observational biases universally present in the data are identified in the workflow. This approach is used for a graphene system to reconstruct optimized geometry and simulate temperature-dependent dynamics including adsorption of Cr as an ad-atom and graphene healing effects. However, it is universal and can be used for other material systems.

36 MATERIALS SCIENCE↗

Imputation of urban environmental sensor data using gated attention bidirectional long short-term memory (GA-BiLSTM): methods, performance, and implications

Urban environmental monitoring networks frequently encounter significant data gaps due to sensor malfunctions, environmental disturbances, and communication failures. Reliable approaches to address these gaps are essential for ensuring the continuity and quality of environmental data streams. In this study, we developed a gated attention bidirectional long short-term memory (GA-BiLSTM) model to impute missing data in a dense urban monitoring network. Using observations from the CROCUS network in Chicago, we evaluated GA-BiLSTM against widely used approaches (XGBoost and K-nearest neighbors) under scenarios of both short-term intermittent gaps and prolonged outages. GA-BiLSTM consistently outperformed comparative methods, particularly during extended outages of up to ten days, demonstrating its ability to capture spatiotemporal dependencies across sensor nodes. Beyond performance metrics, feature importance and spatial network analyses highlighted the unexpected but critical predictive role of peripheral rural nodes, underlining their strategic value for maintaining robust urban monitoring systems. These results emphasize that advanced imputation methods can substantially improve the reliability of environmental monitoring networks and support more resilient data infrastructures for urban sustainability.

Data imputation↗

Enhancements and Deployment of the TDAQ System for the Mu2e Experiment

The Real Time Processing Systems Division at Fermilab has deployed new features to the Off-The-Shelf Data Acquisition framework (otsdaq) for the Mu2e experiment. The Mu2e experiment will search for the coherent neutrino-less conversion of a muon into an electron in the field of an aluminum nucleus with a sensitivity improvement of 10,000 times over existing limits. Such a charged lepton flavor-violating reaction probes new physics at a scale unavailable at present or planned high-energy colliders. The Mu2e Trigger and Data Acquisition (TDAQ) system uses otsdaq as its online Data Acquisition System (DAQ) framework. otsdaq integrates the artdaq and art frameworks for event transfer, filtering, and processing. otsdaq is a web-based DAQ software suite focusing on flexibility and scalability and provides a multi-user interface accessible through a web browser. artdaq handles the entire data stream, which is read over the peripheral component interconnect express (PCIe) bus to a software filter algorithm that selects events combined with the data flux coming from a cosmic-ray veto (CRV) system. Detector front-ends are configured through the PCIe bus by customized otsdaq plugins. The otsdaq slow controls infrastructure has been further developed using the experimental physics and industrial control system (EPICS) open-source platform for monitoring, controlling, alarming, and archiving. The detector control system (DCS) for Mu2e has been integrated into otsdaq. The production TDAQ and DCS system has been deployed at the experimental hall and is being debugged and optimized for experiment operations. We report on the feature enhancements and deployment of otsdaq for Mu2e.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Recommendations for Updating Liquid Discharged Inventory and Transport Modeling Parameters for Cumulative Impacts Evaluation of Hexavalent Chromium in the 200 West Area

The purpose of this environmental calculation file (ECF) is to document information regarding hexavalent chromium (Cr(VI)) inventory discharged in 200 West Area at the Hanford Site and provide data to support predictive transport through the vadose zone and saturated zone for modeling efforts. This document provides a focused evaluation of historical waste stream data and studies to develop estimates of Cr(VI) inventory, discharge fractions, and transport parameters for the 200 West Area waste sites and tank farms during discharge events and for long-term contaminant releases.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Materials Characterization, Prediction and Control Project: Summary Report on Data Analytics Framework

This report summarizes the activities performed under the data analytics Vertex in the Materials Characterization, Prediction and Control Project funded under laboratory directed research and development at Pacific Northwest National Laboratory. The data analytics Vertex developed models for associating global or local process parameters, microstructural features, and performance properties of friction-stir-processed 316L stainless steel plates. Statistical, machine learning, and deep learning models, as well as generative artificial intelligence approaches, were used to develop the associations between the process-structure-property data streams. These associations formed the basis for predicting global properties of parts manufactured under different process envelopes, providing a basis for predicting performance using data driven as well as physics-informed and physics-constrained approaches. Additionally, the associations were used to predict local process parameters and microstructural features of the product, predictive relationships that have the potential to form the basis of a control framework that could eventually modulate a friction-stir process to maintain product quality.

316L stainless steel↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

Contextual Active Online Model Selection with Expert Advice

How can we collect the most useful labels to learn a model selection policy, when presented with arbitrary heterogeneous data streams? In this paper, we formulate this task as a contextual active model selection problem, where at each round the learner receives an unlabeled data point along with a context. The goal is to output the best model for any given context without obtaining an excessive amount of labels. In particular, we focus on the task of selecting pre-trained classifiers, and propose a contextual active model selection algorithm (CAMS), which relies on a novel uncertainty sampling query criterion defined on a given policy class for adaptive model selection. In comparison to prior art, our algorithm does not assume a globally optimal model. We provide rigorous theoretical analysis for the regret and query complexity under both adversarial and stochastic settings. Our experiments on several benchmark classification datasets demonstrate the algorithm’s effectiveness in terms of both regret and query complexity. Notably, to achieve the same accuracy, CAMS incurs less than 10% of the label cost when compared to the best online model selection baselines on CIFAR10.

Liu, Xuefeng↗

The Phase-2 Upgrade of the CMS Data Acquisition

The High Luminosity LHC (HL-LHC) will start operating in 2027 after the third Long Shutdown (LS3), and is designed to provide an ultimate instantaneous luminosity of 7:5 × 10$^{34}$ cm$^{-2}$ s$^{-1}$, at the price of extreme pileup of up to 200 interactions per crossing. The number of overlapping interactions in HL-LHC collisions, their density, and the resulting intense radiation environment, warrant an almost complete upgrade of the CMS detector. The upgraded CMS detector will be read out by approximately fifty thousand highspeed front-end optical links at an unprecedented data rate of up to 80 Tb/s, for an average expected total event size of approximately 8 - 10 MB. Following the present established design, the CMS trigger and data acquisition system will continue to feature two trigger levels, with only one synchronous hardware-based Level-1 Trigger (L1), consisting of custom electronic boards and operating on dedicated data streams, and a second level, the High Level Trigger (HLT), using software algorithms running asynchronously on standard processors and making use of the full detector data to select events for offline storage and analysis. The upgraded CMS data acquisition system will collect data fragments for Level-1 accepted events from the detector back-end modules at a rate up to 750 kHz, aggregate fragments corresponding to individual Level- 1 accepts into events, and distribute them to the HLT processors where they will be filtered further. Events accepted by the HLT will be stored permanently at a rate of up to 7.5 kHz. This paper describes the baseline design of the DAQ and HLT systems for the Phase-2 of CMS.

Badaro, Gilbert↗

Towards a study protocol: A data-driven workflow to identify error sources in direct ink write mechatronics

Abstract Using Direct Ink Write (DIW) technology in a rapid and large-scale production requires reliable quality control for printed parts. Data streams generated during printing, such as print mechatronics, are massive and diverse which impedes extracting insights. In our study protocol approach, we developed a data-driven workflow to understand the behavior of sensor-measured X- and Y- axes positional errors with process parameters, such as print velocity and velocity control. We uncovered patterns showing that instantaneous changes in the velocity, when the build platform accelerates and decelerates, largely influence the positional errors, especially in the X- axis due to the hardware architecture. Since DIW systems share similar mechatronic inputs and outputs, our study protocol approach is broadly applicable and scalable across multiple systems. Graphical abstract

36 MATERIALS SCIENCE↗

Apparatus and amendment of wind turbine blade impact detection and analysis

A multisensory system provides both temporal and spatial coverage capacities for auto-detection of bird collision events. The system includes an apparatus having a first circuitry to capture and store a series of images or video of a blade of a wind turbine; and a memory to store the images from the first circuitry. The apparatus also has one or more sensors to continuously sense vibration of the blade or for acoustic recordings; and a second circuitry to analyze the sensor data stream and/or the series of images or video to identify a cause of the vibration and to trigger the camera(s). A communication interface transmits data from the second circuitry to another device, wherein the second circuitry applies artificial intelligence or machine learning to control sensitivity of the one or more sensors.

Johnston, Matthew↗

Real-time Anomaly Detection for Liquid Argon Time Projection Chambers

We present a real-time anomaly detection framework for liquid argon time projection chambers (LArTPCs), targeting applications in particle physics experiments such as the Short Baseline Near Detector (SBND) or the future Deep Underground Neutrino Experiment (DUNE). These experiments employ detectors that generate and stream high-resolution but sparse images of neutrino and other particle interactions. Our approach utilizes anomaly detection with autoencoders, compressed through knowledge distillation (KD), to enable the detection of anomalous signals in the data through efficient inference on resource-constrained hardware. The framework is targeted for deployment on computing platforms equipped with field-programmable gate arrays (FPGAs), GPUs, or CPUs, allowing low-latency selection of relevant activity directly from the raw detector data stream. We demonstrate that our approach is suitable for the detection and localization of anomalously "high-multiplicity" activity, and outline promising applications for LArTPC online data filtering and triggering.

FOS: Physical sciences↗

Machine Learning Enabled Sensor Fusion for In-Situ Defect Detection in Laser Powder Bed Fusion

Laser Powder Bed Fusion (L-PBF) Additive Manufacturing (AM) is among the metal 3D printing technologies most broadly adopted by the manufacturing industry. The current industry qualification paradigm for critical-application L-PBF parts relies heavily on expensive non-destructive inspection techniques such as X-Ray Computed Tomography (XCT), which significantly limits the use-cases of L-PBF. In situ monitoring of the process promises a less expensive alternative to ex situ testing, but existing sensor technologies and data analysis techniques struggle to detect sub-surface flaws (e.g., porosity and cracking) on production-scale L-PBF printers. RTX Technologies Research Center (RTRC) has licensed ORNL’s Peregrine software package – a printer- and camera-agnostic data analytics tool designed specifically for detecting process anomalies using in situ data collected during powder bed printing. The goal of this project was to feed temporally rich, multi-modal sensor data, including visible light, integrated near infrared (NIR), and spatially mapped co-axial melt pool thermal emission data into Peregrine to enable detection of subsurface flaws. XCT data was used as ground truth training data to allow Peregrine’s deep learning algorithms to recognize anomalies in these complex data streams in both test artifacts and industrially relevant geometries. Completion of this program has seen the successful implementation of multi-modal, multi-layer sensor data footprints for training of machine learning models in Peregrine. Flaws detected in XCT data have been successfully detected directly from this in situ data footprint, and initial analyses of the in situ probability-of-detection has been conducted, showing performance levels commensurate with traditional non-destructive evaluation (NDE) methods. The in situ monitoring methodology was then applied to an industrially relevant component that was using post-build NDE, highlighting the utility of the proposed method for hard-to-inspect AM components. As a direct result of this program, two journal manuscripts [1], [2] have been published in Additive Manufacturing, with additional manuscripts planned following program completion.

36 MATERIALS SCIENCE↗

Measurement of charged-current $\nu_{\mu}$ and $\nu_e$ interactions with wire-cell in MicroBooNE towards a search for low-energy $\nu_e$ excess

This technote summarizes the existing work in searching for $\nu_e$ low-energy excess (eLEE) in MicroBooNE Booster Neutrino Beam (BNB) data stream based on the Wire-Cell event reconstruction paradigm. The charged-current $\nu_{\mu}$ and $\nu_e$ events are selected from the 5.3e19 POT data from the BNB beam and 2.06e20 POT data from Neutrinos at the Main Injector (NuMI) beam. The charged-current $\nu_e$ selection results from the BNB data that are sensitive to the eLEE search are not included. Various comparisons between data and Monte Carlo predictions are performed to validate the overall model and demonstrate the power of the analysis techniques.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

ESnet/JLab FPGA Accelerated Transport

To increase the science rate for high data rates/volumes, Thomas Jefferson National Accelerator Facility (JLab) has partnered with Energy Sciences Network (ESnet) to define an edge to data center traffic shaping / steering transport capability featuring data event aware network shaping and forwarding. The keystone of this ESnet+JLab FPGA Accelerated Transport (EJFAT) is the joint development of an AI/ML directed dynamic compute work Load Balancer (LB) of UDP streamed data. The LB is a suite consisting of a Field Programmable Gate Array (FPGA) executing the dynamically configurable, low fixed latency LB data plane featuring real-time packet redirection and high throughput, and a control plane running on the FPGA host computer that monitors network and compute farm telemetry in order to make dynamic AI/ML guided decisions for destination compute host redirection/load balancing and destination resource provisioning. The LB provides for three-tier horizontal scaling across LB suites, core compute hosts, and CPUs within a host. The LB effectively provides seamless integration of edge/core computing to support direct experimental data processing for immediate use by JLab science programs and others such as the EIC as well as data centers of the future requiring high throughput and low latency for both hot and cooled data for both running experiment data acquisition systems and data center use cases.

97 MATHEMATICS AND COMPUTING↗

Measurement of Charged-Current $ν_μ$ and $ν_e$ Interactions Towards a Search for Low-Energy $ν_Ε$ Excess Using Wire-Cell In Microboone

This technote summarizes the existing work in searching for $ν_e$ low-energy excess (eLEE) in MicroBooNE Booster Neutrino Beam (BNB) data stream based on the Wire-Cell event reconstruction paradigm. The charged-current $ν_μ$ and $ν_e$ events are selected from the 5.3e19 POT open data from the BNB beam, 6.37e20 POT far sideband from the BNB beam data, and 2.10e20 POT data from Neutrinos at the Main Injector (NuMI) beam. The charged-current $ν_e$ selection results from the BNB data that are sensitive to the eLEE search are not included. Various comparisons between data and Monte Carlo predictions are performed to validate the overall model and demonstrate the power of the analysis techniques. Physics sensitivities in terms of the exclusion and the discovery potential are presented.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Scalable Data-Intensive Geocomputation: A Design for Real-Time Continental Flood Inundation Mapping

The convergence of data-intensive and extreme-scale computing enables an integrated software and data ecosystem for scientific discovery. Developments in this realm will fuel transformative research in data-driven interdisciplinary domains. Geocomputation provides computing paradigms in Geographic Information Systems (GIS) for interactive computing of geographic data, processes, models, and maps. Because GIS is data-driven, the computational scalability of a geocomputation workflow is directly related to the scale of the GIS data layers, their resolution and extent, as well as the velocity of the geo-located data streams to be processed. Unique in high user interactivity and low end-to-end latency requirements, geocomputation applications will dramatically benefit from the convergence of high-end data analytics (HDA) and high-performance computing (HPC). The application level challenge, however, is to identify and eliminate computational bottlenecks that arise along a geocomputation workflow. Indeed, poor scalability at any of the workflow components is detrimental to the entire end-to-end pipeline. Here, we study a large geocomputation use case in flood inundation mapping that handles multiple national-scale geospatial datasets and targets low end-to-end latency. We discuss benefits and challenges for harnessing both HDA and HPC for data-intensive geospatial data processing and intensive numerical modeling of geographic processes. We propose an HDA+HPC geocomputation architecture design that couples HDA (e.g., Spark)-based spatial data handling and HPC-based parallel data modeling. Key techniques for coupling HDA and HPC to bridge the two different software stacks are reviewed and discussed.

Liu, Yan↗