Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

LHC physics dataset for unsupervised New Physics detection at 40 MHz

In the particle detectors at the Large Hadron Collider, hundreds of millions of proton-proton collisions are produced every second. If one could store the whole data stream produced in these collisions, tens of terabytes of data would be written to disk every second. The general-purpose experiments ATLAS and CMS reduce this overwhelming data volume to a sustainable level, by deciding in real-time whether each collision event should be kept for further analysis or be discarded. We introduce a dataset of proton collision events that emulates a typical data stream collected by such a real-time processing system, pre-filtered by requiring the presence of at least one electron or muon. This dataset could be used to develop novel event selection strategies and assess their sensitivity to new phenomena. In particular, we intend to stimulate a community-based effort towards the design of novel algorithms for performing unsupervised new physics detection, customized to fit the bandwidth, latency and computational resource constraints of the real-time event selection system of a typical particle detector.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

WHONDRS Surface Water Geochemistry and Organic Matter Characterization Data from Streams Distributed across Latin America

This dataset supports a broader study examining global transferability of stream biogeochemistry and was generated in collaboration with the MicroSudAqua (µSudAqua) network (https://microsudaqua.netlify.app/en/). The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen, cations) and organic matter characterization (FTICR-MS) from streams in Argentina, Brazil, Chile, and Colombia. Samples were collected across stream orders (1st to 6th order) within five basins. Related data were collected and will be published separately in collaboration with the µSudAqua network. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos; (2) a folder of surface water sample data, (3) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data; (4) file-level metadata; (5) data dictionary; (6) field metadata; (7) readme; (8) international generic sample number (IGSN) mapping file; and (9) field protocol. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total dissolved nitrogen data and averages; (3) anions and averages; (4) methods codes; (5) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4.

Anions↗

Machine Learning Digital Twin for Lithium Ion Battery State of Health Predictions

A digital twin system has been established to model the long term degradation of the state of health of lithium ion batteries. Two data streams result from the computational model of the system and from the physical experiment measurements. A machine learning pipeline has been developed at the nexus of these data streams. Leveraging the unique data sources in multiple transfer learning approaches has lead to the development of multiple cell specific machine learning digital twins. Our discussion on this first of its kind technology will cover challenges in data ingestion, scalability, architecture and orchestration, and research findings.

25 - ENERGY STORAGE↗

915rwpprecipcor

Radar Wind Profiler (RWP) has operational modes: short duration high resolution (referred to as 'high' or 'high mode') and long duration low resolution (referred to as 'low' or 'low mode'). The motivation for this is to create a VAP that quality controls (noise, sig echo filtering), merges, co-grids, and performs initial analyses on these RWP modes. This is done to streamline downstream VAP development on the RWP, as well as enable incorporation of echo properties into the ARSCL chain. The proposed VAP will use the RWPPRECIP b0/b1 level data stream as input. Additional consideration for the incorporation of other ARM data streams to improve this VAP was performed, but not included at this time.

54 ENVIRONMENTAL SCIENCES↗

Geochemistry and Multiomics Data Differentiate Streams in Pennsylvania Based on Unconventional Oil and Gas Activity

Unconventional oil and gas (UOG) extraction is increasing exponentially around the world, as new technological advances have provided cost-effective methods to extract hard-to-reach hydrocarbons. While UOG has increased the energy output of some countries, past research indicates potential impacts in nearby stream ecosystems as measured by geochemical and microbial markers. Here, we utilized a robust data set that combines 16S rRNA gene amplicon sequencing (DNA), metatranscriptomics (RNA), geochemistry, and trace element analyses to establish the impact of UOG activity in 21 sites in northern Pennsylvania. These data were also used to design predictive machine learning models to determine the UOG impact on streams. We identified multiple biomarkers of UOG activity and contributors of antimicrobial resistance within the order Burkholderiales. Furthermore, we identified expressed antimicrobial resistance genes, land coverage, geochemistry, and specific microbes as strong predictors of UOG status. Of the predictive models constructed (n = 30), 15 had accuracies higher than expected by chance and area under the curve values above 0.70. The supervised random forest models with the highest accuracy were constructed with 16S rRNA gene profiles, metatranscriptomics active microbial composition, metatranscriptomics active antimicrobial resistance genes, land coverage, and geochemistry (n = 23). The models identified the most important features within those data sets for classifying UOG status. These findings identified specific shifts in gene presence and expression, as well as geochemical measures, that can be used to build robust models to identify impacts of UOG development.

16S rRNA↗

DEVELOPMENT AND DEMONSTRATION TESTBED FOR THE REMOTE OPERATIONS AND MONITORING OF MICROREACTORS

The nuclear industry is rapidly developing many advanced-reactor concepts for near-term deployment in both traditional and non-traditional nuclear-powered applications. One such category of advanced reactor is the microreactor, a class of reactor with less than 20MWth power output, intended for applications where the economics or logistics of traditional power sources are difficult. This includes applications such as remote communities, mining sites, defense installations, or humanitarian and disaster-relief missions. One key enabling feature for the successful deployment of microreactors is a remote operations capability. Remote operations provide monitoring and control capabilities which can significantly reduce staffing costs by eliminating the need for licensed operators at each reactor facility and improve the economic viability for microreactor deployment. A remote concept of operations is not currently an established capability in the nuclear industry. In addition, no demonstration microreactor is expected to complete construction or go critical until at least 2026. This leaves two major capability gaps: the successful demonstration of a remote concept of operations for microreactors and a test bed suitable for said demonstration. Both gaps must be addressed in order to advance the remote concepts of nuclear operation and, more broadly, microreactors themselves from paper to reality. This paper aims to fill these gaps and describes a test bed that would support development and deployment of a remote concept of nuclear operations, initial experimental results from that test bed, and the application of the test bed and experimental results for a digital-twin-based remote concept of operations underdevelopment at Idaho National Laboratory (INL). The platform chosen as a remote concept of nuclear operations test bed is the Single Primary Heat Extraction and Removal Emulator, known as SPHERE, located at INL. SPHERE is a small-scale non-nuclear test bed that emulates thermal behavior of a microreactor. The small-scale and non-nuclear nature of SHPERE limit safety concerns associated with remote operations while still providing the physical response representative of a microreactor. A network connection was added to SPHERE that enables remote-monitoring capability. This allows for real-time data streaming to networked workstations, data historians, and human-machine interfaces (HMIs). These are all critical components in a remote concept of operations, thus providing a robust development and demonstration platform. An initial experiment was performed using the SPHERE remote operations testbed. This included running a comprehensively instrumented SPHERE through a series of steady-state and transient operating scenarios in both normal and abnormal operating conditions, all while streaming live test data to a remote HMI and data warehouse. This initial experiment served three purposes: (1) characterizing the response of SPHERE, (2) demonstrating the remote connection to SPHERE, and (3) providing a baseline data set for development of a digital-twin-based remote concept of operations that is under development at INL.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

DEVELOPMENT AND DEMONSTRATION TESTBED FOR THE REMOTE OPERATIONS AND MONITORING OF MICROREACTORS

The nuclear industry is rapidly developing many advanced-reactor concepts for near-term deployment in both traditional and non-traditional nuclear-powered applications. One such category of advanced reactor is the microreactor, a class of reactor with less than 20MWth power output, intended for applications where the economics or logistics of traditional power sources are difficult. This includes applications such as remote communities, mining sites, defense installations, or humanitarian and disaster-relief missions. One key enabling feature for the successful deployment of microreactors is a remote operations capability. Remote operations provide monitoring and control capabilities which can significantly reduce staffing costs by eliminating the need for licensed operators at each reactor facility and improve the economic viability for microreactor deployment. A remote concept of operations is not currently an established capability in the nuclear industry. In addition, no demonstration microreactor is expected to complete construction or go critical until at least 2026. This leaves two major capability gaps: the successful demonstration of a remote concept of operations for microreactors and a test bed suitable for said demonstration. Both gaps must be addressed in order to advance the remote concepts of nuclear operation and, more broadly, microreactors themselves from paper to reality. This paper aims to fill these gaps and describes a test bed that would support development and deployment of a remote concept of nuclear operations, initial experimental results from that test bed, and the application of the test bed and experimental results for a digital-twin-based remote concept of operations underdevelopment at Idaho National Laboratory (INL). The platform chosen as a remote concept of nuclear operations test bed is the Single Primary Heat Extraction and Removal Emulator, known as SPHERE, located at INL. SPHERE is a small-scale non-nuclear test bed that emulates thermal behavior of a microreactor. The small-scale and non-nuclear nature of SHPERE limit safety concerns associated with remote operations while still providing the physical response representative of a microreactor. A network connection was added to SPHERE that enables remote-monitoring capability. This allows for real-time data streaming to networked workstations, data historians, and human-machine interfaces (HMIs). These are all critical components in a remote concept of operations, thus providing a robust development and demonstration platform. An initial experiment was performed using the SPHERE remote operations testbed. This included running a comprehensively instrumented SPHERE through a series of steady-state and transient operating scenarios in both normal and abnormal operating conditions, all while streaming live test data to a remote HMI and data warehouse. This initial experiment served three purposes: (1) characterizing the response of SPHERE, (2) demonstrating the remote connection to SPHERE, and (3) providing a baseline data set for development of a digital-twin-based remote concept of operations that is under development at INL.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Near-Real-Time Material Tracking: Combining Vis–NIR Spectroscopy with Flow Sensing for Accurate Nd(III) Quantification

A fiber-optic visible–near-infrared (vis–NIR) absorption spectroscopy and flow sensor system has been developed for near-real-time tracking of Nd mass in the effluent stream from a column in a fume hood. The approach leverages two unique data streams and a partial least-squares regression (PLSR) model trained on vis–NIR absorption spectra of Nd(III) (0–1.5 M) in 1 M HNO 3 . In-line volumetric flow rate and vis–NIR spectra are measured in sequence after a chromatography column. The time stamps from each data stream are then synchronized, which allows integrated volumes to be combined with Nd(III) molarities predicted by a PLSR model to accurately calculate the Nd mass flowing through the column. This integrated measurement provides instantaneous mass flow and accumulates these data over time to obtain the total mass processed. The methodology developed in this study contributes critical technical infrastructure to improve monitoring capabilities to support chemical separations and the production of strategic materials and isotopes.

Irvine, Sawyer B. [Oak Ridge National Laboratory (↗

EJFAT: Towards Intelligent Compute Destination Load Balancing

To handle increased data flow, Jefferson Lab (JLab) is partnering with ESnet for development of an AI/ML directed compute work Load Balancer (LB) of UDP streamed data. The LB is FPGA based featuring dynamically configurable, low latency and high throughput destination address switching. The LB provides integration of edge and core computing to support JLab experimental programs, the Electron-Ion Collider, as well as data centers of the future. In the ESnet/JLab FPGA Accelerated Transport (EJFAT) initiative, the function of the LB Data Plane (DP) is to redirect data streams to selectable (but unknown to sender) destination hosts based on current worload and within that host to destination ports as a function of sub- stream id. This effects hierarchical scaling, first across compute machines for processing over a series of events and second, across ports so different data source sub-streams may be assigned to different processors for further parallelization. The LB Control Plane (CP) programs the DP using compute farm telemetry to direct and balance workloads across a compute cluster as the operating conditions require. While Proportional/Integrative/Derivative (PID) controllers are often seen in similar applications, here we investigate the feasibility of a Reinforcement Learning (RL) based schedule manager running in the CP to provide dynamic updates to the DP scheduling policy.

Lawrence, David↗

My vehicle is a data mine

In this talk we explore how analysis of vehicle data provides information of individual vehicle behaviors, information of other vehicles in the flow of traffic, and insights into the behavior of drivers. Over the last two decades traditional passenger vehicles have been transformed from integrated two-port electrical nodes to cyber-physical systems of communicating computational nodes whose individual state and control variables are shared on a standard controller area network (CAN) bus. As driver assistance systems have crept into vehicles as safety features, driver behaviors can be observed through analysis of the data streams on the CAN bus as these new nodes communicate with one another. The properties of these data streams, as well as architectures and approaches to gather the data, are important to consider when drawing conclusions on the relevance of the data in making decisions at varying levels of the information hierarchy. We will demonstrate several technical challenges associated with these data collection processes, as well as preliminary results that demonstrate application relevance of the data to behavior, traffic, and systems domains.

42 ENGINEERING↗

IPC-Fusion (Infrastructure Perception and Control (IPC): Multisensor Data Fusion Software) [SWR-25-153]

As part of the National Laboratory of the Rockies' (NLR’s) Infrastructure Perception and Control Laboratory, the IPC-Fusion toolkit provides a probabilistic, scalable, multi-sensor fusion framework that integrates (late-stage fusion) heterogeneous object detection data from traffic sensors to enable robust, real-time tracking of roadway occupants. The algorithmic design of the toolkit is motivated by the need for creating a digital twin of traffic at the edge in a scalable and affordable manner. The software operates by combining object-level measurements (such as position and velocity) from a suite of sensors (such as radar, lidar, camera) using Kalman filtering and probabilistic data association techniques to overcome individual sensor limitations and achieve superior tracking performance in complex traffic zones. The framework addresses key challenges including heterogeneous measurement uncertainties, asynchronous data streams, varying spatiotemporal data resolutions, robust data association, and adaptive object lifecycle management. Validated on real-world traffic intersection data including vehicles and pedestrians, IPC-Fusion demonstrates enhanced tracking reliability across scenarios involving occlusions, sensor failures, and varying traffic densities, supporting the broader IPC initiative's goal of transforming transportation infrastructure through advanced perception capabilities for intelligent transportation systems, traffic safety applications, and autonomous vehicle support.

Sandhu, Rimple [National Laboratory of the Rockies↗

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation↗

krowkee

krowkee is a toolkit for scalably and efficiently summarizing many data streams in distributed memory. krowkee is intended for applications where one needs to summarize huge loosely structured data, such as matrices or graphs, where individual components such rows/columns or vertex adjacency information are impractical to store and directly inspect. krowkee ingests these objects as data streams - unstructured, arbitrarily ordered lists of updates - and accumulates summaries thereof in the form of data sketches.

Dunton, AlecM.↗

Watermarks in stream processing systems: semantics and comparative analysis of Apache Flink and Google cloud dataflow

Streaming data processing is an exercise in taming disorder: from oftentimes huge torrents of information, we hope to extract powerful and timely analyses. But when dealing with streaming data, the unbounded and temporally disordered nature of real-world streams introduces a critical challenge: how does one reason about the completeness of a stream that never ends? In this paper, we present a comprehensive definition and analysis of watermarks, a key tool for reasoning about temporal completeness in infinite streams.First, we describe what watermarks are and why they are important, highlighting how they address a suite of stream processing needs that are poorly served by eventually-consistent approaches:• Computing a single correct answer, as in notifications.• Reasoning about a lack of data, as in dip detection.• Performing non-incremental processing over temporal subsets of an infinite stream, as in statistical anomaly detection with cubic spline models.• Safely and punctually garbage collecting obsolete inputs and intermediate state.• Surfacing a reliable signal of overall pipeline health.Second, we describe, evaluate, and compare the semantically equivalent, but starkly different, watermark implementations in two modern stream processing engines: Apache Flink and Google Cloud Dataflow.

Akidau, Tyler↗

Towards a self-driving trigger at the LHC: adaptive response in real time

Real-time data filtering and selection—or trigger—systems at high-throughput scientific facilities such as the experiments at the Large Hadron Collider must process extremely high-rate data streams under stringent bandwidth, latency, and storage constraints. Yet these systems are typically designed as static, hand-tuned menus of selection criteria grounded in prior knowledge and simulation. In this work, we further explore the concept of a self-driving trigger, an autonomous data-filtering framework that reallocates resources and adjusts thresholds dynamically in real-time to optimize signal efficiency, rate stability, and computational cost as instrumentation and environmental conditions evolve. We introduce a benchmark ecosystem to emulate realistic collider scenarios and demonstrate real-time optimization of a menu including canonical energy sum triggers as well as modern anomaly-detection algorithms that target non-standard event topologies using machine learning. Using simulated data streams and publicly available collision data from the Compact Muon Solenoid experiment, we demonstrate the capability to dynamically and automatically optimize trigger performance under specific cost objectives without manual retuning. Our adaptive strategy shifts trigger design from static menus with heuristic tuning to intelligent, automated, data-driven control, unlocking greater flexibility and discovery potential in future high-energy physics analyses.

Emami, Shaghayegh [Michigan U.] (ORCID:00090007589↗

CANShield: Signal-based Intrusion Detection for Controller Area Networks

Modern vehicles rely on complex cyber-physical systems made up of hundreds of electronic control units (ECUs) connected through controller area network (CAN) buses. However, the CAN bus attack surface is increasing due to advanced features in automobiles, making it prone to injection attacks. The ordinary injection attacks disrupt the typical timing properties of the CAN data stream, and the rule-based intrusion detection systems (IDS) can easily detect them. However, advanced attackers can inject false data to the signal level, maintaining the regular pattern/frequency of the CAN messages. Such attacks can bypass the rule-based IDS or any anomaly-based IDS built on binary payload data. To make the vehicles robust against such intelligent attacks, we propose CANShield, a signal-based intrusion detection framework for the CAN bus that consists of three modules. A data preprocessing module handles the high-dimensional CAN data stream at the signal level and make them suitable for any machine learning model. A data analyzer module consists of multiple deep autoencoder networks, each analyzing the time series data from a different perspective. Finally, an attack detection module uses an ensemble method to make the final decision. Evaluation results on a standard signal-based dataset show the effectiveness of the CANShield in detecting five advanced attacks.

Shahriar, Md Hasan↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗