Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Kernelized approaches to streaming compression of scientific data

In this paper three algorithms are developed for the streaming compression of scientific data. The algorithms presented are reliant on the theory of vector-valued reproducing kernel Hilbert spaces and operator valued kernel. Further, the scientific data is modeled as a snapshot of time dependent vector field F(x, t) over a manifold M and the recovery of the data is framed as a learning problem. These processes are then appropriately modified and ana lyzed for the streaming scenario in which data is generated without the ability to revisit past entries.

97 MATHEMATICS AND COMPUTING↗

via machinae : Searching for stellar streams using unsupervised machine learning

ABSTRACT We develop a new machine learning algorithm, via machinae, to identify cold stellar streams in data from the Gaia telescope. via machinae is based on ANODE, a general method that uses conditional density estimation and sideband interpolation to detect local overdensities in the data in a model agnostic way. By applying ANODE to the positions, proper motions, and photometry of stars observed by Gaia, via machinae obtains a collection of those stars deemed most likely to belong to a stellar stream. We further apply an automated line-finding method based on the Hough transform to search for line-like features in patches of the sky. In this paper, we describe the via machinae algorithm in detail and demonstrate our approach on the prominent stream GD-1. Though some parts of the algorithm are tuned to increase sensitivity to cold streams, the via machinae technique itself does not rely on astrophysical assumptions, such as the potential of the Milky Way or stellar isochrones. This flexibility suggests that it may have further applications in identifying other anomalous structures within the Gaia data set, for example debris flow and globular clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

An Incremental Tensor Train Decomposition Algorithm

We present a new algorithm for incrementally updating the tensor train decomposition of a stream of tensor data. This new algorithm, called the tensor train incremental core expansion (TT-ICE) improves upon the current state-of-the-art algorithms for compressing in tensor train format by developing a new adaptive approach that incurs significantly slower rank growth and guarantees compression accuracy. This capability is achieved by limiting the number of new vectors appended to the TT-cores of an existing accumulation tensor after each data increment. These vectors represent directions orthogonal to the span of existing cores and are limited to those needed to represent a newly arrived tensor to a target accuracy. We provide two versions of the algorithm: TT-ICE and TT-ICE accelerated with heuristics (TT-ICE*). Here, we provide a proof of correctness for TT-ICE and empirically demonstrate the performance of the algorithms in compressing large-scale video and scientific simulation datasets. Compared to existing approaches that also use rank adaptation, TT-ICE* achieves 57× higher compression and up to 95% reduction in computational time.

97 MATHEMATICS AND COMPUTING↗

Large-scale real-time signal processing in physics experiments: the ALICE TPC FPGA pipeline

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB s -1 of raw detector data. This requirement is met by a custom FPGA-based processing pipeline that performs the complete front-end data treatment fully in-stream, including common-mode correction, pedestal subtraction, ion-tail filtering, zero suppression, and dense data packing. A central element of the design is a highly parallel common-mode correction algorithm operating directly on the streaming data. It robustly identifies signal-free readout channels on a time-bin basis and applies pad-dependent scaling to compensate for local variations in capacitive coupling in the GEM readout. In combination with pedestal subtraction and ion-tail filtering, this enables accurate baseline restoration under extreme high-occupancy conditions, preventing signal loss while efficiently suppressing noise prior to zero suppression. The pipeline operates continuously at the full detector bandwidth and reduces the raw input rate of approximately 3 TB s -1 to about 900 GBps for Pb-Pb collisions at the target interaction rate. Overall, it represents a large-scale FPGA-based real-time signal-processing implementation for high-energy physics detector readout.

Digital signal processing (DSP)↗

Automatic Traffic Queue-End Identification using Location-Based Waze User Reports

Traffic queues, especially queues caused by non-recurrent events such as incidents, are unexpected to high-speed drivers approaching the end of queue (EOQ) and become safety concerns. Though the topic has been extensively studied, the identification of EOQ has been limited by the spatial-temporal resolution of traditional data sources. This study explores the potential of location-based crowdsourced data, specifically Waze user reports. It presents a dynamic clustering algorithm that can group the location-based reports in real time and identify the spatial-temporal extent of congestion as well as the EOQ. The algorithm is a spatial-temporal extension of the density-based spatial clustering of applications with noise (DBSCAN) algorithm for real-time streaming data with an adaptive threshold selection procedure. Here, the proposed method was tested with 34 traffic congestion cases in the Knoxville, Tennessee area of the United States. It is demonstrated that the algorithm can effectively detect spatial-temporal extent of congestion based on Waze report clusters and identify EOQ in real-time. The Waze report-based detection are compared to the detection based on roadside sensor data. The results are promising: The EOQ identification time of Waze is similar to the EOQ detection time of traffic sensor data, with only 1.1 min difference on average. In addition, Waze generates 1.9 EOQ detection points every mile, compared to 1.8 detection points generated by traffic sensor data, suggesting the two data sources are comparable in respect of reporting frequency. The results indicate that Waze is a valuable complementary source for EOQ detection where no traffic sensors are installed.

99 GENERAL AND MISCELLANEOUS↗

Simulating the Autonomous Future: A Look at Virtual Vehicle Environments and How to Validate Simulation Using Public Data Sets

The rapid evolution of autonomous vehicles (AVs) has exposed the need for fast-paced development and testing processes of a variety of perception, planning, and control algorithms. To expedite development, the AV industry and researchers leverage virtual vehicle environments to simulate a range of test scenarios that may otherwise be costly or difficult to conduct on a real test track. However, the various virtual environments may have different results depending on the fidelity of various simulation features, such as vehicle dynamics, sensor simulation, and environment recreation. Herein, this tutorial article examines a proposed framework for constructing, parameterizing, and validating a virtual vehicle environment using an existing AV data set. First, an overview of several open source and commercially available simulation tools, including their associated workflows, for scene and scenario creation is presented. Next, various open AV data sets are examined to inform the data set selection for the validation framework. Then, an example workflow of recreating a real-world scene from the selected data set in a simulation tool with various emulated sensors parameterized to match the data set is demonstrated. Finally, an example AV-perception algorithm is subjected to data streams from virtual and real-world environments and suggested metrics for analyzing the results are discussed.

42 ENGINEERING↗

The Pristine survey – XVI. The metallicity of 26 stellar streams around the Milky Way detected with the STREAMFINDER in Gaia EDR3

ABSTRACT We use the photometric metallicities provided by the panoramic Pristine survey to study the veracity and derive the metallicities of the numerous stellar streams found by the application of the STREAMFINDER algorithm to the Gaia Early Data Release 3 data. All 26 streams present in Pristine show a clear metallicity distribution function, which provides an independent check of the reality of these structures, supporting the reliability of STREAMFINDER in finding streams and the power of Pristine to measure precise metallicities. We further present six candidate structures with coherent phase-space and metallicity signals that are very likely streams. The majority of studied streams are very metal-poor (14 structures with [Fe/H] < −2.0) and include three systems with [Fe/H] < −2.9 (C-11, C-19, and C-20). These streams could be the closest debris of low-luminosity dwarf galaxies or may have originated from globular clusters of significantly lower metallicity than any known current Milky Way globular cluster. Our study shows that the promise of the Gaia data for Galactic Archeology studies can be substantially strengthened by quality photometric metallicities, allowing us to peer back into the earliest epochs of the formation of our Galaxy and its stellar halo constituents.

79 ASTRONOMY AND ASTROPHYSICS↗

Streaming Readout and Data-Stream Processing With ERSAP

With the exponential growth in the volume and complexity of data generated at high-energy physics and nuclear physics research facilities, there is an imperative demand for innovative strategies to process this data in real or near-real-time. Given the surge in the requirement for high-performance computing, it becomes pivotal to reassess the adaptability of current data processing architectures in integrating new technologies and managing streaming data. This paper introduces the ERSAP framework, a modern solution that synergizes flow-based programming with the reactive actor model, paving the way for distributed, reactive, and high performance in data stream processing applications. Additionally, we unveil a novel algorithm focused on time-based clustering and event identification in data streams. The efficacy of this approach is further exemplified through the data-stream processing outcomes obtained from the recent beam tests of the EIC prototype calorimeter at DESY.

Vardan, Gyurjyan↗

Autonomous Anomaly Detection For Continuous Streams

The code implements the Isolation Forest (IFML) algorithm within the digital twin (DT) of the AGN-201 nuclear reactor. The DT captures real-time operational data including control rod positions, reactor power, and temperature. The IFML model isolates anomalies by detecting patterns that deviate from expected operational behavior. The algorithm recursively partitions the data and assigns anomaly scores based on the isolation of rare and different events. By tuning parameters specific to the reactor’s operational data, the IFML identifies deviations such as unauthorized material insertions or reactor reactivity shifts. The system streams data using LabView and integrates with the DeepLynx data warehouse for anomaly processing.

Trevino, Eduardo↗

F-Hash: Feature-Based Hash Design for Time-Varying Volume Visualization via Multi-Resolution Tesseract Encoding

Interactive time-varying volume visualization is challenging due to its complex spatiotemporal features and sheer size of the dataset. Recent works transform the original discrete time-varying volumetric data into continuous Implicit Neural Representations (INR) to address the issues of compression, rendering, and super-resolution in both spatial and temporal domains. However, training the INR takes a long time to converge, especially when handling large-scale time-varying volumetric datasets. In this work, we proposed F-Hash, a novel feature-based multi-resolution Tesseract encoding architecture to greatly enhance the convergence speed compared with existing input encoding methods for modeling time-varying volumetric data. The proposed design incorporates multi-level collision-free hash functions that map dynamic 4D multi-resolution embedding grids without bucket waste, achieving high encoding capacity with compact encoding parameters. Our encoding method is agnostic to time-varying feature detection methods, making it a unified encoding solution for feature tracking and evolution visualization. Experiments show the F-Hash achieves state-of-the-art convergence speed in training various time-varying volumetric datasets for diverse features. We also proposed an adaptive ray marching algorithm to optimize the sample streaming for faster rendering of the time-varying neural representation.

deep learning↗

Developing a Decision Support Engine to Enable Irrigation Modernization - Poster

Irrigation systems in the United States are some of the oldest continually utilized infrastructure in existence today, with some systems exceeding 100 years in age. They are operated to meet farming demands but are managed through a balance of varying influence: policies limiting water usage, stakeholder interests, and environmental impacts. Irrigation modernization is defined as a set of activities that update and improve existing irrigation systems, including, but not limited to improving water quantity, development of distributed energy resources for surrounding communities, ecosystem services, and improved agricultural yields. Modernizing an existing irrigation system can enable stakeholders to combat changing environmental and population demands but is difficult because the complexity involved in determining the potential benefits and consequences of irrigation modernization is high. We are combining a large amount of various geospatial, tabular, and temporal data with subject matter expertise into a decision support engine that will enable stakeholders to determine the benefits and consequences of irrigation modernization in their irrigation systems. A web-based GIS will allow the user to construct the modifications out of a palette of modernization options, which will be sent to the analytics engine for computations, and back to the web client for a graphical display and comparison of relevant metrics. Our development process involves four phases: 1) identify mechanisms of modernization, 2) identify data requirements, data streams, first principles and applicable algorithms necessary to quantify modernization mechanisms, 3) create ‘modules’ for each modernization mechanism, these modules will form the decision support engine, each capable of performing independently but can also inform other modules when needed, 4) Merge the decision support engine with a user interface, capable of ingesting user inputs and returning insights into the impacts of a modernization project as they relate to economic, environmental, monetary, and energy generation. Once complete, it is our intention that this tool will be fundamental in irrigation modernization projects, providing a strong analytical basis from which stakeholders can quickly make informed decisions regarding project development.

13 HYDRO ENERGY↗

Timely Reporting of Heavy Hitters Using External Memory

Given an input stream S of size N, a Φ-heavy hitter is an item that occurs at least ΦN times in S. The problem of finding heavy-hitters is extensively studied in the database literature. In this work, we study a real-time heavy-hitters variant in which an element must be reported shortly after we see its T = Φ N-th occurrence (and hence it becomes a heavy hitter). We call this the Timely Event Detection (TED) Problem. The TED problem models the needs of many real-world monitoring systems, which demand accurate (i.e., no false negatives) and timely reporting of all events from large, high-speed streams with a low reporting threshold (high sensitivity). Like the classic heavy-hitters problem, solving the TED problem without false-positives requires large space (Ω (N) words). Thus in-RAM heavy-hitters algorithms typically sacrifice accuracy (i.e., allow false positives), sensitivity, or timeliness (i.e., use multiple passes). We show how to adapt heavy-hitters algorithms to external memory to solve the TED problem on large high-speed streams while guaranteeing accuracy, sensitivity, and timeliness. Our data structures are limited only by I/O-bandwidth (not latency) and support a tunable tradeoff between reporting delay and I/O overhead. With a small bounded reporting delay, our algorithms incur only a logarithmic I/O overhead. We implement and validate our data structures empirically using the Firehose streaming benchmark. Multi-threaded versions of our structures can scale to process 11M observations per second before becoming CPU bound. In comparison, a naive adaptation of the standard heavy-hitters algorithm to external memory would be limited by the storage device’s random I/O throughput, i.e., ≈100K observations per second.

97 MATHEMATICS AND COMPUTING↗

Qubit Lattice Algorithms based on the Schrodinger-Dirac representation of Maxwell Equations and their Extensions

It is well known that Maxwell equations can be expressed in a unitary Schrodinger-Dirac representation for homogeneous media. However, difficulties arise when considering inhomoge- neous media. A Dyson map points to a unitary field qubit basis, but the standard qubit lattice algorithm of interleaved unitary collision-stream operators must be augmented by some sparse non-unitary potential operators that recover the derivatives on the refractive indices. The effect of the steepness of these derivatives on two dimensional scattering is examined with simulations showing quite complex wavefronts emitted due to transmissions/reflections within the dielectric objects. Maxwell equations are extended to handle dissipation using Kraus operators. Then, our theoretical algorithms are extended to these open quantum systems. A quantum circuit diagram is presented as well as estimates on the required number of quantum gates for implementation on a quantum computer.

Vahala, George↗

Qubit Lattice Algorithms Based on the Schrödinger-Dirac Representation of Maxwell Equations and Their Extensions

It is well known that Maxwell equations can be expressed in a unitary Schrodinger-Dirac representation for homogeneous media. However, difficulties arise when considering inhomogeneous media. A Dyson map points to a unitary field qubit basis, but the standard qubit lattice algorithm of interleaved unitary collision-stream operators must be augmented by some sparse non-unitary potential operators that recover the derivatives on the refractive indices. Here, the effect of the steepness of these derivatives on two-dimensional scattering is examined with simulations showing quite complex wavefronts emitted due to transmissions/reflections within the dielectric objects. Maxwell equations are extended to handle dissipation using Kraus operators. Then, our theoretical algorithms are extended to these open quantum systems. A quantum circuit diagram is presented as well as estimates on the required number of quantum gates for implementation on a quantum computer.

2D electromagnetic scattering↗

Online Power System Event Detection via Bidirectional Generative Adversarial Networks

Accurate and speedy detection of power system events is critical to enhancing the reliability and resiliency of power systems. Although supervised deep learning algorithms show great promise in identifying power system events, they require a large volume of high-quality event labels for training. This paper develops a bidirectional anomaly generative adversarial network (GAN)-based algorithm to detect power system events using streaming PMU data, which does not rely on a huge amount of event labels. By introducing conditional entropy constraint in the objective function of GAN and graph signal processing-based PMU sorting technique, our proposed algorithm significantly outperforms state-of-the-art event detection algorithms in terms of accuracy. To facilitate the adoption of the proposed algorithm, a prototype online platform is also developed using Apache Hadoop, Kafka, and Spark to enable real-time event detection. Here, the accuracy and computational efficiency of the proposed algorithm are validated using a large-scale real-world PMU dataset from the Eastern Interconnection of the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Constituent Data Replacement Tool

The purpose of this tool is to estimate key parameters that may be missing in public wastewater composition datasets. The tool can be applied to develop complete treatment and critical mineral extraction profiles for leachate, produced water and other aqueous waste streams. The tool applies machine learning algorithms to replace missing data in a user’s water data set that are adjusted based on user preferences for options including algorithm type, number of features, and classification variables. The tool can use the user’s data alone or combine user data with the NEWTS USGS Produced Water Database for more robust training. This research was funded by the U.S. Department of Energy’s Office Fossil Energy and Carbon Management (FECM) through National Energy Technology Laboratory’s ongoing research under the Water Management for Power System Field Work Proposal, DE-FECM 1022428 and Critical Minerals Field Work Proposal, DE-FECM 1022420.

Aqueous Chemistry↗

Energy-Efficient Adaptive Cruise Control Using Eco-Driving Algorithm

SwRI Eco-Driving Algorithm – Objectives and Method – Required Information Streams and Velocity Profile Realization ▪ Impact on Energy Consumption – Study 1 – NEXTCAR II • Individual (ego/eco) vehicle on Blanco Road, San Antonio, TX – Study 2 – DOE EEMS • Entire Urban Corridor – N High Street, Columbus, OH

Bhagdikar, Piyush↗