Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Real-time simulator and controller of power system using distributed data streaming server

Systems and methods for simulating and controlling a power system in real time, using a controller, are provided. The controller includes a simulation layer to simulate an operation of the power system, a disturbance generation layer to provide data to the simulation layer to disturb the simulated operation of the power system, an application layer to display the simulated operation of the power system and generate a control signal to control, based on the simulated operation of the power system, at least one element of the power system, and a distributed data streaming server (D2S2) to allow interoperability among the simulation layer, the disturbance layer, and the application layer.

Li, Fangxing↗

A High-level Design for Bidirectional Data Streaming to High-Performance Computing Systems from External Science Facilities

Cutting-edge science is increasingly data-driven due to the emergence of scientific machine learning models that can guide scientists toward fruitful areas of exploration. Experimental science facilities such as light and neutron sources, particle colliders, and radio astronomy telescopes are also producing raw measurement data at rates that exceed available data storage and computing capacity at those facilities. As a result, scientific workflows are being developed that concurrently couple experiments at science facilities with high-performance computing (HPC) facilities to enable analysis of experimental data while the experiment is ongoing, and where analysis results are potentially fed back to the experiment for use in experimental control and/or steering in a time-sensitive manner. Our goal is to design, prototype, and deploy a new capability for the Oak Ridge Leadership Computing Facility (OLCF) that enables such workflows through support for bidirectional, memory-based streaming of data from external experiments into and out of OLCF HPC systems. This high-level design document describes the related work and motivating use cases that inform our understanding of the technical requirements for this capability, and describes a proposed architectural solution that meets these requirements and our plans for demonstrating the capability.

97 MATHEMATICS AND COMPUTING↗

A tale of two towers: comparing NEON and AmeriFlux data streams at Bartlett Experimental Forest

Long-term ecological data are essential for detecting impacts of climate change and other global change factors, and for making informed predictions about future change. However, long-term measurements are rarely replicated at the site level, which raises questions about their representativeness. We used a multiscale approach to evaluate the agreement of parallel observations from AmeriFlux and NEON (National Ecological Observatory Network) towers at Bartlett Experimental Forest, New Hampshire, USA. The two towers are separated by a horizontal distance of 93 m. Here, we focused our analysis on standard meteorological variables; fluxes of CO 2 , sensible heat, and latent heat measured by eddy covariance; and phenology derived from PhenoCam imagery. Results suggest excellent agreement between AmeriFlux and NEON in meteorology and phenology, and good agreement in fluxes at the half-hourly scale. However, large disagreements in CO 2 and latent heat fluxes occurred at the annual scale, with implications especially for the forest carbon balance. The AmeriFlux tower measurements indicate a site that is close to carbon-neutral (-8 ± 65 g C m -2 y -1 , mean ± 1 SD), whereas the NEON tower measurements indicate a forest that is a carbon sink (-137 ± 10 g C m -2 y -1 ). Causes of this disagreement may include measurement height (26 m vs. 35 m), which resulted in different flux footprints being measured by the two towers, and differences in the flux measurement systems. Our results suggest the need for caution when attempting to merge long-term flux data from two different measurement platforms, and when using measurements from any one measurement platform to inform decision-making on issues related to carbon accounting or natural climate solutions.

Carbon cycle↗

Architecture Of A Multi-channel Data Streaming Device With An Fpga As A Coprocessor

The design of a data acquisition system often involves the integration of a Field Programmable Gate Array (FPGA) with analog front-end components to achieve precise timing and control. Reuse of these hardware systems can be difficult since they need to be tightly coupled to the communications interface and timing requirements of the specific ADC used. A hybrid design exploring the use of FPGA as a coprocessor to a traditional CPU in a dataflow architecture is presented. Reduction in the volume of data and gradual transitioning of data processing away from a hard real-time environment are both discussed. Chief design concerns, including data throughput and precise synchronization with external stimuli, are addressed. The discussion is illustrated by the implementation of a multi-channel digital integrator, a device based entirely on commercial off-the-shelf (COTS) equipment.

Nogiec, Jerzy M.↗

Transitioning from File-Based HPC Workflows to Streaming Data Pipelines with openPMD and ADIOS2

This paper aims to create a transition path from file-based IO to streaming-based workflows for scientific applications in an HPC environment. By using the openPMP-api, traditional workflows limited by filesystem bottlenecks can be overcome and flexibly extended for in situ analysis. The openPMD-api is a library for the description of scientific data according to the Open Standard for Particle-Mesh Data (openPMD). Its approach towards recent challenges posed by hardware heterogeneity lies in the decoupling of data description in domain sciences, such as plasma physics simulations, from concrete implementations in hardware and IO. The streaming backend is provided by the ADIOS2 framework, developed at Oak Ridge National Laboratory. This paper surveys two openPMD-based loosely-coupled setups to demonstrate flexible applicability and to evaluate performance. In loose coupling, as opposed to tight coupling, two (or more) applications are executed separately, e.g. in individual MPI contexts, yet cooperate by exchanging data. This way, a streaming-based workflow allows for standalone codes instead of tightly-coupled plugins, using a unified streaming-aware API and leveraging high-speed communication infrastructure available in modern compute clusters for massive data exchange. We determine new challenges in resource allocation and in the need of strategies for a flexible data distribution, demonstrating their influence on efficiency and scaling on the Summit compute system. The presented setups show the potential for a more flexible use of compute resources brought by streaming IO as well as the ability to increase throughput by avoiding filesystem bottlenecks.

Poeschel, Franz↗

EdgeAI: Machine learning via direct attached accelerator for streaming data processing at high shot rate x-ray free-electron lasers

We present a case for low batch-size inference with the potential for adaptive training of a lean encoder model. We do so in the context of a paradigmatic example of machine learning as applied in data acquisition at high data velocity scientific user facilities such as the Linac Coherent Light Source-II x-ray Free-Electron Laser. We discuss how a low-latency inference model operating at the data acquisition edge can capitalize on the naturally stochastic nature of such sources. We simulate the method of attosecond angular streaking to produce representative results whereby simulated input data reproduce high-resolution ground truth probability distributions. By minimizing the mean-squared error between the decoded output of the latent representation and the ground truth distributions, we ensure that the encoding layers and resulting latent representation maintains full fidelity for any downstream task, be it classification or regression. We present throughput results for data-parallel inference of various batch sizes, some with throughput exceeding 100 k images per second. We also show in situ training below 10 s per epoch for the full encoder–decoder model as would be relevant for streaming and adaptive real-time data production at our nation’s scientific light sources.

97 MATHEMATICS AND COMPUTING↗

Search for long-lived particles decaying into muon pairs in proton-proton collisions at $\sqrt{s}$ = 13 TeV collected with a dedicated high-rate data stream

A search for long-lived particles decaying into muon pairs is performed using proton-proton collisions at a center-of-mass energy of 13 TeV, collected by the CMS experiment at the LHC in 2017 and 2018, corresponding to an integrated luminosity of 101 fb -1 . The data sets used in this search were collected with a dedicated dimuon trigger stream with low transverse momentum thresholds, recorded at high rate by retaining a reduced amount of information, in order to explore otherwise inaccessible phase space at low dimuon mass and nonzero displacement from the primary interaction vertex. No significant excess of events beyond the standard model expectation is found. Upper limits on branching fractions at 95% confidence level are set on a wide range of mass and lifetime hypotheses in beyond the standard model frameworks with the Higgs boson decaying into a pair of long-lived dark photons, or with a long-lived scalar resonance arising from a decay of a b hadron. The limits are the most stringent to date for substantial regions of the parameter space. These results can be also used to constrain models of displaced dimuons that are not explicitly considered in this paper.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Digital Twin for an Inverter-Based Resource Power Plant: Real-time data streaming unlocks situation awareness

Here, this study presents the development and successful implementation of a digital twin specifically designed for a grid-connected IBR power plant. By integrating a reduced-order model of the IBR system and dynamically updating the grid impedance with real-time data, the digital twin effectively captures and replicates the behavior of the physical system. Its accuracy and reliability are validated through critical test scenarios, including a three-phase fault and a line-tripping event. The results confirm that the digital twin closely emulates its physical counterpart, demonstrating its strong potential for real-time analysis, system monitoring, and predictive decision making in modern power systems.

Digital twins↗

Investigation of Multiple Data Streams for Gearbox Bearing Fault Prediction Through Machine-Learning Models

Operations and maintenance (O&M) cost of wind plant accounts up to 30% of total energy cost, which can be reduced through continuous monitoring and successfully detecting incipient wind turbine failures. To accomplish this, condition monitoring and predictive maintenance systems are being implemented in wind industry to support O&M decision making. A wide range of approaches for condition monitoring and fault prediction have been developed. These approaches generally use historical data of wind turbines collected by Supervisory Control and Data Acquisition (SCADA) system to identify patterns that lead to failure. These SCADA data show the overall condition of a wind turbine and can be leveraged to detect when the turbine's performance is degrading and to identify if a fault is developing. However, it becomes challenging to predict the failure of a specific wind turbine gearbox bearing, because the SCADA data are often not directly linked to the component. To bridge the gap, we have investigated features calculated from SCADA data using physics-based models and the gearbox design over the years. The damaged metric we used in the physics domain is frictional energy. Combining these physics domain variables with SCADA data as inputs to various machine learning models for gearbox bearing fault prediction, we have demonstrated the benefits of leveraging both physics and data domain models. It was an attempt to improve frictional-energy-based damage metric by adding data domain inputs, as we had learned that the frictional-energy-based damage metric alone is not sufficient to single out failed bearings from healthy. As condition monitoring data (either vibration or oil debris data) has become available at more and more wind plants, we would like to evaluate whether by adding the condition monitoring data can help further improve the performance of frictional-energy-based damage metric for gearbox bearing fault prediction. Both cases by modeling through various machine learning algorithms are discussed in this study along with some observations.

fault prediction↗

Online real-time learning of dynamical systems from noisy streaming data

Abstract Recent advancements in sensing and communication facilitate obtaining high-frequency real-time data from various physical systems like power networks, climate systems, biological networks, etc. However, since the data are recorded by physical sensors, it is natural that the obtained data is corrupted by measurement noise. In this paper, we present a novel algorithm for online real-time learning of dynamical systems from noisy time-series data, which employs the Robust Koopman operator framework to mitigate the effect of measurement noise. The proposed algorithm has three main advantages: (a) it allows for online real-time monitoring of a dynamical system; (b) it obtains a linear representation of the underlying dynamical system, thus enabling the user to use linear systems theory for analysis and control of the system; (c) it is computationally fast and less intensive than the popular extended dynamic mode decomposition (EDMD) algorithm. We illustrate the efficiency of the proposed algorithm by applying it to identify the Van der Pol oscillator, the chaotic attractor of the Henon map, the IEEE 68 bus system, and a ring network of Van der Pol oscillators.

97 MATHEMATICS AND COMPUTING↗

Calculation of Nuclear Reactor Cooling Tower Performance With Limited Data Streams

Monitoring of cooling tower performance in a nuclear reactor facility is necessary to ensure safe operation; however, instrumentation for measuring performance characteristics can be difficult to install and may malfunction or break down over long duration experiments. This paper describes employing a thermodynamic approach to quantify cooling tower performance, the Merkel model, which requires only five parameters, namely, inlet water temperature, outlet water temperature, liquid mass flowrate, gas mass flowrate, and wet bulb temperature. Using this model, a general method to determine cooling tower operation for a nuclear reactor was developed in situations when neither the outlet water temperature nor gas mass flowrate are available, the former being a critical piece of information to bound the Merkel integral. Furthermore, when multiple cooling tower cells are used in parallel (as would be in the case of large-scale cooling operations), only the average outlet temperature of the cooling system is used as feedback for fan speed control, increasing the difficulty of obtaining the outlet water temperature for each cell. To address these shortcomings, this paper describes a method to obtain individual cell outlet water temperatures for mechanical forced-air cooling towers via parametric analysis and optimization. In this method, the outlet water temperature for an individual cooling tower cell is acquired as a function of the liquid-to-gas ratio (L/G). Leveraging the tight tolerance on the average outlet water temperature, an error function is generated to describe the deviation of the parameterized L/G to the highly controlled average outlet temperature. The method was able to determine the gas flowrate at rated conditions to be within 3.9% from that obtained from the manufacturer’s specification, while the average error for the four individual cooling cell outlet water temperatures were 1.6 °C, -0.5 °C, -1.0 °C, and 0.3 °C.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Streaming Data Reorganization at Scale with DeltaFS Indexed Massive Directories

We report complex storage stacks providing data compression, indexing, and analytics help leverage the massive amounts of data generated today to derive insights. It is challenging to perform this computation, however, while fully utilizing the underlying storage media. This is because, while storage servers with large core counts are widely available, single-core performance and memory bandwidth per core grow slower than the core count per die. Computational storage offers a promising solution to this problem by utilizing dedicated compute resources along the storage processing path. We present DeltaFS Indexed Massive Directories (IMDs), a new approach to computational storage. DeltaFS IMDs harvest available (i.e., not dedicated) compute, memory, and network resources on the compute nodes of an application to perform computation on data. We demonstrate the efficiency of DeltaFS IMDs by using them to dynamically reorganize the output of a real-world simulation application across 131,072 CPU cores. DeltaFS IMDs speed up reads by 1,740x while only slightly slowing down the writing of data during simulation I/O for in situ data processing.

97 MATHEMATICS AND COMPUTING↗