Big data in nuclear comes in the form of many structures, systems, components, data streams, data types, and more
Big Data in nuclear comes in the form of many structures, systems, components, data streams, data types, and more.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Big Data in nuclear comes in the form of many structures, systems, components, data streams, data types, and more.
Various Internet of Things (IoT) devices generate complex, dynamically changed, and infinite data streams. Adversaries can cause harm if they can access the user’s sensitive raw streaming data. For this reason, protecting the privacy of the data streams is crucial. In this paper, we explore local differential privacy techniques for streaming data. We compare the techniques and report the advantages and limitations. We also present the effect on component (e.g., smoother, perturber) variations of distribution-based local differential privacy. We find that combining distribution-based noise during perturbation provides more flexibility to the interested entity.
Streaming data is data that is emitted at variable volumes in a continuous, incremental manner with the goal of low-latency processing often at a different physical location. Network infrastructure is used to facilitate the connection between data sources and sinks, and must be robust to handle the requirements of the workflow. The U.S. Department of Energy Office of Science (DOE SC) a federal agency supporting fundamental scientific research for energy and the Nation’s largest supporter of basic research in the physical sciences. DOE SC has the responsibility for operating $\mathbf{1 0}$ National Laboratories, and 28 scientific user facilities supporting advanced supercomputers, particle accelerators, large x-ray light sources, neutron scattering sources, and other specialized facilities for nanoscience and genomics. This paper investigates the state of streaming data workfows, and details some of the approaches to this challenging problem.
In this paper we demonstrate direct data streaming from instruments and detectors at a large-scale experimental facility to a supercomputer for real-time data processing and feedback. Streaming data to supercomputers introduces the potential for novel scientific applications and workflow models, including the ability to provide real-time feedback from very large datasets during an experiment and the integration of real-time ML training and inference at scale. We discuss a successful demonstration for real-time processing of data from the Advanced Photon Source (APS) on the Polaris supercomputer using an EPICS-based streaming framework. We describe the capabilities of the streaming framework itself, and outline the architecture that allows us to process experimentally derived data on a supercomputer without file-based data transfers. We present throughput measurements that are indicative of system performance capable of sustaining the expected data production rates of the facility, as well as discuss some outstanding challenges and our future directions.
With the exponential growth in the volume and complexity of data generated at high-energy physics and nuclear physics research facilities, there is an imperative demand for innovative strategies to process this data in real or near-real-time. Given the surge in the requirement for high-performance computing, it becomes pivotal to reassess the adaptability of current data processing architectures in integrating new technologies and managing streaming data. This paper introduces the ERSAP framework, a modern solution that synergizes flow-based programming with the reactive actor model, paving the way for distributed, reactive, and high performance in data stream processing applications. Additionally, we unveil a novel algorithm focused on time-based clustering and event identification in data streams. The efficacy of this approach is further exemplified through the data-stream processing outcomes obtained from the recent beam tests of the EIC prototype calorimeter at DESY.
Apparatus for doubling the data density rate of an analog to digital converter or doubling the data density storage capacity of a memory deviced is discussed. An interstitial data point midway between adjacent data points in a data stream having an even number of equal interval data points is generated by applying a set of predetermined one-dimensional convolute integer coefficients which can include a set of multiplier coefficients and a normalizer coefficient. Interpolator means apply the coefficients to the data points by weighting equally on each side of the center of the even number of equal interval data points to obtain an interstital point value at the center of the data points. A one-dimensional output data set, which is twice as dense as a one-dimensional equal interval input data set, can be generated where the output data set includes interstitial points interdigitated between adjacent data points in the input data set. The method for generating the set of interstital points is a weighted, nearest-neighbor, non-recursive, moving, smoothing averaging technique, equivalent to applying a polynomial regression calculation to the data set.
Implementing a physics data processing application is relatively straightforward with the use of current containerization technologies and container image runtime services, which are prevalent in most high-performance computing (HPC) environments. However, the process is complicated by the challenges associated with data provisioning and migration, impacting the ease of workflow migration and deployment. Transitioning from traditional file-based batch processing to data-stream processing workflows is suggested as a method to streamline these workflows. This transition not only simplifies file provisioning and migration but also significantly reduces the necessity for extensive disk space. Data-stream processing is particularly effective for real-time processing during data acquisition, thereby enhancing data quality assurance. This paper introduces the integration of the JLAB CLAS12 event reconstruction application within the ERSAP data-stream processing framework that facilitates the execution of streaming event reconstruction at a remote data center and enables the return streaming of reconstructed events to JLAB while circumventing the need for temporary data storage throughout the process.
A stream of raw data is compressed prior to transmissio in a communication channel by a system which includes modules for choosing a current segment of the raw data stream for processing and defining a set of operators for representing data segments by a mathematical operation and parameters thereof. The system performs a competitive evaluation of different tools comprising different combinations of one or more of the operators and the parameters threrof with respect to the current data segment in order to determine relative abilities among the different tools to reduce the number of bits required to represent the current data segment. The system then selects a tool and a set of parameters thereof found in the competitive evaluation to have a superior ability relative to others of the different tools to reduce a number of bits required to represent the current data segment.
We consider two streams of data or measurements with disparate qualities and time resolutions that need to be classified. The first stream consists of higher quality data at a coarser time resolution, and the other consists of lower quality data at a finer time resolution. We present a fuser-switch method that fuses the set of classifiers of each stream separately and switches between them. We show that this method provides classification decisions at a finer time resolution with superior detection and false alarm probabilities compared to individual classifiers, under the statistical independence and time resolution ratio conditions. When classifiers are trained using machine learning methods, we show that this superior performance is guaranteed with a confidence probability specified by the classifiers' generalization equations. We use these results to provide analytical foundations for previous practical results that achieved significant performance improvements in classifying Pu/Np target dissolution events at a radiochemical processing facility.
In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.
This paper presents an reactive, actor-model and FBP paradigm based framework that we develop to design data-stream processing applications for HEP and NP. This framework encourages a functional decomposition of the overall data processing application into small mono-functional artifacts. Artifacts that are easy to understand, develop, deploy and debug. The fact that these artifacts (actors) are programmatically independent they can be scaled and optimized independently, which is impossible to do for components of the monolithic application. One of the important advantages of this approach is fault tolerance where independent actors can come and go on the data-stream without forcing the entire application to crash. Furthermore, it also makes it is easy to locate the faulty actor in the data pipeline. Due the fact that the actors are loosely coupled, and that the data (inevitably) carries the context, they can run on heterogeneous environments, utilizing different accelerators. This paper describes the main design concepts of the framework and presents a ?proof of concept? application design and deployment results obtain processing on-beam calorimeter streaming data.
With ongoing automation and digitization of the electric power system, several Phasor Measurement Units(PMUs) have been deployed for monitoring and control. PMU data can have multiple anomalies, and many of the researchers in the past have concentrated on training machine/deep learning algorithms offline for anomaly detection over PMU data (i.e., not in real time). These machine/deep learning algorithms, when trained offline on a sample rather than a population of the dataset, fail to consider the dynamic behavior of the power grid in real-time, resulting in low accuracy. Considering the dynamic behavior of the power grid (e.g., change in load, generation, distributed energy resources (DERs) switching, network, controls), the definition of data anomalies varies in time and requires online training. A fundamental challenge is to enable online (i.e., real-time) training of machine/deep learning algorithms for anomaly detection over streaming PMU data. While machine/deep learning is often desirable to manage data streams, training a deep learning algorithm over streaming PMU data is nontrivial due to changes in data statistics caused by dynamic streaming data. This paper proposes PMUNET: a novel device-level deep learning-based data-driven approach for anomaly detection, localization, and classification over streaming PMU data, using online learning and multivariate data-drift detection algorithm .Two variants of PMUNET, Dynamic data Change Driven Learning (DCDL) and Continuity Driven Learning (CDL), are proposed and compared. DCDL aims to train the deep learning algorithm whenever the definition of anomaly changes due to the power grid dynamics. On the other hand, CDL continuously trains the deep learning algorithm over the PMU data-stream. The experimental results verify that DCDL outperforms CDL and other efficient anomaly detection methods over multiple events such as faults and load/ generator/capacitor/DERs variations/switching for IEEE 14 and 39 Bus test system as well as real PMU industrial data. The result verifies that DCDL variant of PMUNET improves over existing approach with a gain of 2% - 10% in terms of accuracy, false-positive rate, and false-negative rate.
One of the richest potential sources of insight into fundamental physics that LISA will be capable of observing is the inspiral of supermassive black hole binaries (BHBs). However, the data analysis challenge presented by the LISA data stream is quite unlike the situation for present day gravitational wave detectors. In order to make the precision measurements necessary to achieve LISA's science goals, the BHB signal must be distinguished from a data stream that not only contains instrumental noise, but potentially thousands of other signals as well, so that the "background" we wish to separate out to focus on the BHB signal is likely to be highly nonstationary and nongaussian, as well as being of scientific interest in its own right. In addition, whereas the theoretical templates that we calculate in order to ultimately estimate the parameters can afford to be somewhat inaccurate and still be effective for present day and near future detectors, this is not the case for LISA, and extremely high fidelity of the theoretical templates for high signal-to-noise signals will be required to prevent theoretical errors from dominating the parameter estimates. NVe, will describe efforts in the community of LISA data analysts to address the challenges regarding the specific issue of BHB signals. These efforts include using a Markov Chain Monte Carlo approach with the freedom to model the BHB and the other signals present in the data stream simultaneously, rather than trying to remove other signals and risk biasing the remaining data. The Mock LISA Data Challenge is a community of LISA scientists who generate rounds of simulated LISA noise with increasingly difficult signal content, and invite the LISA data analysis community to exercise their methods, or develop new methods, in an attempt to extract the parameters for the signals embedded in the mock data. In addition to practical approaches such ,is this to assess the level of parameter accuracy, one can apply the Fisher matrix formalism to assess both the statistical errors from noise and the theoretical errors
One of the most pernicious cause for spurious emission and performance degradation in space telemetry systems is the imperfect data stream at the input of the modulator. The imperfection of the data stream can be caused by impalance between -1s and +1s (unbalanced data) and/or by data asymmetry.
There has been an emerging interest in developing and applying dictionary learning (DL) to process massive datasets in the last decade. Many of these efforts, however, focus on employing DL to compress and extract a set of important features from data, while considering restoring the original data from this set a secondary goal. On the other hand, although several methods are able to process streaming data by updating the dictionary incrementally as new snapshots pass by, most of those algorithms are designed for the setting where the snapshots are randomly drawn from a probability distribution. In this paper, we present a new DL approach to compress and denoise massive dataset in real time, in which the data are streamed through in a preset order (instances are videos and temporal experimental data), so at any time, we can only observe a biased sample set of the whole data. Here, our approach incrementally builds up the dictionary in a relatively simple manner: if the new snapshot is adequately explained by the current dictionary, we perform a sparse coding to find its sparse representation; otherwise, we add the new snapshot to the dictionary, with a Gram-Schmidt process to maintain the orthogonality. To compress and denoise noisy datasets, we apply the denoising to the snapshot directly before sparse coding, which deviates from traditional dictionary learning approach that achieves denoising via sparse coding. Compared to full-batch matrix decomposition methods, where the whole data is kept in memory, and other mini-batch approaches, where unbiased sampling is often assumed, our approach has minimal requirement in data sampling and storage: i) each snapshot is only seen once then discarded, and ii) the snapshots are drawn in a preset order, so can be highly biased. Through experiments on climate simulations and scanning transmission electron microscopy (STEM) data, we demonstrate that the proposed approach performs competitively to those methods in data reconstruction and denoising.
This paper compares the performances of two different types of data-derived symbol synchronizers, namely, Filter and Square (FS) and Digital Data Transition Tracking Loop (DTTL), in the presence of unbalanced data streams.
Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.
Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.