Big data in nuclear comes in the form of many structures, systems, components, data streams, data types, and more
Big Data in nuclear comes in the form of many structures, systems, components, data streams, data types, and more.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Big Data in nuclear comes in the form of many structures, systems, components, data streams, data types, and more.
Various Internet of Things (IoT) devices generate complex, dynamically changed, and infinite data streams. Adversaries can cause harm if they can access the user’s sensitive raw streaming data. For this reason, protecting the privacy of the data streams is crucial. In this paper, we explore local differential privacy techniques for streaming data. We compare the techniques and report the advantages and limitations. We also present the effect on component (e.g., smoother, perturber) variations of distribution-based local differential privacy. We find that combining distribution-based noise during perturbation provides more flexibility to the interested entity.
Streaming data is data that is emitted at variable volumes in a continuous, incremental manner with the goal of low-latency processing often at a different physical location. Network infrastructure is used to facilitate the connection between data sources and sinks, and must be robust to handle the requirements of the workflow. The U.S. Department of Energy Office of Science (DOE SC) a federal agency supporting fundamental scientific research for energy and the Nation’s largest supporter of basic research in the physical sciences. DOE SC has the responsibility for operating $\mathbf{1 0}$ National Laboratories, and 28 scientific user facilities supporting advanced supercomputers, particle accelerators, large x-ray light sources, neutron scattering sources, and other specialized facilities for nanoscience and genomics. This paper investigates the state of streaming data workfows, and details some of the approaches to this challenging problem.
In this paper we demonstrate direct data streaming from instruments and detectors at a large-scale experimental facility to a supercomputer for real-time data processing and feedback. Streaming data to supercomputers introduces the potential for novel scientific applications and workflow models, including the ability to provide real-time feedback from very large datasets during an experiment and the integration of real-time ML training and inference at scale. We discuss a successful demonstration for real-time processing of data from the Advanced Photon Source (APS) on the Polaris supercomputer using an EPICS-based streaming framework. We describe the capabilities of the streaming framework itself, and outline the architecture that allows us to process experimentally derived data on a supercomputer without file-based data transfers. We present throughput measurements that are indicative of system performance capable of sustaining the expected data production rates of the facility, as well as discuss some outstanding challenges and our future directions.
With the exponential growth in the volume and complexity of data generated at high-energy physics and nuclear physics research facilities, there is an imperative demand for innovative strategies to process this data in real or near-real-time. Given the surge in the requirement for high-performance computing, it becomes pivotal to reassess the adaptability of current data processing architectures in integrating new technologies and managing streaming data. This paper introduces the ERSAP framework, a modern solution that synergizes flow-based programming with the reactive actor model, paving the way for distributed, reactive, and high performance in data stream processing applications. Additionally, we unveil a novel algorithm focused on time-based clustering and event identification in data streams. The efficacy of this approach is further exemplified through the data-stream processing outcomes obtained from the recent beam tests of the EIC prototype calorimeter at DESY.
Implementing a physics data processing application is relatively straightforward with the use of current containerization technologies and container image runtime services, which are prevalent in most high-performance computing (HPC) environments. However, the process is complicated by the challenges associated with data provisioning and migration, impacting the ease of workflow migration and deployment. Transitioning from traditional file-based batch processing to data-stream processing workflows is suggested as a method to streamline these workflows. This transition not only simplifies file provisioning and migration but also significantly reduces the necessity for extensive disk space. Data-stream processing is particularly effective for real-time processing during data acquisition, thereby enhancing data quality assurance. This paper introduces the integration of the JLAB CLAS12 event reconstruction application within the ERSAP data-stream processing framework that facilitates the execution of streaming event reconstruction at a remote data center and enables the return streaming of reconstructed events to JLAB while circumventing the need for temporary data storage throughout the process.
We consider two streams of data or measurements with disparate qualities and time resolutions that need to be classified. The first stream consists of higher quality data at a coarser time resolution, and the other consists of lower quality data at a finer time resolution. We present a fuser-switch method that fuses the set of classifiers of each stream separately and switches between them. We show that this method provides classification decisions at a finer time resolution with superior detection and false alarm probabilities compared to individual classifiers, under the statistical independence and time resolution ratio conditions. When classifiers are trained using machine learning methods, we show that this superior performance is guaranteed with a confidence probability specified by the classifiers' generalization equations. We use these results to provide analytical foundations for previous practical results that achieved significant performance improvements in classifying Pu/Np target dissolution events at a radiochemical processing facility.
In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.
This paper presents an reactive, actor-model and FBP paradigm based framework that we develop to design data-stream processing applications for HEP and NP. This framework encourages a functional decomposition of the overall data processing application into small mono-functional artifacts. Artifacts that are easy to understand, develop, deploy and debug. The fact that these artifacts (actors) are programmatically independent they can be scaled and optimized independently, which is impossible to do for components of the monolithic application. One of the important advantages of this approach is fault tolerance where independent actors can come and go on the data-stream without forcing the entire application to crash. Furthermore, it also makes it is easy to locate the faulty actor in the data pipeline. Due the fact that the actors are loosely coupled, and that the data (inevitably) carries the context, they can run on heterogeneous environments, utilizing different accelerators. This paper describes the main design concepts of the framework and presents a ?proof of concept? application design and deployment results obtain processing on-beam calorimeter streaming data.
With ongoing automation and digitization of the electric power system, several Phasor Measurement Units(PMUs) have been deployed for monitoring and control. PMU data can have multiple anomalies, and many of the researchers in the past have concentrated on training machine/deep learning algorithms offline for anomaly detection over PMU data (i.e., not in real time). These machine/deep learning algorithms, when trained offline on a sample rather than a population of the dataset, fail to consider the dynamic behavior of the power grid in real-time, resulting in low accuracy. Considering the dynamic behavior of the power grid (e.g., change in load, generation, distributed energy resources (DERs) switching, network, controls), the definition of data anomalies varies in time and requires online training. A fundamental challenge is to enable online (i.e., real-time) training of machine/deep learning algorithms for anomaly detection over streaming PMU data. While machine/deep learning is often desirable to manage data streams, training a deep learning algorithm over streaming PMU data is nontrivial due to changes in data statistics caused by dynamic streaming data. This paper proposes PMUNET: a novel device-level deep learning-based data-driven approach for anomaly detection, localization, and classification over streaming PMU data, using online learning and multivariate data-drift detection algorithm .Two variants of PMUNET, Dynamic data Change Driven Learning (DCDL) and Continuity Driven Learning (CDL), are proposed and compared. DCDL aims to train the deep learning algorithm whenever the definition of anomaly changes due to the power grid dynamics. On the other hand, CDL continuously trains the deep learning algorithm over the PMU data-stream. The experimental results verify that DCDL outperforms CDL and other efficient anomaly detection methods over multiple events such as faults and load/ generator/capacitor/DERs variations/switching for IEEE 14 and 39 Bus test system as well as real PMU industrial data. The result verifies that DCDL variant of PMUNET improves over existing approach with a gain of 2% - 10% in terms of accuracy, false-positive rate, and false-negative rate.
There has been an emerging interest in developing and applying dictionary learning (DL) to process massive datasets in the last decade. Many of these efforts, however, focus on employing DL to compress and extract a set of important features from data, while considering restoring the original data from this set a secondary goal. On the other hand, although several methods are able to process streaming data by updating the dictionary incrementally as new snapshots pass by, most of those algorithms are designed for the setting where the snapshots are randomly drawn from a probability distribution. In this paper, we present a new DL approach to compress and denoise massive dataset in real time, in which the data are streamed through in a preset order (instances are videos and temporal experimental data), so at any time, we can only observe a biased sample set of the whole data. Here, our approach incrementally builds up the dictionary in a relatively simple manner: if the new snapshot is adequately explained by the current dictionary, we perform a sparse coding to find its sparse representation; otherwise, we add the new snapshot to the dictionary, with a Gram-Schmidt process to maintain the orthogonality. To compress and denoise noisy datasets, we apply the denoising to the snapshot directly before sparse coding, which deviates from traditional dictionary learning approach that achieves denoising via sparse coding. Compared to full-batch matrix decomposition methods, where the whole data is kept in memory, and other mini-batch approaches, where unbiased sampling is often assumed, our approach has minimal requirement in data sampling and storage: i) each snapshot is only seen once then discarded, and ii) the snapshots are drawn in a preset order, so can be highly biased. Through experiments on climate simulations and scanning transmission electron microscopy (STEM) data, we demonstrate that the proposed approach performs competitively to those methods in data reconstruction and denoising.
Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.
Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.
Here, this work introduces an online greedy method for constructing quadratic manifolds from streaming data, designed to enable in situ analysis of numerical simulation data on the Petabyte scale. Unlike traditional batch methods, which require all data to be available upfront and take multiple passes over the data, the proposed online greedy method incrementally updates quadratic manifolds in one pass as data points are received, eliminating the need for expensive disk input/output operations as well as storing and loading data points once they have been processed. A range of numerical examples demonstrate that the online greedy method learns accurate quadratic manifold embeddings while being capable of processing data that far exceed common disk input/output capabilities and volumes as well as main-memory sizes.
Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.
Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.
The mapping of computational needs onto execution resources is, by and large, a manual task, and users are frequently guided simply by intuition and past experiences. We present a queueing theory based performance model for streaming data applications that takes steps towards a better understanding of resource mapping decisions, thereby assisting application developers to make good mapping choices. The performance model (and associated cost model) are agnostic to the specific properties of the compute resource and application, simply characterizing them by their achievable data throughput. We illustrate the model with a pair of applications, one chosen from the field of computational biology and the second is a classic machine learning problem.
The “IO Wall” problem, in which the gap between computation rate and data access rate grows continuously, poses significant problems to scientific workflows which have traditionally relied upon using the filesystem for intermediate storage between workflow stages. One way to avoid this problem in scientific workflows is to stream data directly from producers to consumers and avoiding storage entirely. However, the manner in which this is accomplished is key to both performance and usability. This paper presents the Sustainable Staging Transport, an approach which allows direct streaming between traditional file writers and readers with few application changes. SST is an ADIOS “engine”, accessible via standard ADIOS APIs, and because ADIOS allows engines to be chosen at run-time, many existing file-oriented ADIOS workflows can utilize SST for direct application-to-application communication without any source code changes. This paper describes the design of SST and presents performance results from various applications that use SST, for feeding model training with simulation data with substantially higher bandwidth than the theoretical limits of Frontier’s file system, for strong coupling of separately developed applications for multiphysics multiscale simulation, or for in situ analysis and visualization of data to complete all data processing shortly after the simulation finishes.