Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Integrating LEO and GEO Observations: Toward Optimal Summertime Satellite Precipitation Retrieval

Abstract Reliable quantitative precipitation estimation with a rich spatiotemporal resolution is vital for understanding the Earth’s hydrological cycle. Precipitation estimation over land and coastal regions is necessary for addressing the high degree of spatial heterogeneity of water availability and demand, and for resolving the extremes that modulate and amplify hazards such as flooding and landslides. Advancements in computation power along with unique high spatiotemporal and spectral resolution data streams from passive meteorological sensors aboard geosynchronous Earth-orbiting (GEO) and low Earth-orbiting (LEO) satellites offer exciting opportunities to retrieve information about surface precipitation phenomena using data-driven machine learning techniques. In this study, the capabilities of U-Net–like architecture are investigated to map instantaneous, summertime surface precipitation intensity at the spatial resolution of 2 km. The calibrated brightness temperature products from the Global Precipitation Measurement (GPM) Microwave Imager (GMI) radiometer are combined with multispectral images (visible, near-infrared, and infrared bands) from the Advanced Baseline Imager (ABI) aboard the GOES-R satellites as main inputs to the U-Net–like precipitation algorithm. Total precipitable water and 2-m temperature from the Global Forecast System (GFS) model are also used as auxiliary inputs to the model. The results show that the U-Net–like algorithm can capture fine-scale patterns and intensity of surface precipitation at high spatial resolution over stratiform and convective precipitation regimes. The evaluations reveal the potential of extracting relevant, high spatial features over complex surface types such as mountainous regions and coastlines. The algorithm allows users to interpret the inputs’ importance and can serve as a starting point for further exploration of precipitation systems within the field of hydrometeorology.

Meteorology & Atmospheric Sciences↗

Fusing time-varying mosquito data and continuous mosquito population dynamics models

Climate change is arguably one of the most pressing issues affecting the world today and requires the fusion of disparate data streams to accurately model its impacts. Mosquito populations respond to temperature and precipitation in a nonlinear way, making predicting climate impacts on mosquito-borne diseases an ongoing challenge. Data-driven approaches for accurately modeling mosquito populations are needed for predicting mosquito-borne disease risk under climate change scenarios. Many current models for disease transmission are continuous and autonomous, while mosquito data is discrete and varies both within and between seasons. This study uses an optimization framework to fit a non-autonomous logistic model with periodic net growth rate and carrying capacity parameters for 15 years of daily mosquito time-series data from the Greater Toronto Area of Canada. The resulting parameters accurately capture the inter-annual and intra-seasonal variability of mosquito populations within a single geographic region, and a variance-based sensitivity analysis highlights the influence each parameter has on the peak magnitude and timing of the mosquito season. This method can easily extend to other geographic regions and be integrated into a larger disease transmission model. This method addresses the ongoing challenges of data and model fusion by serving as a link between discrete time-series data and continuous differential equations for mosquito-borne epidemiology models.

97 MATHEMATICS AND COMPUTING↗

Dissolved Oxygen and Temperature Data from the Hyporheic Zone of the East River Watershed July 2017 to October 2018

Dissolved oxygen (DO) is critical for aquatic ecosystems. Our focus is on the long-term DO dynamics in hyporheic zone of rivers, which are a function of both transport (hydrologic exchange between river and hyporheic zone) and uptake by biogeochemical reactions or respiration. The study site is the alpine East River watershed in Colorado, USA, meander A downstream from the pump house. Opti O2 probes were deployed in the water column and directly within the river-bed at 10, 20, and 35 cm depth (38°55'23.38"N, 106°57'4.21"W, 9046.88m) to monitor DO and temperature. A continuous data stream from July 24, 2017 to Oct 24, 2018 was autonomously telemetered to the cloud. This 14-month DO and temperature time series were obtained without any servicing for maintenance or data downloads; additionally the ability to remotely verify probe performance during field deployment was essential to confirm data validity during winter freeze-in and hydrological/weather events, such as spring melt and summer monsoons. We investigate the variations in dissolved oxygen dynamics of this snow-pack dominated watershed during a comparatively low flow water year (2018) and a relatively normal water year (2017), enabled by distinctive, in-situ, high frequency (∆t = 5min) sensors that provided a continuous time-series from the undisturbed study site over multiple seasons.

54 ENVIRONMENTAL SCIENCES↗

Congruity of genomic and epidemiological data in modelling of local cholera outbreaks

Cholera continues to be a global health threat. Understanding how cholera spreads between locations is fundamental to the rational, evidence-based design of intervention and control efforts. Traditionally, cholera transmission models have used cholera case-count data. More recently, whole-genome sequence data have qualitatively described cholera transmission. Integrating these data streams may provide much more accurate models of cholera spread; however, no systematic analyses have been performed so far to compare traditional case-count models to the phylodynamic models from genomic data for cholera transmission. Here, we use high-fidelity case-count and whole-genome sequencing data from the 1991 to 1998 cholera epidemic in Argentina to directly compare the epidemiological model parameters estimated from these two data sources. We find that phylodynamic methods applied to cholera genomics data provide comparable estimates that are in line with established methods. Our methodology represents a critical step in building a framework for integrating case-count and genomic data sources for cholera epidemiology and other bacterial pathogens.

59 BASIC BIOLOGICAL SCIENCES↗

Automated Integration of Continental-Scale Observations in Near-Real Time for Simulation and Analysis of Biosphere–Atmosphere Interactions

The National Ecological Observatory Network (NEON) is a continental-scale observatory with sites across the US collecting standardized ecological observations that will operate for multiple decades. To maximize the utility of NEON data, we envision edge computing systems that gather, calibrate, aggregate, and ingest measurements in an integrated fashion. Edge systems will employ machine learning methods to cross-calibrate, gap-fill and provision data in near-real time to the NEON Data Portal and to High Performance Computing (HPC) systems, running ensembles of Earth system models (ESMs) that assimilate the data. For the first time gridded EC data products and response functions promise to offset pervasive observational biases through evaluating, benchmarking, optimizing parameters, and training new machine learning parameterizations within ESMs all at the same model-grid scale. Leveraging open-source software for EC data analysis, we are already building software infrastructure for integration of near-real time data streams into the International Land Model Benchmarking (ILAMB) package for use by the wider research community. We will present a perspective on the design and integration of end-to-end infrastructure for data acquisition, edge computing, HPC simulation, analysis, and validation, where Artificial Intelligence (AI) approaches are used throughout the distributed workflow to improve accuracy and computational performance.

Durden, David J.↗

Machine learning on FPGA for event selection

Real-time data processing is a frontier field in experimental particle physics. The application of FPGAs at the trigger level is used by many current and planned experiments (CMS, LHCb, Belle2, PANDA). Usually they use conventional processing algorithms. LHCb has implemented Machine Learning (ML) elements for real-time data processing with a triggered readout system that runs most of the ML algorithms on a computer farm. The work described in this article aims to test the ML-FPGA algorithms for streaming data acquisition. Herein, there are many experiments working in this area and they have a lot in common, but there are many specific solutions for detector and accelerator parameters that are worth exploring further. This report describes the purpose of the work and progress in evaluating the ML-FPGA application.

47 OTHER INSTRUMENTATION↗

An Integrated Platform for Collaborative Data Analytics

While collaboration among data scientists is a key to organizational productivity, data analysts face significant barriers to achieving this end, including data sharing, accessing and configuring the required computational environment, and a unified method of sharing knowledge. Each of these barriers to collaboration is related to the fundamental question of knowledge management “how can organizations use knowledge more effectively?”. In this paper, we consider the problem of knowledge management in collaborative data analytics and present ShareAL, an integrated knowledge management platform, as a solution to that problem. The ShareAL platform consists of three core components: a full stack web application, a dashboard for analyzing streaming data and a High Performance Computing (HPC) cluster for performing real time analysis. Prior research has not applied knowledge management to collaborative analytics or developed a platform with the same capabilities as ShareAL. ShareAL overcomes the barriers data scientists face to collaboration by providing intuitive sharing of data and analytics via the web application, a shared computing environment via the HPC cluster and knowledge sharing and collaboration via a real time messaging application.

Oesch, T↗

Modular Autonomous Experimentation for Biological Applications (Full Report)

The Modular Autonomous Research System (MARS) was developed to address the pressing need for faster, more reliable, and more adaptable scientific discovery. Traditional experimentation is limited by manual labor, long cycle times, and fragmented data streams, which constrain the ability to explore complex chemical and materials design spaces. To overcome these limitations, we created an integrated, modular platform that combines laboratory robotics, diverse measurement instruments, and a central data infrastructure with artificial intelligence–driven decision-making. The system links liquid handling robots, robotic arms, and optical plate readers into a closed loop where experiments are executed automatically, data is analyzed in real time, and subsequent experimental conditions are adaptively chosen to maximize information gain. Over the course of the project, MARS was validated on two primary test cases—spectroscopic metal–ligand binding assays and peptide-directed mineralization—which highlighted the system’s ability to handle uncertainty and variability in experimental measurements. To further demonstrate modularity and extensibility, we also established additional testbeds in electrochemistry for catalyst discovery and electrolyte formulation for advanced batteries. The results show that MARS can reliably conduct autonomous campaigns with minimal human intervention, adapt to distinct scientific domains, and provide a scalable model for future self-driving laboratories. This work establishes new capabilities for modular, uncertainty-aware automation and directly supports the need for advanced, data-driven research platforms capable of accelerating discovery across a wide range of scientific and national security missions.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning at the edge enables real-time streaming ptychographic imaging

Abstract Coherent imaging techniques provide an unparalleled multi-scale view of materials across scientific and technological fields, from structural materials to quantum devices, from integrated circuits to biological cells. Driven by the construction of brighter sources and high-rate detectors, coherent imaging methods like ptychography are poised to revolutionize nanoscale materials characterization. However, these advancements are accompanied by significant increase in data and compute needs, which precludes real-time imaging, feedback and decision-making capabilities with conventional approaches. Here, we demonstrate a workflow that leverages artificial intelligence at the edge and high-performance computing to enable real-time inversion on X-ray ptychography data streamed directly from a detector at up to 2 kHz. The proposed AI-enabled workflow eliminates the oversampling constraints, allowing low-dose imaging using orders of magnitude less data than required by traditional methods.

36 MATERIALS SCIENCE↗

HERMES

HERMES - High-speed Event Retrieval and Management for Enhanced Spectral imaging code. This code is meant to unpack and process neutron imaging data from the TPX3Cam made by Amsterdam Scientific Instruments. The TPX3Cam utilizes the Timepix3 chip in a single photon counting mode of image acquisition. The photon counting data streaming off the TPX3Cam needs to be process and analyzed in order to create final images. Therefore we are developing both python and cpp codes to allow for users of the TPX3Cam to efficiently analyze data and create images.

Long, Alexander↗

Data Reduction for Science: Brochure from the Advanced Scientific Computing Research Workshop

Data reduction for science holds promise for addressing the challenges of moving, storing, and processing massive data sets produced by the scientific community. Pursuing the PRDs outlined here will enable advances in data streaming, fast feedback and/or autonomous control of experiments, and faster time to scientific insight. These advances will result in a significant improvement in the ability to transport, store, process and interpret experimental, observational, and computational data.

97 MATHEMATICS AND COMPUTING↗

Status of the data acquisition, trigger, and slow control systems of the Mu2e experiment at Fermilab

The Mu2e experiment at the Fermilab will search for a coherent neutrinoless conversion of a muon into an electron in the field of an aluminum nucleus with a sensitivity improvement by a factor of 10,000 over existing limits. In this work, the Mu2e Trigger and Data Acquisition System (TDAQ) uses otsdaq framework as the online Data Acquisition System (DAQ) solution. Developed at Fermilab, otsdaq integrates several framework components — an artdaq-based DAQ, an art-based event processing, and an EPICS-based detector control system (DCS), and provides a uniform multi-user interface to its components through a web browser. Data streams from the Mu2e tracker and calorimeter are handled by the artdaq-based DAQ and processed by a one-level software trigger implemented within the art framework. Events accepted by the trigger have their data combined, post-trigger, with the separately read out data from the Mu2e Cosmic Ray Veto system. Foundation of the Mu2e DCS, EPICS – an Experimental Physics and Industrial Control System – is an open-source platform for monitoring, controlling, alarming, and archiving. A prototype of the TDAQ and the DCS systems has been built and tested over the last three years at Fermilab’s Feynman Computing Center, and now the production system installation is underway. This work presents their status and focus on the installation plans and procedures for racks, workstations, network switches, gateway computers, DAQ hardware, slow controls implementation, and testing.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Unsupervised anomaly clustering via offset alignment in multivariate grid sensing data

Modern industries increasingly rely on multi-sensor technologies to acquire complex, high-dimensional data streams, enabling advanced monitoring and control systems. One critical application is online anomaly detection in electrical smart grids, where multivariate and multimodal sensing technologies play a vital role. However, detecting anomalies in such time-series data is challenging due to their inherent temporal dependencies and stochastic behavior. Traditional approaches based on supervised and semi-supervised learning methods depend on labeled datasets, which are often unavailable in real-world scenarios. While unsupervised methods have emerged as promising alternatives, these methods are highly susceptible to noise and outliers commonly present in sensing applications. Furthermore, deep learning-based anomaly detection methods, despite their performance, are often criticized for their black-box nature, limiting their applicability in safety-critical and online environments where interpretability and explainability are paramount. In this work, we propose an unsupervised anomaly clustering method leveraging a cyclic alignment-based offset detection algorithm for multivariate time-series signals. The proposed method is applied to multivariate data collected from vibrational, voltage, and magnetic field sensors deployed in a local grid substation. Our results demonstrate the robustness of the algorithm in accurately clustering various anomalies/events across different sensing modalities. Additionally, we compare the effectiveness of the proposed approach against a simple pattern-based anomaly detection method, which performs well for univariate data but fails to generalize to multivariate and multimodal time-series data.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Slow control and TDAQ systems installation and tests in the Mu2e experiment

The Mu2e experiment at Fermilab will attempt to detect a coherent neutrinoless conversion of a muon into an electron in the field of an aluminum nucleus, with a sensitivity that is 10,000 times greater than existing limits. The Mu2e trigger and data acquisition system (TDAQ) uses the otsdaqframework as its online Data Acquisition System (DAQ) solution. Developed at Fermilab, otsdaq integrates several components, such as an artdaq-based DAQ, an art-based event processing, and an EPICS-based detector control system (DCS), and provides a uniform multi-user interface toits components through a web browser. The data streams from the Mu2e tracker and calorimeter are handled by the artdaq-based DAQ and processed by a one-level software trigger implemented within the art framework. Events accepted by the trigger have their data combined, post-trigger, with the separately read-out data from the Mu2e Cosmic Ray Veto system. The foundation of Mu2e DCS, EPICS, an Experimental Physics and Industrial Control System, is an open-source platform for monitoring, controlling, alarming, and archiving. Over the last three years, a prototype ofthe TDAQ and DCS systems has been built and tested at Fermilab’s Feynman Computing Center.Currently, the production system installation is underway. At the end, this work presents a brief update on the installation of racks and DAQ hardware.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Operational Focused Data Analytics for Optimizing Radiation Portal Monitor-Based Nuclear Smuggling Detection Systems at Global Ports of Entry

The National Nuclear Security Administration’s Office of Nuclear Smuggling Detection and Deterrence has deployed a fleet of radiation portal monitors (RPMs) across the world at global ports of entry including seaports, airports, and land border crossings. These RPMs are integrated into radiation detection systems (RDS) that also include fixed cameras, optical character recognition (OCR) systems, primary scanning systems (e.g., X-ray or gamma-ray), and secondary scanning systems (e.g., spectroscopic radiation portal monitors, portable radiation detection systems). The data from these sensing technologies is collected at the Central Alarm Station (CAS) where servers and computers reside to control and operate the system. Operators utilize the data collected by the CAS and declared cargo information to make decisions on how to respond to an alarm.This work explores the use of CAS-located data, looking at both the sensor data streams and operator inputs, to perform analysis which supports customs and border protection agencies to improve training capability and operational effectiveness. We focus on analyzing site level effectiveness and behavior by rolling up CAS-located data collected from individual occurrences. To-date, more than 15 sites (e.g., seaports, airports, border crossings) have been analyzed in this manner with the goal of understanding system operations to verify effectiveness and recommend potential improvements. This work first aims to provide background information on relevant CAS-located data sources and our current operational system analytics process including example results. After summarizing our current analytic techniques, we discuss how the future data analytics systems can provide key benefits to improving operational performance while minimizing the burden these detection systems place on operators.

Kuhn, Michael↗

Hysteresis Patterns of Watershed Nitrogen Retention and Loss Over the Past 50 years in United States Hydrological Basins

Patterns of watershed nitrogen (N) retention and loss are shaped by how watershed biogeochemical processes retain, biogeochemically transform, and lose incoming atmospheric deposition of N. Loss patterns represented by concentration, discharge, and their associated stream exports are important indicators of integrated watershed N retention behaviors. We examined continental United States (CONUS) scale N deposition (e.g., wet and dry atmospheric deposition), vegetation trends, and stream trends as potential indicators of watershed N-saturation and retention conditions, and how watershed N retention and losses vary over space and time. By synthesizing changes and modalities in watershed nitrogen loss patterns based on stream data from 2200 U.S. watersheds over a 50 years record, our work revealed two patterns of watershed N-retention and loss. One was a hysteresis pattern that reflects the integrated influence of hydrology, atmospheric inputs, land-use, stream temperature, elevation, and vegetation. The other pattern was a one-way transition to a new state. We found that regions with increasing atmospheric deposition and increasing vegetation health/biomass patterns have the highest N-retention capacity, become increasingly N-saturated over time, and are associated with the strongest declines in stream N exports—a pattern, that is, consistent across all land cover categories. We provide a conceptual model, validated at an unprecedented scale across the CONUS that links instream nitrogen signals to upstream mechanistic landscape processes. Our work can aid in the future interpretation of in-stream concentrations of DOC and DIN as indicators of watershed N-retention status and integrators of watershed hydrobiogeochemical processes.

58 GEOSCIENCES↗

Quillinan, et al 2018 DOE Geothermal Technology Office REE Report for NEWTS Database and Case Studies

Produced water data processed into the NEWTS data format for easy input into aqueous chemistry modeling software, including oil & gas and coal bed methane produced waters, and geothermal waters. Includes information on rare earth elements (REEs) and critical minerals (CMs). Case studies are included to demonstrate solved streams using aqueous chemistry software. An input template is provided for modeling stream data in OLI Studio. Original data from: Quillinan, Scott, Nye, Charles, Engle, Mark, Bartos, Timothy T., Neupane, Ghanashyam, Brant, Jonathan, Bagdonas, Davin, McLing, Travis, McLaughlin, J. Fred, Phillips, Erin, Hallberg, Laura L., Shahabadi, Mahdi, and Johnson, Matthew. Assessing rare earth element concentrations in geothermal and oil and gas produced waters: A potential domestic source of strategic mineral commodities (Final Report). United States: N. p., 2018. Web. doi:10.2172/1509037.

Aqueous Chemistry↗