Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Simple and Scalable Streaming: The GRETA Data Pipeline

The Gamma Ray Energy Tracking Array (GRETA) is a state of the art gamma-ray spectrometer being built at Lawrence Berkeley National Laboratory to be first sited at the Facility for Rare Isotope Beams (FRIB) at Michigan State University. A key design requirement for the spectrometer is to perform gamma-ray tracking in near real time. To meet this requirement we have used an inline, streaming approach to signal processing in the GRETA data acquisition system, using a GPU-equipped computing cluster. The data stream will reach 480 thousand events per second at an aggregate data rate of 4 gigabytes per second at full design capacity. We have been able to simplify the architecture of the streaming system greatly by interfacing the FPGA-based detector electronics with the computing cluster using standard network technology. A set of highperformance software components to implement queuing, flow control, event processing and event building have been developed, all in a streaming environment which matches detector performance. Prototypes of all high-performance components have been completed and meet design specifications.

Cromaz, Mario↗

Real-time nuclear activation detectors for measuring neutron angular distributions at the National Ignition Facility (invited)

The Real Time Nuclear Activation Detector (RTNAD) array at NIF measures the distribution of 14 MeV neutrons emitted by deuterium-tritium (DT) fueled inertial confinement fusion implosions. The uniformity of the neutron distribution is an important indication of implosion symmetry and DT shell integrity. The array consists of 48 LaBr 3 (Ce) crystal gamma-ray spectrometers mounted outside the NIF target chamber, which continuously monitor the slow decay of the 909 keV gamma-ray line from activated 89 Zr located in Zr cups surrounding each crystal. The measured decay rate dramatically increases during a DT implosion in proportion to the number of 14 MeV neutrons striking each Zr cup. The neutrons produce activated 89 Zr through an (n, 2n) reaction on 90 Zr, which is insensitive to low energy neutrons. The neutron flux along the detector line-of-sight at shot time is determined by extrapolating the fitted 909 keV decay curve back to shot time. Automatic analysis algorithms were developed to handle the non-stop data stream. The large number of detectors and the high statistical accuracy of the array enable the spherical harmonic modes of the neutron angular distribution to be measured up to L ≤ 4 to provide a better understanding of implosion dynamics. In addition, these data combined with measurements of the down-scattered neutrons can be used to derive fuel areal density distributions. This paper will describe the RTNAD hardware and analysis procedures.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Multimodal sensor fusion framework for residential building occupancy detection

For several years now, smart building energy systems have been a research area of intensive activity. In light of the increasing need for sustainable buildings and energy systems, this trend motivates an increasing need for a solution to reduce carbon dioxide emissions and improve energy efficiency. This work proposes a high-performing and transferable occupancy detection framework that combines sensor data from different data modalities, including time series environmental data (temperature, humidity, and illuminance), image data, and acoustic energy data using ensemble method. To draw out the best prediction performance in each modality, the proposed framework was developed, including various models that were designed to learn the occupancy patterns reflected in the physical data streams. To tackle the time series environmental data, we designed two variants of an occupancy detection spatiotemporal pattern network (Occ-STPN) that performs both feature level and decision level fusion, respectively. We also propose a new metric; the fading memory mean square error (FMMSE), that provides a fair evaluation and penalization of delayed occupancy predictions. Multiple open-sourced datasets, including the Electricity Consumption and Occupancy and the University of California, Irvine's (UCI) building occupancy detection dataset, along with our own real data collected from six different houses, were used to validate the algorithms' performance. The experimental results presented herein break down the performance for each sensing modality, and a detailed analysis of the performance is also discussed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nuclear data uncertainty propagation and modeling uncertainty impact evaluation in neutronics core simulation

Uncertainty analysis is a critical requirement in reactor simulation as it is used to quantify the reliability of best-estimate calculation. A comprehensive uncertainty analysis should characterize all sources of uncertainties in a computationally-feasible and scientifically-defendable manner. Here we employ a well-established reduced order modeling (ROM) based uncertainty quantification methodology to propagate uncertainties throughout neutronic calculations. ROM relies on recent advances in randomized data mining techniques applied to large data streams. In our proposed implementation, the nuclear data uncertainties are first propagated from multi-group level through lattice physics calculation to generate few-group parameter uncertainties, described using a vector of mean values and a covariance matrix. Employing an ROM-based compression of the covariance matrix, the few-group uncertainties are then propagated through downstream core simulation in a computationally efficient manner. This straightforward approach, albeit efficient as compared to brute force forward and/or adjoint-based methods, often employs a number of assumptions that have been unquestioned in the literature of neutronic uncertainty analysis. This manuscript argues that these assumptions could introduce another source of uncertainty referred to as modeling uncertainties, whose magnitude needs to be quantified in tandem with nuclear data uncertainties. Thus, our primary goal is to explore the interactions between these two uncertainty sources in order to assess whether modeling uncertainties have an impact on parameter uncertainties. To explore this endeavor, the impact of a number of modeling assumptions on core attributes uncertainties is quantified. The study employs a CANDU reactor model, with Serpent and NEWT as lattice physics solvers and NESTLE-C as core simulator. The modeling assumptions investigated include those related with the uncertainty propagation method employed, e.g., deterministic vs. stochastic, the few-group energy structure employed to represent the cross-sections, the resonance treatment in lattice physics calculation, the reference values for the cross-section, and the number of samples employed to render ROM compression. Results indicate that some of the modeling assumptions could have a non-negligible impact on the core responses propagated uncertainties, highlighting the need for a more comprehensive approach to combine parameter and modeling uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Wattile: Probabilistic Deep Learning-based Forecasting of Building Energy Consumption [SWR-20-94]

Accurate energy forecasting is becoming critical due to many reasons: i ) optimal distributed energy resources operations and dispatch, ii) fault detection and diagnostics, and iii) meeting operational energy efficiency targets. Wattile uses deep learning (DL) for the building's short-term load forecasting application. Two specific types of neural networks called, Long Short Term Memory (LSTM) and Sequence-to-Sequence (S2S) models are used to make predictions. Forecasting models are trained using online historical weather and occupancy indicator data streams from the Intelligent Campus Program's data acquisition systems at the National Renewable Energy Laboratory (NREL) for main meters and sub-meters of multiple building types. These models use probabilistic methods to provide quantile-based forecasts in addition to nominal conditional median predictions of electricity consumption.

Frank, Stephen↗

Data fusion enhanced multi-modality wellbore integrity inspection system

A downhole multi-modality inspection system includes a first imaging device operable to generate first imaging data and a second imaging device operable to generate second imaging data. The first imaging device includes a first source operable to emit energy of a first modality, and a first detector operable to detect returning energy induced by the emitted energy of the first modality. The second imaging device includes a second source operable to emit energy of a second modality, and a second detector operable to detect returning energy induced by the emitted energy of the second modality. The system further includes a processor configured to receive the first imaging data and the second imaging data, and integrate the first imaging data with the second imaging data into an enhanced data stream. The processor correlates the first imaging data and the second imaging data to provide enhanced data for detecting potential wellbore anomalies.

Kasten, Ansas Matthias↗

FTICR-MS, Sensor, and Environmental Data from 5 Streams Impacted by the 2020 Holiday Farm Fire Associated with: "Spatiotemporal controls on the delivery of dissolved organic matter to streams following a wildfire"

This data package is associated with the publication "Spatiotemporal Controls on the Delivery of Dissolved Organic Matter to Streams Following a Wildfire" submitted to Geophysical Research Letters (Roebuck et al., 2022). The study aims to understand storm induced transport of pyrogenic materials to streams impacted by varying degrees of burn severity. Time series samples (24 samples in 1-hour intervals) were collected at 5 sites within the McKenzie River Watershed (Oregon, USA) whose catchment were each completely engulfed by the 2020 Holiday Farm Fire. The samples were collected in November 2020 during the first major storm pulse following the conclusion of the wildfire. Samples were characterized for dissolved organic carbon, total dissolved nitrogen, and by ultra-high resolution mass spectrometry. In situ turbidity data also collected.This data package contains 4 primary folders that include the following: 1) Metadata, 2) EnvData (Environmental Data), 3) SensorData, and 4) FTICR_SupportingData. The package contains a single file-level metadata (flmd) file. Each primary folder also contains individual data dictionaries (dd) to define and provide descriptors of column/row headers and data flags. The FTICR_SupportingData, folder 4, contains raw, unprocessed FTICR-MS Data files in addition to a csv containing processed FTICR-MS data. This package contains the following file types: csv, xml, pdf.

54 ENVIRONMENTAL SCIENCES↗

Real-time data processing for serial crystallography experiments

We report the use of streaming data interfaces to perform fully online data processing for serial crystallography experiments, without storing intermediate data on disk. The system produces Bragg reflection intensity measurements suitable for scaling and merging, with a latency of less than 1 s per frame. Our system uses the CrystFEL software in combination with the ASAP::O data framework. In a series of user experiments at PETRA III, frames from a 16 megapixel Dectris EIGER2 X detector were searched for peaks, indexed and integrated at the maximum full-frame readout speed of 133 frames per second. The computational resources required depend on various factors, most significantly the fraction of non-blank frames ('hits'). The average single-thread processing time per frame was 242 ms for blank frames and 455 ms for hits, meaning that a single 96-core computing node was sufficient to keep up with the data, with ample headroom for unexpected throughput reductions. Further significant improvements are expected, for example by binning pixel intensities together to reduce the pixel count. We discuss the implications of real-time data processing on the `data deluge' problem from recent and future photon-science experiments, in particular on calibration requirements, computing access patterns and the need for the preservation of raw data.

47 OTHER INSTRUMENTATION↗

Evaluating lightweight unsupervised online IDS for masquerade attacks in CAN

Vehicular controller area networks (CANs) are susceptible to masquerade attacks by malicious adversaries. In masquerade attacks, adversaries silence a targeted ID and then send malicious frames with forged content at the expected timing of benign frames. As masquerade attacks could seriously harm vehicle functionality and are the stealthiest attacks to detect in CAN, recent work has devoted attention to compare frameworks for detecting masquerade attacks in CAN. However, most existing works report offline evaluations using CAN logs already collected using simulations that do not comply with the domain’s real-time constraints. Here we contribute to advance the state of the art by presenting a comparative evaluation of four different non-deep learning (DL)-based unsupervised online intrusion detection systems (IDS) for masquerade attacks in CAN. Our approach differs from existing comparative evaluations in that we analyze the effect of controlling streaming data conditions in a sliding window setting. In doing so, we use realistic masquerade attacks being replayed from the ROAD dataset. We show that although evaluated IDS are not effective at detecting every attack type, the method that relies on detecting changes in the hierarchical structure of clusters of time series produces the best results at the expense of higher computational overhead. We discuss limitations, open challenges, and how the evaluated methods can be used for practical unsupervised online CAN IDS for masquerade attacks.

Anomaly detection↗

HAIMOS Ensemble Forecasts for Intra-day and Day- Ahead GHI, DNI and Ramps

The objective of this research is to develop a hybrid physics-based/data-driven forecast model to improve direct normal and global horizontal irradiance (DNI and GHI) prediction for horizons ranging from 1 to 72 hours. Project objectives also address key gaps in state-of-the-art solar forecasting: accurate probabilistic solar forecasts and the forecasting of large irradiance ramps (ramp onset and magnitude). The proposed model ensembles Numerical Weather Prediction (NWP) forecasts, determinist physics-based algorithms, and new-generation cloud cover products (high-resolution rapid refresh satellite images and Large Eddy Simulations). The result is the Hybrid Adaptive Input Model Objective Selection (HAIMOS) ensemble model. HAIMOS blends state of the art machine learning methodologies with physics-based models for cloud cover and cloud optical depth forecasts. The technical activities followed a two-pronged strategy. First, the preprocessing of data, the selection of inputs to the nonlinear approximators, the type of approximator and objective functions, and post-processing ensembling techniques included in HAIMOS were all optimized adaptively to find the best model for a specific goal (reduce DNI/GHI forecast error, improve the prediction of ramp onset, etc.). Second, a large effort was put in improving cloud identification and the forecast of cloud cover and cloud optical depth. To this end, new-generation cloud parametrization products were developed in this work. These include improved algorithms to assist in cloud identification, cloud classification and cloud parametrization from satellite images – three key factors in the accuracy of 1 to 6-hours irradiance forecasts and prediction of ramp onset. Furthermore, we also included cloud information extracted from high resolution rapid refresh satellite images (GOES-16) and Large Eddy Simulations (LES). LES was used to model the atmosphere in detail over locations of interest and produce cloud optical depth forecasts. Once these data streams were validated, they were used as input data to the HAIMOS forecast. The model was developed using data from several climatologically distinct locations with potential for high solar penetration. In the last year of the project, we conducted a validation campaign according to the guidelines stipulated by the Topic Area 1 project as described in the FOA. This effort brings, for the first time, proven machine-learning methodologies for generating state-of-the-art solar forecasts interweaved with detailed physics-based models for cloud detection, and cloud optical depth forecasts. HAIMOS will generate accurate irradiance probabilistic forecast to assist in reducing solar generation prediction error. Globally optimized solar forecast models are more likely to impact solar energy stakeholders. The goal of this project was to increase the state-of-the-art forecast skill from their present values of 10 to 35%. At the end of the project, we achieved between 30% and 50% forecast skill across a wide range of horizons for both GHI and DNI.

14 SOLAR ENERGY↗

Streaming Compression of Scientific Data via Weak-SINDy

Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.

97 MATHEMATICS AND COMPUTING↗

In Situ Transmission Electron Microscopy: Signal processing challenges and examples

Transmission electron microscopy (TEM) is a powerful tool for imaging material structure and characterizing material chemistry. Recent advances in data collection technology for TEM have enabled high-volume and high-resolution data collection at a microsecond frame rate. Here, taking advantage of these advances in data collection rates requires the development and application of data processing tools, including image analysis, feature extraction, and streaming data processing techniques. In this article, we highlight a few areas in materials science that have benefited from combining signal processing and statistical analysis with data collection capabilities in TEM and present a future outlook on opportunities of integrating signal processing with automated TEM data analysis.

36 MATERIALS SCIENCE↗

The LAKE model input dataset for three Arctic lakes

This dataset contains meteorological data collected for three Arctic lakes and compiled to satisfy input requirements of the LAKE 2.0 model. The dataset was generated to act as a benchmarking dataset for future model-data inter-comparisons. The LAKE 2.0 model simulates temperatures within the water later and the sedimentary layer of a lake. The LAKE2.0. is an open-source code and available to download via this weblike http://tesla.parallel.ru/Viktor/LAKE/-/wikis/LAKE-model (last visit July 14, 2021). The meteorological data are required to simulate the surface energy balance at the surface of a lake. This dataset includes a compilation of the meteorological data pulled from multiple data streams, including National Oceanic and Atmospheric Administration (NOAA) climate data, Circumarctic Lakes Observation Network (CALON) data, and the United States Geological Survey (USGS) data. The data were compiled for three Arctic lakes: FoxDen (66.55877, -164.45670), Atqasuk (70.452497, -156.951984), and Toolik (68.63150, -149.60740). Each meteorological data is in comma-delimited format (file extension ‘.dat’) and includes eight columns: Temperature [K], Pressure [Pa], longwave downward radiation [W/m2], shortwave downward radiation [W/m2], “U” wind speed [m/s], ”V” wind speed [m/s], humidity [kg/kg], precipitation [m/s]. In addition to the meteorological data file, we included setup and driver files. The Toolik lake is the deepest out of three lakes and has inflowing and outflowing groundwater data. InflowOutflowREADME.txt has more information about inflow and outflow flies. The other two lakes are much shallower and modeled as a closed system (i.e. no water inflow or outflow).

54 ENVIRONMENTAL SCIENCES↗

Dendrometer data at The Morton Arboretum Forestry Plots 2019-2023

We are collecting long-term dendrometer data at The Morton Arboretum to determine seasonal growth patterns in trees. This data on stem growth patterns will eventually be integrated with other ongoing data streams to paint a broader picture of plant phenology. This data package contains the outputs from three types of dendrometer devices: ICT band dendrometer data is in the "ICT_data2020_2023.csv" file, TreeHugger band dendrometer data is in the "TreeHugger_data2019_2020.csv" file, and TOMST point dendrometer data is in the "TOMST_data2021_2023.csv" file. The TreeHugger and TOMST files contain corrected and raw uncorrected values, whereas the ICT data file contains only raw data. Initial tree size data at device installation is located in the "Initial_Tree_Size.csv" file. Additional information on units are contained within each data file's respective data dictionary, and the location metadata file contains the geographic locations of the 23 forestry plots at The Morton Arboretum as well as the measured tree species at each location.

54 ENVIRONMENTAL SCIENCES↗

Process Design and Techno-Economic Analysis of the Modular Staged Pressurized Oxy-Combustion (SPOC) Power Plant for Biomass

This work describes the process design and techno-economic analysis (TEA) of the modular SPOC power plant for biomass firing and coal-biomass co-firing. Two Rankine cycles were considered: a supercritical steam cycle (242 bar, 593°C, 593°C) with 550 MWe net output and a subcritical cycle (166 bar, 566°C, 566°C) with 200 MWe net output. For both cases, 95% carbon capture was modeled, and hybrid poplar biomass was chosen to generate carbon-negative power. In addition, the supercritical 500 MWe case included a 25% biomass co-firing (carbon neutral) case. For both cycles, a 100% Powder River Basin coal firing case was used for comparison purposes. In the SPOC process, oxygen is produced via a cryogenic air separation unit (ASU) and the heat generated from the compression of air is integrated into the steam cycle and utilized for boiler feed water pre-heating. Unique to the SPOC process, the boilers are pressurized and arranged in a series-parallel configuration, with minimized flue gas recirculation. The flue gas is cooled and scrubbed in the direct-contact cooler (DCC) column, and the moisture in the flue gas is condensed, leaving the bottom of the DCC at a sufficiently high temperature such that it can be used for boiler feed water pre-heating, improving plant thermal efficiency. Following drying and purification, CO2 in the flue gas is at the purity required for storage or utilization. The performance data were obtained from process modelling via Aspen Plus®. The stream data from Aspen Plus® were used as an input for the AACE Class 5 cost study. Ultimately, the capital costs, Levelized Cost of Electricity (LCOE), and cost of CO2 captured and avoided were obtained. The HHV efficiency of the carbon negative 550 MWe supercritical SPOC case (34.8%) was clearly above those reported by NETL for the BECCS baseline cases of supercritical pulverized coal with capture (B12B, 31.5%) and the 49% biomass co-firing case with capture (PA3, 29.2%). The HHV efficiency of the carbon-negative subcritical plant is also higher than the subcritical baseline PC plant with capture (case B11B.95) presented by NETL (32% vs 29.7%). The LCOE for the SPOC 100% biomass case was similar to the LCOE for the BECCS 49% biomass with carbon capture case ($147/MWh), and the SPOC carbon neutral case LCOE was lower ($110/MWh) than the cost for the NETL baseline SC coal firing case with 90% carbon capture ($114/MWh).

Magalhaes, Duarte↗

California - Leosphere Windcube 866 (120), Humboldt / Reviewed Data

The purpose of this dataset is to provide filtered, averaged lidar data and standardize the data format of various data streams from the buoy into NetCDF. The attached Lidar Buoy Data Dictionary provides further details on the various instruments mounted on the buoys, parameters measured by each instrument, and the frequency of data collection.

17 WIND ENERGY↗

Buoy - California - Wind Sentinel (130), Morro Bay - Processed Data

The purpose of the dataset is to provide preliminary filtered, averaged buoy data and standardize the data format of various data streams from the buoy into NetCDF. The attached Lidar Buoy Data Dictionary provides further details on the various instruments mounted on the buoys, parameters measured by each instrument, and the frequency of data collection.

17 WIND ENERGY↗

Lidar - California - Leosphere Windcube 866 (130), Morro Bay - Processed Data

The purpose of this dataset is to provide preliminary filtered, averaged lidar data and standardize the data format of various data streams from the buoy into NetCDF. The attached Lidar Buoy Data Dictionary provides further details on the various instruments mounted on the buoys, parameters measured by each instrument, and the frequency of data collection.

17 WIND ENERGY↗