Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Arroyo Stream Processing Toolset (arroyopy) v0.1.0

Processing event or streaming data presents several technological challenges. A variety of technologies are often used by scientific user facilities. ZMQ is used to stream data and messages in a peer-to-peer fashion. Message brokers like Kafka, Redis Pubsub, EPICS PVA and RabbitMQ are often employed to route and pass messages from instruments to processing workflows. Arroyopy provides an API and structure to flexibly integrate with these tools and incorporate arbitrarily complex processing workflows, letting the hooks to the workflow code be independent of the connection code and hence reusable at a variety of instruments.

Chavez Esparza, Tanny Andrea [Lawrence Berkeley Na↗

Wind Energy Forecasting with the Weather Research and Forecasting Model

This was a collaborative effort between Lawrence Livermore National Security, LLC as manager and operator of Lawrence Livermore National Laboratory (LLNL) and Siemens Energy, Inc. (Siemens) to develop a wind resource forecasting tool. LLNL was to develop an independent high-resolution mesoscale modeling capability forecasting tool that could be implemented in conjunction with existing wind farm control and monitoring software to provide forecasting of wind resources using local observations of winds and temperature. Research with LLNL’s state-of-the-art large-eddy simulation meteorological prediction model, based on the community WRF model and innovative turbulence parameterizations, would improve that model’s applicability to large wind farms offshore and in complex terrain. The modeling capability would include uncertainty quantification. Finally, the application of the modeling tool and existing global climate change predictions would enable the delineation of the likely effects of climate change on wind resources. Siemens was to provide high time resolution hub-height wind speed and other meteorological data streams, including temperature profiles from wind farms, for LLNL to incorporate into the modeling system, to validate and tune this forecasting model for their locations of interest. These data streams would also be used for longer-term studies of correlations of wind resources to climate oscillations to indicate how long-term climate change trends may affect the available wind resource. Siemens would also provide information and observations of turbine wakes for incorporation into the modeling tool. By implementing state-of-the-art turbulence parameterizations into a simulation model and/or ensembles of simulation models, and by integrating real-time hub height wind speed and other meteorological datastreams from wind farms into that model or ensemble of models, LLNL would develop a forecasting tool that could be implemented by Siemens as an add-on to existing wind farm control and monitoring software to provide owners with useful resource forecasting. The desired outcome was that the accuracy level of the output would be sufficient to substantiate power output commitments. The final deliverable for this work would consist of a document outlining the algorithms and software tools that could be integrated into Siemens Wind Park Supervisor.

17 WIND ENERGY↗

MDDC Multi-Length Scale Data Architecture Contribution Report – PNNL, INL, ANL, LANL and ORNL

This report offers a comprehensive view of data streams currently generated at Pacific Northwest National Laboratory, Idaho National Laboratory, Argonne National Laboratory, Los Alamos National Laboratory, and Oak Ridge National Laboratory set to integrate into the evolving Multi-Dimensional Data Correlation framework at Oak Ridge National Laboratory. Developed by the Advanced Materials and Manufacturing Technologies program, the Multi-Dimensional Data Correlation framework serves as a cutting-edge software to manage data relevant to advanced manufacturing and material behavior in advanced reactors. The report defines data streams, highlights their generation methods and visualization methods both for experimental and computational aspects relevant to the Advanced Materials and Manufacturing Technologies project. A logical next step for this work is to integrate the MDDC framework into PNNL’s, INL’s, ANL’s, LANL’s and ORNL’s fabrication, experimentation, and modelling workflows. This would require setting up the MDDC framework at PNNL, INL, ANL, and LANL and integrating it into the data collection and storage for these different activities.

36 MATERIALS SCIENCE↗

KAZRARSCL-c0-Cloud Boundaries subset

The KAZR-ARSCL VAP provides cloud boundaries and best-estimate time-height fields of radar moments. The VAP merges corrected, but uncalibrated, KAZR moments from all active radar modes with cloud base and cloud mask observations from the micropulse lidar (MPL), cloud base from the ceilometer, as well information from soundings, rain gauge, and microwave radiometer instruments to produce two data streams, one with best-estimate cloud base and cloud layer boundaries, and another which also includes best-estimate time-height fields of radar moments. This DOI is for the data stream that contain cloud layer boundaries only. Please note that the reflectivity used in this level c0 product is uncalibrated.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS 2016 Sediment Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon

This dataset supports a broader synoptic effort to map morphological, hydrological, chemical, and biological conditions across a fifth-order mountain stream network. Samples were generated through a collaborative synoptic sampling effort in 2016. The dataset provides sediment Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) from 60 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Related data were collected as part of the event and were published separately in collaboration with other team members. The data are available at http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1567-2019 (Ward et al., 2019). The hydroshare data package contains processed FTICR-MS data from the samples included in this data package. The data were processed via Formultitude (previously called Formularity; https://github.com/PNNL-Comp-Mass-Spec/Formultitude). However, we have re-processed the data using Core-MS and included it in this data package. Additional related data collected in 2025 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3023310 and http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of sample data; (2) data dictionary; (3) file-level metadata; (4); (5) coordinates; and (6) readme. The sample data subfolder contains 12 Tesla (12T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, or .json.

Biogeochemistry↗

Assimilation of multiple datasets results in large differences in regional- to global-scale NEE and GPP budgets simulated by a terrestrial biosphere model

In spite of the importance of land ecosystems in offsetting carbon dioxide emissions released by anthropogenic activities into the atmosphere, the spatiotemporal dynamics of terrestrial carbon fluxes remain largely uncertain at regional to global scales. Over the past decade, data assimilation (DA) techniques have grown in importance for improving these fluxes simulated by terrestrial biosphere models (TBMs), by optimizing model parameter values while also pinpointing possible parameterization deficiencies. Although the joint assimilation of multiple data streams is expected to constrain a wider range of model processes, their actual benefits in terms of reduction in model uncertainty are still under-researched, also given the technical challenges. In this study, we investigated with a consistent DA framework and the ORCHIDEE-LMDz TBM–atmosphere model how the assimilation of different combinations of data streams may result in different regional to global carbon budgets. To do so, we performed comprehensive DA experiments where three datasets (in situ measurements of net carbon exchange and latent heat fluxes, spaceborne estimates of the normalized difference vegetation index, and atmospheric CO 2 concentration data measured at stations) were assimilated alone or simultaneously. We thus evaluated their complementarity and usefulness to constrain net and gross C land fluxes. We found that a major challenge in improving the spatial distribution of the land C sinks and sources with atmospheric CO 2 data relates to the correction of the soil carbon imbalance.

54 ENVIRONMENTAL SCIENCES↗

Beyond Energy Efficiency: A clustering approach to embed demand flexibility into building energy benchmarking

The intermittency of carbon-free renewables and the demand changes associated with the widespread push for electrifying the transportation and building sectors provides an opportunity for buildings to go beyond energy efficiency and push towards providing demand flexibility to the electricity grid. The duality of energy efficiency and demand flexibility is necessary for success in a sustainable and reliable energy transition. Current building energy benchmarking models are limited in their ability to integrate concepts of demand flexibility and/or utilize granular smart meter data. Thus, current benchmarking methods are focused annual energy usage and fail to incorporate how the time of use of energy consumption impacts emissions in a quickly changing energy grid. Without a more comprehensive view of energy usage and associated real-time emissions, current benchmarking methods are unlikely to realize the full decarbonization potential of buildings. New emerging data streams provide an opportunity to develop a new generation of benchmarking energy models that embed dimensions of energy efficiency, grid interactivity, and demand flexibility into their analysis. In this paper, we propose a four-step method for embedding grid interactivity and demand flexibility into building benchmarking models that utilizes emerging building and time-series electricity data streams. We first engineer features to produce a mix-type dataset that encompasses many attributes of grid-interactive and efficient buildings, and then we apply K-medoids using Gower's Distance to produce peer-group clusters. We apply the method to a case study of 306 primary and secondary schools in southern California, USA. The results show that the method effectively clusters buildings by attributes of demand flexibility and energy efficiency. The clustering results reveal patterns in inefficient building operations and demand inflexibility at the building peer group level. In conclusion, the interpretation of clusters can serve as an integrated energy efficiency and demand flexibility benchmarking model and inform performance-specific policy targeting for buildings that go beyond traditional efficiency measures.

24 POWER TRANSMISSION AND DISTRIBUTION↗

WHONDRS Surface Water and Sediment Geochemistry and Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon (v2)

This dataset supports a broader study developing conceptual models for river corridor critical zone processes across spatial scales and was generated in collaboration with the HJ Andrews River Corridor Critical Zone Workshop in 2025. The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen) from 48 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Some of the sites have been impacted by the Holiday Farm Fire and the Lookout Fire in 2020 and 2023, respectively. Related data were collected as part of the workshop and will be published separately in collaboration with other workshop attendees and available at http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. Related genomic data can be found on the National Center for Biotechnology Information (NCBI) under BioProject PRJNA1503030 (see critical details section below for more information). Additional related data collected in 2016 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3377027 and http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1-2019 (Ward et al., 2019). This data package was originally published in March 2026. It was updated in August 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos, (2) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, (3) a data checks report, (4) a folder of sample data, (5) file-level metadata, (6) data dictionary, (7) field metadata, (8) readme, (9) international generic sample number (IGSN) mapping file; and (10) field protocol. The sample data subfolder contains surface water and sediment (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages, (2) total dissolved nitrogen data and averages, (3) methods codes, (4) FTICR-MS methods; and (5) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the CoreMS processed data and seven subfolders, thee containing .xml files for each sample type (sediment, surface water and blank samples), three containing the sediment CoreMS output files for each sample type (sediment, surface water and blank samples), and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, .json, .jpg, or .jpeg.

Biogeochemistry↗

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada↗

Investigating Vegetation Responses to Underground Nuclear Explosions Through Integrated Analyses

Vegetation has the potential to respond to underground nuclear explosions, yet these links have not been fully explored. Given the lack of previously described signatures, the changes in vegetation are possibly subtle. The integration of multiple different data streams is potentially a useful approach to improve signal detection. Here, we investigate whether semi-arid vegetation growth patterns responded to eight legacy underground nuclear tests at the Nevada National Security Site in southern Nevada, USA. We tested for spatial and temporal changes in vegetation cover, tree growth patterns, and tree leaf spectral properties using ground-based measurements, including those from tree-rings and hyperspectral surface vegetation reflectance, as well as space-based measurements of Normalized Difference Vegetation Index (NDVI) from Landsat. Multiple data streams suggest a localized (<1.2 km) spatial pattern whereby tree growth is enhanced closer to the source of the underground test relative to sites further away. We also observed a more regional (>1.2–9 km) pattern whereby tree growth is suppressed coincident with a drought beginning 1 year before the 1989 tests, but continuing in the 5 years following the tests, which is anomalous relative to what is expected based on the response of tree growth to previous droughts. furthermore, quantification of the relative effects of the tests on vegetation remains a challenge due to the coincident drought and the potential for other disturbances to have impacted tree growth at this time, but the integration of these data reveals a more nuanced growth response than any other one data set indicates alone.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

BeyondPlanck: V. Minimal ADC Corrections for Planck LFI

We describe the correction procedure for Analog-to-Digital Converter (ADC) differential non-linearities (DNL) adopted in the Bayesian end-to-end BEYONDPLANCK analysis framework. This method is nearly identical to that developed for the official Planck Low Frequency Instrument (LFI) Data Processing Center (DPC) analysis, and relies on the binned rms noise profile of each detector data stream. However, rather than building the correction profile directly from the raw rms profile, we first fit a Gaussian to each significant ADC-induced rms decrement, and then derive the corresponding correction model from this smooth model. The main advantage of this approach is that only samples which are significantly affected by ADC DNLs are corrected, as opposed to the DPC approach in which the correction is applied to all samples, filtering out signals not associated with ADC DNLs. The new corrections are only applied to data for which there is a clear detection of the non-linearities, and for which they perform at least comparably with the DPC corrections. Out of a total of 88 LFI data streams (sky and reference load for each of the 44 detectors) we apply the new minimal ADC corrections in 25 cases, and maintain the DPC corrections in 8 cases. All these corrections are applied to 44 or 70 GHz channels, while, as in previous analyses, none of the 30 GHz ADCs show significant evidence of non-linearity. By comparing the BEYONDPLANCK and DPC ADC correction methods, we estimate that the residual ADC uncertainty is about two orders of magnitude below the total noise of both the 44 and 70 GHz channels, and their impact on current cosmological parameter estimation is small. However, we also show that non-idealities in the ADC corrections can generate sharp stripes in the final frequency maps, and these could be important for future joint analyses with the Planck High Frequency Instrument (HFI), Wilkinson Microwave Anisotropy Probe (WMAP), or other datasets. We therefore conclude that, although the existing corrections are adequate for LFI-based cosmological parameter analysis, further work on LFI ADC corrections is still warranted.

79 ASTRONOMY AND ASTROPHYSICS↗

Exploring Physics of Ferroelectric Domain Walls in Real Time: Deep Learning Enabled Scanning Probe Microscopy

The functionality of ferroelastic domain walls in ferroelectric materials is explored in real-time via the in situ implementation of computer vision algorithms in scanning probe microscopy (SPM) experiment. The robust deep convolutional neural network (DCNN) is implemented based on a deep residual learning framework (Res) and holistically nested edge detection (Hed), and ensembled to minimize the out-of-distribution drift effects. The DCNN is implemented for real-time operations on SPM, converting the data stream into the semantically segmented image of domain walls and the corresponding uncertainty. Further the pre-defined experimental workflows perform piezoresponse spectroscopy measurement on thus discovered domain walls, and alternating high- and low-polarization dynamic (out-of-plane) ferroelastic domain walls in a PbTiO 3 (PTO) thin film and high polarization dynamic (out-of-plane) at short ferroelastic walls (compared with long ferroelastic walls) in a lead zirconate titanate (PZT) thin film is reported. This work establishes the framework for real-time DCNN analysis of data streams in scanning probe and other microscopies and highlights the role of out-of-distribution effects and strategies to ameliorate them in real time analytics.

36 MATERIALS SCIENCE↗

Timeseries Unlabeled and Labeled Photos, Modeled Stream Elevation, and (Meta)Data of Variably Inundated Streams Across The Yakima River Basin, Washington, United States (v2)

This dataset is associated with the “River Monitoring Photos” (RMP) study and subsequent manuscript (Bao et al. 2025. Monitoring river flow status using low-cost wildlife camera and image segmentation artificial intelligence doi: 10.1016/j.envsoft.2025.106715). Game camera timeseries photos were collected to evaluate stream variable inundation via changes in width. A subset of photos was labeled for training the YOLOv8 and Mask2Former models and used to segment water surface fractions from all the game camera photos.This data package was originally published in March 2024. It was updated in October 2025 (v2) to add additional photos and files associated with the manuscript (i.e., processed data, labeled photos, and trained models). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to a readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; and (5) folders containing game camera photos and manuscript-associated files. Each Yakima River Basin site has a folder that contains subfolders for each month photos were collected. There is also a folder for files associated with the manuscript which has subfolders for labeled data, trained models, Yakima River Basin site water surface fractions, and USGS site water surface fractions. All files are .csv, .json, .txt, .yaml, .pth, .pt, or .pdf. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Revealing Local Structures through Machine-Learning-Fused Multimodal Spectroscopy

Atomistic structures of materials offer valuable insights into their functionality. Determining these structures remains a fundamental challenge in materials science, especially for systems with defects. While both experimental and computational methods exist, each has limitations in resolving nanoscale structures. Core-level spectroscopies, such as X-ray absorption (XAS) or electron energy-loss spectroscopies (EELS), have been used to determine the local bonding environment and structure of materials. Recently, machine learning (ML) methods have been applied to extract structural and bonding information from XAS/EELS data. However, frameworks relying solely on a single data stream, defined as characterization data derived from a single element using one technique, are often insufficient because multiple local environments can yield similar spectral features, making it challenging to differentiate between competing structural hypotheses. Here, in this work, we address this challenge by integrating multimodal ab initio simulations, experimental data acquisition, and ML techniques for structure characterization. Our goal is to determine local structures and properties using EELS and XAS data from multiple elements and edges. To showcase our approach, we use various lithium nickel manganese cobalt (NMC) oxide compounds which are used for lithium ion batteries, including those with oxygen vacancies and antisite defects, as the sample material system. We successfully inferred local element content, ranging from lithium to transition metals, with quantitative agreement with experimental data. Beyond local element inference, we find that ML model based on multimodal spectroscopic data is able to determine whether local defects such as oxygen vacancy and antisites are present, a task which is impossible for single mode spectra or other experimental techniques. Furthermore, our framework is able to provide physical interpretability, bridging spectroscopy with the local atomic and electronic structures.

battery↗

Application of Chebyshev’s Inequality in Online Anomaly Detection Driven by Streaming PMU Data

The day-to-day operation of modern power systems is highly reliant on prompt and adequate situational-awareness. This can be achieved via various system monitoring functions such as anomaly detection, in which static thresholds are commonly utilized to distinguish the normal and the abnormal system states. However, a predetermined static threshold usually lacks the flexibility to adapt to unobserved scenarios. In this paper, we propose two self-adaptive synchrophasor data driven anomaly detection approaches based on Chebyshev’s Inequality. The proposed approaches have been evaluated with Kundur’s 2area system and Mini-WECC system. Experimental results verify that the proposed approaches can dynamically adapt to unprecedented scenarios, and detect anomalous events with lower false alarm rate compared to static threshold based detection.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ESnet Secure Copy (EScp) v0.6

EScp is a high speed transfer tool with a similar command line syntax to scp. Unlike SCP it is designed to transfer files at high speed, thus far we have been able to show 100gbit/s transfers, although I expect that the throughput should scale in proportion to the network interface, i.e. I expect 400gbit/s performance on our 400gbit/s test bed. EScp achieves good performance through an innovative design (multithreaded, zero copy transfers), along with pluggable filters and I/O engines. As an example, you can switch from POSIX i/O to UIO by checking a different engine. It also natively supports encryption, and cheksums for file verification and transport security. AAA is through standard SSH (same as SCP). By taking advantage of filters, EScp supports transferring unstructured data and/or I/O to non-posix data sources. Examples include streaming data (i.e. from equipment), transferring data to the cloud, and/or supporting non-posix file systems (like HPSS).

Shiflett, Charles↗

Intelligent Monitoring Systems and Advanced Well Integrity and Mitigation

Long-term seismic monitoring of carbon capture and storage projects is needed to verify that the injected gas is safely stored in the subsurface until permanence can be assured. Conventional surface seismic monitoring techniques are usually expensive, require highly invasive surface operations, and need significant time investments on the part of personnel for both the field effort and processing the acquired data. For these reasons, permanent reservoir monitoring technologies are preferred, as they can offer a cost-effective solution for long-term monitoring. As part of the monitoring program of the Archer Daniels Midland’s large-scale injection of CO 2 in Decatur, Illinois, USA, a continuous seismic monitoring array was installed using a combination of surface orbital vibrator (SOV) sources and fiber-optic cables for distributed acoustic sensing (DAS) acquisition with the objective to build a continuous monitoring array. The aim of the presented project was to build a monitoring array and platform that integrates real-time seismic data with conventional data streams and provides continuous data analysis using dynamic computational models to deliver a comprehensive real-time assessment of subsurface conditions. It is in this context that the Intelligent Monitoring Systems and Advanced Well Integrity and Mitigation project was proposed with the objective to develop an integrated architecture that utilizes a permanent seismic monitoring network, combines the real-time geophysical and process data with reservoir flow and geomechanical models to create a comprehensive monitoring, visualization, and control system that delivers critical information for process surveillance and optimization.

54 ENVIRONMENTAL SCIENCES↗