Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series: Preprint

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as abrupt time series data shifts. These shifts, which are usually unintentional, may be caused by software or hardware configuration changes on a PV system, and do not reflect an actual change in overall system performance. Locating these shifts and segmenting the associated time series aids in more accurate future PV analysis. In this research, an offline changepoint detection (CPD) algorithm that automatically detects these abrupt data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window for the labeled data sets. To promote further research in this space, we are releasing the labeled data shift sets on U.S. Department of Energy's (DOE) DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package.

changepoint detection↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as abrupt time series data shifts. These shifts, which are usually unintentional, may be caused by software or hardware configuration changes on a PV system, and do not reflect an actual change in overall system performance. Locating these shifts and segmenting the associated time series aids in more accurate future PV analysis. In this research, an offline changepoint detection (CPD) algorithm that automatically detects these abrupt data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window for the labeled data sets. To promote further research in this space, we are releasing the labeled data shift sets on U.S. Department of Energy's (DOE) DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package.

changepoint detection↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as time series data shifts. In this research, a changepoint detection (CPD) algorithm that automatically detects data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window across the labeled data sets. Pending approval, we plan to release the labeled data sets for this research on NREL's DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package. By supplying the training sets and algorithm, we hope to encourage further development in this research space.

data shift↗

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

J Lemery↗

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

Medical Operations↗

Earth Independent Medical Operations (EIMO) DATASCOPE Technical Interchange Meeting 21st August 2023: Background and Summary of Discussion

An aspiration for EIMO datascope is to realize artificial intelligence-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A vision proposed to the meeting participants was that of a “system of systems,” whereby EIMO will utilize AI-supported natural language processing and machine learning techniques to synthesize embedded reference databases and real-time data streams [input vectors] from multiple data sources to continuously and seamlessly assess crew health & performance. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will ideally have a degree of mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats.

Artificial Intelligence↗

Integrated rate isolation sensor

In one embodiment, a system for providing fault-tolerant inertial measurement data includes a sensor for measuring an inertial parameter and a processor. The sensor has less accuracy than a typical inertial measurement unit (IMU). The processor detects whether a difference exists between a first data stream received from a first inertial measurement unit and a second data stream received from a second inertial measurement unit. Upon detecting a difference, the processor determines whether at least one of the first or second inertial measurement units has failed by comparing each of the first and second data streams to the inertial parameter.

Brady, Tye↗

Overview and How-to Tutorial Videos for Using NEWTS Data

Overview and How-to Tutorial Videos for Using NEWTS Data Video 1: Overview of NEWTS Database Video 2: How-to tutorial for EPA Flue Gas Desulfurization (FGD) Effluent NEWTS dataset Video 3: How-to tutorial for USGS Produced Waters NEWTS dataset Video 4: How-to tutorial for EPA Ash NEWTS dataset Video 5: How-to tutorial for Quillinan, et al 2018 DOE Geothermal Technology Office REE dataset Video 6: Tutorial video on navigating the NEWTS Dashboard, with an overview of NEWTS and navigating between the NEWTS Dashboard and Datasets (https://netl-doe.maps.arcgis.com/apps/dashboards/a5fa4192f7c6478dab3d6180d9c30b84) Video 7: Additional tutorial video on navigating the NEWTS Dashboard and investigating specific data points in the Dashboard and Datasets Video 8: Re-record of recent webinar giving an overview of the NEWTS Database and Dashboard, including interacting with the NEWTS Dashboard, locating specific data points, and finding the relevant streams in the NEWTS Database and datasets on EDX. Includes overview of the datasets, case studies, and steps for taking stream data from the database and modeling stream data in OLI Studio and Geochemist's Workbench. Note: Video 3 tutorial is also applicable to the USGS Brackish Water NEWTS dataset.

Aqueous Chemistry↗

A variable-data-rate, multimode quadriphase modem.

This paper describes the design and performance of a highly versatile modulator and demodulator recently developed to facilitate the evaluation of various digital communications links. The modem is capable of either PSK or QPSK operation and can accommodate a very wide range of continuously tunable data rates (1 kbps to 30 Mbps in each of two channels). In the QPSK mode, operation is possible using either a single serial data stream (single channel operation) or using two mutually independent, unrelated, and asynchronous data streams (dual-channel operation). Integrate and dump detectors are used at the demodulator for regeneration of the data stream(s). Measurements indicate that the performance of the overall system (including the bit detectors) is within 2 dB of the theoretically optimum performance of either PSK or QPSK at any rate within the range of rates provided by the modem, and is within 1 dB of theoretical over most of the range of rates.

Allen, R. W.↗

The magnetic-field investigation on ISPM

The International Solar Polar Mission (ISPM) onboard instrumentation for the magnetic field experiment to establish, on the basis of in-situ observations, the heliolatitude dependence of the interplanetary magnetic field, is described. The prime output consists of vector measurements, made by two triaxial magnetometers, of the ambient magnetic field along the orbit of the spacecraft. The onboard data processor generates two data streams to be transmitted through the spacecraft telemetry. The low speed, analog data stream consists of averaged and despun vector measurements, digitized in the spacecraft analog to digital converter (ADC). The despinning algorithm used in the analog processor is described. The high speed, digital data stream consists of vector measurements digitized in the instrument ADCs, generating up to two vector samples per sec. Multiple data-path switching is used to increase system reliability and to allow cross calibration of the ADCs. On board facilities for inflight calibration are described.

Balogh, A.↗

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

StOKeDMD: Streaming Occupation kernel dynamic mode decomposition

Dynamic mode decomposition (DMD) has become a common technique for constructing surrogate models for dynamical systems from observed system states. The Occupation Kernel DMD (OKDMD) method proposed in (Rosenfeld et al., 2022) and (Rosenfeld et al., 2024) is a Liouville operator based method that builds surrogate models from system state trajectories. Here, this paper proposes an extension of OKDMD to the case when the system states are observed in a streaming fashion, i.e., only a small fraction of the state trajectory is available at a given time. The developed method, Streaming Occupation Kernel DMD (StOKeDMD), accommodates the streaming data input by leveraging properties of specific choices of kernel functions and occupation kernels. We apply the StoKeDMD method as a compression method for streaming data, analyze the memory complexity, and demonstrate the performance of StoKeDMD in the compression of streaming data generated from a Lorenz system and a fluid flow simulation.

97 MATHEMATICS AND COMPUTING↗

The IRAS faint source survey

The principal features of the IRAS Faint Source Survey (FSS), a new product resulting from the extended IRAS mission, are reviewed. The FSS has achieved an increase in sensitivity of about a factor of 2.5 relative to the IRAS Point Source Catalog by coadding the data before extracting sources. The FSS was produced by point-source filtering the individual detector data streams and then coadding the data streams using a trimmed-average algorithm. The discussion covers FSS production methods; reliability, completeness, and positional accuracy of the FSS; and FSS view of the IR sky.

Moshir, Mehrdad↗

ARM Data as A Resource for Validation of NASA PACE Cloud Retrievals

NASA’s Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission will launch in January 2024 and continue and improve upon satellite data records in its eponymous domains. PACE will carry a broad-swath hyperspectral imager, OCI, which will provide MODIS/VIIRS-type cloud data products (i.e. a cloud mask, top height, visible optical thickness, droplet effective radius, phase, and derived water path). It will also carry two multi-angle polarimeters (HARP2 and SPEXone) which will not only provide the above but also enable retrievals of additional cloud properties (e.g. droplet effective variance, ice crystal asymmetry parameter). Validating satellite-based cloud retrievals is challenging. We plan to use several ARM data streams to evaluate PACE cloud data products and are prototyping our analyses using retrievals from MODIS on the Aqua satellite and OLCI on the Sentinel-3A satellite. This poster shows how we plan to use ARM data to evaluate liquid water path (via MWRRET) and cloud top height retrievals (via KARZASRCL), with example results from these proxy sensors and ARM data from the SGP, ENA, and NSA sites. We seek comments from and collaborations with the ARM community to get the most out of our respective data streams.

ARM↗

Distributing Data to Hand-Held Devices in a Wireless Network

ADROIT is a developmental computer program for real-time distribution of complex data streams for display on Web-enabled, portable terminals held by members of an operational team of a spacecraft-command-and-control center who may be located away from the center. Examples of such terminals include personal data assistants, laptop computers, and cellular telephones. ADROIT would make it unnecessary to equip each terminal with platform- specific software for access to the data streams or with software that implements the information-sharing protocol used to deliver telemetry data to clients in the center. ADROIT is a combination of middleware plus software specific to the center. (Middleware enables one application program to communicate with another by performing such functions as conversion, translation, consolidation, and/or integration.) ADROIT translates a data stream (voice, video, or alphanumerical data) from the center into Extensible Markup Language, effectuates a subscription process to determine who gets what data when, and presents the data to each user in real time. Thus, ADROIT is expected to enable distribution of operations and to reduce the cost of operations by reducing the number of persons required to be in the center.

Hodges, Mark↗

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele↗

The influence of environmental microseismicity on detection and interpretation of small-magnitude events in a polar glacier setting

Glacial environments exhibit temporally variable microseismicity. To investigate how microseismicity influences event detection, we implement two noise-adaptive digital power detectors to process seismic data from Taylor Glacier, Antarctica. We add scaled icequake waveforms to the original data stream, run detectors on the hybrid data stream to estimate reliable detection magnitudes and compare analytical magnitudes predicted from an ice crack source model. We find that detection capability is influenced by environmental microseismicity for seismic events with source size comparable to thermal penetration depths. When event counts and minimum detectable event sizes change in the same direction (i.e. increase in event counts and minimum detectable event size), we interpret measured seismicity changes as ‘true’ seismicity changes rather than as changes in detection. Generally, one detector (two degree of freedom (2dof)) outperforms the other: it identifies more events, a more prominent summertime diurnal signal and maintains a higher detection capability. We conclude that real physical processes are responsible for the summertime diurnal inter-detector difference. One detector (3dof) identifies this process as environmental microseismicity; the other detector (2dof) identifies it as elevated waveform activity. Our analysis provides an example for minimizing detection biases and estimating source sizes when interpreting temporal seismicity patterns to better infer glacial seismogenic processes.

54 ENVIRONMENTAL SCIENCES↗

Towards Resilient Near Real-Time Analysis Workflows in Fusion Energy Science

Nuclear fusion holds the promise of an endless source of energy. Several research experiments across the world and joint modeling and simulation efforts between the nuclear physics and high performance computing communities are actively preparing the operation of the International Thermonuclear Experimental Reactor (ITER). Both experimental reactors and their simulated counterparts generate data that must be analyzed quickly and in a resilient way to support decision making for the configuration of subsequent runs or prevent a catastrophic failure. However, the cost if the traditional techniques used to improve the resilience of analysis workflows, i.e., replicating datasets and computational tasks, becomes prohibitive with explosion of the volume of data produced by modern instruments and simulations. Therefore, we advocate in this paper for an alternate approach based on data reduction and data streaming. The rationale is that by allowing for a reasonable, controlled, and guaranteed loss of accuracy it becomes possible to transfer smaller amounts of data, shorten the execution time of analysis workflows, and lower the cost of replication to increase resilience. We develop our research and development roadmap towards resilient near real-time analysis workflows in fusion energy science and present early results showing that data streaming and data reduction is a promising way to speed up the execution and improve the resilience of analysis workflows.

Suter, Fred↗