Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “streaming data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

The MAMS Quick View System-2 (QVS2) - A workstation for NASA aircraft scanner data evaluation

This paper describes a ground-based data-evaluation workstation named Quick View System-2 (QVS2) developed to support postflight evaluation of data supplied by the Multispectral Atmospheric Mapping Sensor (MAMS), one of the four spectrometers that can be used with the Daedalus scanner flown on the ER-2 aircraft. The QVS2 provides advanced analysis capabilities and can be applied to other airborne scanners used throughout NASA for earth-system-science investigations, because of the commonality in the data stream and in the generalized data structure.

Jedlovec, Gary J.↗

Sustainable Biosphere Initiative Project

The goal of the Advanced Technology in Ecological Sciences project is to gain broad participation within the environmental scientific community in developing a research agenda addressing the development and refinement of technologies instrumental to research that responds to these challenges (e.g. global climate change, unsustainable resource use, and threats to biological diversity). The following activities have been completed: (1) A listserve 'eco-tech was set up to serve as a clearinghouse of information about activities and events relating to advanced technologies; (2) A series of conference calls were organized on specific topics including data visualization and spatial analysis, and remote sensing; and (3) Two meetings were organized at the 19% ESA Annual Meeting in Providence, Rhode Island. Topics covered included concerns about tool and data sharing; interest in expanded development of ground-based remote sensing technologies for monitoring; issues involved in training for using new technologies and increasing data streams, and- associated implications of data processing capabilities; questions about how to develop appropriate standards (i.e. surface morphology classification standards) that facilitate the exchange and comparison of analytical results; and some thoughts about remote sensing platforms and vehicles.

Source record↗

Conquering Data Chaos: Research Data Management with Kubernetes

Managing massive volumes of data and effectively making it accessible to researchers poses significant challenges and is a barrier to scientific discovery. In many cases, critical data is locked up in unwieldy file formats or one-off databases and is too large to effectively process on a single machine. This talk explores the role of Kubernetes, an open-source container orchestration platform, in addressing research data management challenges. I will discuss how we are using a set of publicly available open-source and home-grown tools in the National Renewable Energy Lab (NREL) Data, Analysis, and Visualization (DAV) group to help researchers overcome data-related bottlenecks. The talk will begin by providing an overview of the data challenges faced in research data management, including data storage, processing, and analysis. I will highlight Kubernetes' ability to handle large-scale data by leveraging containerization and distributed computing, including distributed storage. Kubernetes allows researchers to encapsulate data processing infrastructure and workflows into portable containers, enabling reproducibility and ease of deployment. Kubernetes can then schedule and manage the resource allocation of these containers to enable efficient utilization of limited computing resources, leading to more efficient data processing and analysis. I will discuss some limitations of traditional, siloed approaches to dealing with data and emphasize the need for solutions which foster collaboration. I will highlight how we are using Kubernetes at NREL to facilitate data sharing and cooperation among research teams. Kubernetes' flexible architecture enables the deployment of shared computing environments, such as Apache Superset, where researchers can seamlessly access and analyze shared datasets. Providing the ability to have one research team easily consume data generated by another, utilizing Kubernetes' as a central data platform, is one of the major wins we've encountered by adopting the platform. Finally, I will showcase real-world use cases from NREL where we have used Kubernetes to solve some persistent data challenges involving large volumes of sensor and monitoring data. I will discuss the challenges we encountered when creating our cluster and making it available as a production-ready resource. I will also discuss the specific suite of tools, including Postgres and Apache Druid for columnar and timeseries data, and Redpanda Kafka for streaming data we have deployed in our infrastructure, and the process that went into the selection of these tools.

collaborative environment↗

1.2 Mfps standalone X-ray detector for Time-Resolved Experiments

We present a standalone and autonomous X-ray detector capable of operation with the speed of up to1.2Mfps. The detector utilizes UFXC32k hybrid pixel detectors for sensing X-rays, Spartan-6 LX45 FPGA placed in commercially available sbRIO 9628 controller for data acquisition and processing including a compression with zero-suppression algorithm. A Linux-RT system working on the 400 MHz Dual-Core CPU is used for FPGA control and data streaming to the higher-level system over 1 Gbps Ethernet connection. 1.2 M frames per second is achieved in so-called burst mode of operation while in zerodead-time mode 70 kfps is possible. Due to efficient data compression in FPGA there’s no need of using high-speed transceivers and Frame-Grabber cards on the data server side and the detector can stream the data infinitely over standard 1 Gbps network connection. Operation modes were tested at Advanced Photon Source Synchrotron at Argonne National Laboratory.

47 OTHER INSTRUMENTATION↗

Driving Next-Generation Workflows from the Data Plane

We observe the emergence of a new generation of scientific workflows that process data produced at a sustained rate by scientific instruments and large scale numerical simulations. This data is consumed by multiple analysis, visualization, or Machine Learning components not only to enable inference and justify the scientific program, but also to monitor and steer the evolution of these experiments. In such workflows, moving intermediate data efficiently is key to performance, more than efficiently scheduling computational tasks. However, most traditional workflow management systems focus on optimizing task scheduling and then deal with data management, assuming a “move little, compute for long” model, which makes them unfit to the efficient management of this new generation of workflows. Therefore, we advocate for a new way to manage scientific workflows. We propose to consider an efficiently and independently managed data plane that can store and stream data. Workflows compute components, in the application plane can then interact with the data plane, abstracted from complexities of data management. Then, the role of a workflow management system would become that of a control plane that allows users to connect services together to execute the workflow and manages connections between the application and data planes. In this position paper, we characterize several next-generation workflow motifs and describe how their interaction with the data plane is a challenge to traditional workflow management systems. Then, we express a set of requirements that a workflow management system should meet to efficiently manage next-generation workflows at different scales. Based on these requirements, we expose our vision of driving next-generation workflows from the data plane and list remaining open challenges.

Suter, Fred↗

Attack on Grid Event Cause Analysis: An Adversarial Machine Learning Approach

With the ever-increasing reliance on data for data-driven applications in power grids, such as event cause analysis, the authenticity of data streams has become crucially important. The data can be prone to adversarial stealthy attacks aiming to manipulate the data such that residual-based bad data detectors cannot detect them, and the perception of system operators or event classifiers changes about the actual event. This paper investigates the impact of adversarial attacks on convolutional neural network-based event cause analysis frameworks. We have successfully verified the ability of adversaries to maliciously misclassify events through stealthy data manipulations. The vulnerability assessment is studied with respect to the number of compromised measurements. Furthermore, a defense mechanism to robustify the performance of the event cause analysis is proposed. The effectiveness of adversarial attacks on changing the output of the framework is studied using the data generated by real-time digital simulator (RTDS) under different scenarios such as type of attacks and level of access to data.

Niazazari, Iman↗

Advanced Computing, Data Science, and Artificial Intelligence Research Opportunities for Energy-Focused Transportation Science

The Energy Efficient Mobility Systems (EEMS) technology landscape is complex and rapidly evolving, which provides both tremendous opportunities and formidable challenges. Significant alterations to the mobility landscape are underway due to the advent of vehicle and infrastructure connectivity, autonomous driving, and rapid passenger- and freight-vehicle electrification. Advanced computing will play an increasingly important role in enabling the EEMS program to understand and identify the most important levers to improve the energy productivity of future integrated mobility systems. It is also driving new approaches to mobility and the research to unlock an affordable, efficient, safe, and accessible transportation future. Driving much of this change is the collection, analysis, and strategic use of massive amounts of diverse, complex data from infrastructure and vehicles with on-board sensors and data storage and transmission capabilities. Diverse and representative data are key to implementing approaches to maximize mobility energy productivity. While high-fidelity modeling of integrated transportation networks has strengthened our understanding of dynamic movement and behavior patterns, existing tools must be expanded beyond their current focus. This work necessitates data infrastructure investments (e.g., secure-streaming data platforms driven by ubiquitous sensors and video analytics) as well as investments in critical capabilities for large-scale automated analysis and organization using modern machine learning, statistics, and artificial intelligence. Other chief needs include agile, large-scale storage that can be quickly searched and queried for relevant data to support validation and model development, data-sharing agreements, and formatting standards for key data types. The future of public transit must be explored in greater detail, research must inform design, and opportunities must be identified for improving the mobility productivity of public transit in both urban and rural America.

33 ADVANCED PROPULSION SYSTEMS↗

The ABCs of On-Demand Transit (ODT)

On-demand mobility - also referred to as on-demand transit (ODT) - is a form of public mobility that is flexible with respect to where and when service is provided, and ODT deployments have increased significantly in recent years. Transit agencies are becoming increasingly interested in ODT, and due to differing definitions and various service design and business model options, it can be difficult to learn about the emerging industry. This work provides an overview of the definitions of ODT, recent trends internationally and in the U.S., ODT's benefits and challenges (particularly compared to fixed-route transit), three primary service design options, system costs and funding considerations, and a metrics framework for evaluating ODT systems to ensure continued successful performance. Identified benefits include increased service areas, short ride and wait times, increased user flexibility, potential to reduce energy consumption and emissions through shared trips and smaller, right-sized vehicles, increased safety and comfort through door-to-door service, and rich data streams including granular spatio-temporal data that can be analyzed to continuously improve the service. Challenges include scaling ODT service up as small increases in ridership require additional supply to keep service quality high, serving peak times including keeping low wait times, the lack of fixed schedule being challenging for commuters, integrating ODT services with nearby transit systems, and equity for riders without smartphones who cannot track the vehicle in a mobile app. Finally, an overview of seven ODT case studies (in Texas, Missouri, New York, and Ontario, Canada) performed by NREL and related analysis of travel time, energy and emissions, costs, and equity are presented. Initial key findings include: ODT can be cost- and energy-effective compared to fixed-route transit, ODT serves more people than other transit options, and ODT system deployments can be followed by rapid growth.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Memory-based frame synchronizer

A frame synchronizer for use in digital communications systems wherein data formats can be easily and dynamically changed is described. The use of memory array elements provide increased flexibility in format selection and sync word selection in addition to real time reconfiguration ability. The frame synchronizer comprises a serial-to-parallel converter which converts a serial input data stream to a constantly changing parallel data output. This parallel data output is supplied to programmable sync word recognizers each consisting of a multiplexer and a random access memory (RAM). The multiplexer is connected to both the parallel data output and an address bus which may be connected to a microprocessor or computer for purposes of programming the sync word recognizer. The RAM is used as an associative memory or decorder and is programmed to identify a specific sync word. Additional programmable RAMs are used as counter decoders to define word bit length, frame word length, and paragraph frame length.

Stattel, R. J.↗

Soft Decision Analyzer

The Soft Decision Analyzer (SDA) is an instrument that combines hardware, firmware, and software to perform realtime closed-loop end-to-end statistical analysis of single- or dual- channel serial digital RF communications systems operating in very low signal-to-noise conditions. As an innovation, the unique SDA capabilities allow it to perform analysis of situations where the receiving communication system slips bits due to low signal-to-noise conditions or experiences constellation rotations resulting in channel polarity in versions or channel assignment swaps. SDA s closed-loop detection allows it to instrument a live system and correlate observations with frame, codeword, and packet losses, as well as Quality of Service (QoS) and Quality of Experience (QoE) events. The SDA s abilities are not confined to performing analysis in low signal-to-noise conditions. Its analysis provides in-depth insight of a communication system s receiver performance in a variety of operating conditions. The SDA incorporates two techniques for identifying slips. The first is an examination of content of the received data stream s relation to the transmitted data content and the second is a direct examination of the receiver s recovered clock signals relative to a reference. Both techniques provide benefits in different ways and allow the communication engineer evaluating test results increased confidence and understanding of receiver performance. Direct examination of data contents is performed by two different data techniques, power correlation or a modified Massey correlation, and can be applied to soft decision data widths 1 to 12 bits wide over a correlation depth ranging from 16 to 512 samples. The SDA detects receiver bit slips within a 4 bits window and can handle systems with up to four quadrants (QPSK, SQPSK, and BPSK systems). The SDA continuously monitors correlation results to characterize slips and quadrant change and is capable of performing analysis even when the receiver under test is subjected to conditions where its performance degrades to high error rates (30 percent or beyond). The design incorporates a number of features, such as watchdog triggers that permit the SDA system to recover from large receiver upsets automatically and continue accumulating performance analysis unaided by operator intervention. This accommodates tests that can last in the order of days in order to gain statistical confidence in results and is also useful for capturing snapshots of rare events.

Steele, Glen↗

A Heuristic Approach to Correlating ERAM Flight Data from Twenty Centers

Among its many other functions, the Federal Aviation Administration’s En Route Automation Modernization (ERAM) provides external systems with real-time air traffic data for flights in enroute airspace in the National Airspace System. It replaced the En Route Host computer and backup system used at 20 FAA Air Route Traffic Control Centers (Centers) nationwide. Among the new features of ERAM, its output data stream of flight plan and track data includes a unique identifier for a flight originating in any one of the 20 ERAM Centers. The unique identifier, called the Global Unique Flight Identifier (GUFI), is persistent across all the Centers that track the flight. However, certain factors make it difficult to correlate data using the GUFI. First, the value of the GUFI is only unique within a time window of seven days. Second, the GUFI is attached only to flight-plan related data messages. Finally, track positions reported by ERAM do not reference the GUFI. In order to correlate historical as well as real time flight-plan and position related ERAM data, an efficient, heuristic approach was developed, and a prototype was developed. The approach showed that the processing speed, through parallel processing, is sufficient to correlate ERAM data in real-time. As described in this paper, when there are multiple track positions reported from multiple Centers within a few seconds, each position is assigned with a weighted score to indicate the quality of the position relative to its last know position. The weighted score can be used to eliminate potentially duplicate track positions. The approach is database-agnostic, and can be implemented in a Big Data system such as an Apache Hadoop system, as well as in traditional database systems.

Correlating ERAM Flight Data↗

DRIPS: Dynamic Rebalancing of Pipelined Streaming Applications on CGRAs

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit data streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across application kernels. However, emerging streaming applications at the edge (scientific instruments, sensor networks, network processing) perform much more than digital signal processing and often are data and input dependent. This leads to extremely variable kernel execution times, severely impacting the throughput of the entire pipeline if resources are only statically allocated. Therefore, in this paper, we propose DRIPS — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We present a unified compiler framework to facilitate the mapping of a given streaming application onto the DRIPS CGRA architecture. The experimental results show that DRIPS achieves an average throughput improvement of 1.46$\times$ across a set of representative applications over a statically partitioned solution. The additional area overhead to enable dynamic rebalancing consumes 16.34% of the entire area for a 5x5 CGRA prototype.

Tan, Cheng↗

Google Health Trends performance reflecting dengue incidence for the Brazilian states

Abstract Background Dengue fever is a mosquito-borne infection transmitted by Aedes aegypti and mainly found in tropical and subtropical regions worldwide. Since its re-introduction in 1986, Brazil has become a hotspot for dengue and has experienced yearly epidemics. As a notifiable infectious disease, Brazil uses a passive epidemiological surveillance system to collect and report cases; however, dengue burden is underestimated. Thus, Internet data streams may complement surveillance activities by providing real-time information in the face of reporting lags. Methods We analyzed 19 terms related to dengue using Google Health Trends (GHT), a free-Internet data-source, and compared it with weekly dengue incidence between 2011 to 2016. We correlated GHT data with dengue incidence at the national and state-level for Brazil while using the adjusted R squared statistic as primary outcome measure (0/1). We used survey data on Internet access and variables from the official census of 2010 to identify where GHT could be useful in tracking dengue dynamics. Finally, we used a standardized volatility index on dengue incidence and developed models with different variables with the same objective. Results From the 19 terms explored with GHT, only seven were able to consistently track dengue. From the 27 states, only 12 reported an adjusted R squared higher than 0.8; these states were distributed mainly in the Northeast, Southeast, and South of Brazil. The usefulness of GHT was explained by the logarithm of the number of Internet users in the last 3 months, the total population per state, and the standardized volatility index. Conclusions The potential contribution of GHT in complementing traditional established surveillance strategies should be analyzed in the context of geographical resolutions smaller than countries. For Brazil, GHT implementation should be analyzed in a case-by-case basis. State variables including total population, Internet usage in the last 3 months, and the standardized volatility index could serve as indicators determining when GHT could complement dengue state level surveillance in other countries.

59 BASIC BIOLOGICAL SCIENCES↗

Simple and Scalable Streaming: The GRETA Data Pipeline

The Gamma Ray Energy Tracking Array (GRETA) is a state of the art gamma-ray spectrometer being built at Lawrence Berkeley National Laboratory to be first sited at the Facility for Rare Isotope Beams (FRIB) at Michigan State University. A key design requirement for the spectrometer is to perform gamma-ray tracking in near real time. To meet this requirement we have used an inline, streaming approach to signal processing in the GRETA data acquisition system, using a GPU-equipped computing cluster. The data stream will reach 480 thousand events per second at an aggregate data rate of 4 gigabytes per second at full design capacity. We have been able to simplify the architecture of the streaming system greatly by interfacing the FPGA-based detector electronics with the computing cluster using standard network technology. A set of highperformance software components to implement queuing, flow control, event processing and event building have been developed, all in a streaming environment which matches detector performance. Prototypes of all high-performance components have been completed and meet design specifications.

Cromaz, Mario↗

Real-time nuclear activation detectors for measuring neutron angular distributions at the National Ignition Facility (invited)

The Real Time Nuclear Activation Detector (RTNAD) array at NIF measures the distribution of 14 MeV neutrons emitted by deuterium-tritium (DT) fueled inertial confinement fusion implosions. The uniformity of the neutron distribution is an important indication of implosion symmetry and DT shell integrity. The array consists of 48 LaBr 3 (Ce) crystal gamma-ray spectrometers mounted outside the NIF target chamber, which continuously monitor the slow decay of the 909 keV gamma-ray line from activated 89 Zr located in Zr cups surrounding each crystal. The measured decay rate dramatically increases during a DT implosion in proportion to the number of 14 MeV neutrons striking each Zr cup. The neutrons produce activated 89 Zr through an (n, 2n) reaction on 90 Zr, which is insensitive to low energy neutrons. The neutron flux along the detector line-of-sight at shot time is determined by extrapolating the fitted 909 keV decay curve back to shot time. Automatic analysis algorithms were developed to handle the non-stop data stream. The large number of detectors and the high statistical accuracy of the array enable the spherical harmonic modes of the neutron angular distribution to be measured up to L ≤ 4 to provide a better understanding of implosion dynamics. In addition, these data combined with measurements of the down-scattered neutrons can be used to derive fuel areal density distributions. This paper will describe the RTNAD hardware and analysis procedures.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Multimodal sensor fusion framework for residential building occupancy detection

For several years now, smart building energy systems have been a research area of intensive activity. In light of the increasing need for sustainable buildings and energy systems, this trend motivates an increasing need for a solution to reduce carbon dioxide emissions and improve energy efficiency. This work proposes a high-performing and transferable occupancy detection framework that combines sensor data from different data modalities, including time series environmental data (temperature, humidity, and illuminance), image data, and acoustic energy data using ensemble method. To draw out the best prediction performance in each modality, the proposed framework was developed, including various models that were designed to learn the occupancy patterns reflected in the physical data streams. To tackle the time series environmental data, we designed two variants of an occupancy detection spatiotemporal pattern network (Occ-STPN) that performs both feature level and decision level fusion, respectively. We also propose a new metric; the fading memory mean square error (FMMSE), that provides a fair evaluation and penalization of delayed occupancy predictions. Multiple open-sourced datasets, including the Electricity Consumption and Occupancy and the University of California, Irvine's (UCI) building occupancy detection dataset, along with our own real data collected from six different houses, were used to validate the algorithms' performance. The experimental results presented herein break down the performance for each sensing modality, and a detailed analysis of the performance is also discussed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nuclear data uncertainty propagation and modeling uncertainty impact evaluation in neutronics core simulation

Uncertainty analysis is a critical requirement in reactor simulation as it is used to quantify the reliability of best-estimate calculation. A comprehensive uncertainty analysis should characterize all sources of uncertainties in a computationally-feasible and scientifically-defendable manner. Here we employ a well-established reduced order modeling (ROM) based uncertainty quantification methodology to propagate uncertainties throughout neutronic calculations. ROM relies on recent advances in randomized data mining techniques applied to large data streams. In our proposed implementation, the nuclear data uncertainties are first propagated from multi-group level through lattice physics calculation to generate few-group parameter uncertainties, described using a vector of mean values and a covariance matrix. Employing an ROM-based compression of the covariance matrix, the few-group uncertainties are then propagated through downstream core simulation in a computationally efficient manner. This straightforward approach, albeit efficient as compared to brute force forward and/or adjoint-based methods, often employs a number of assumptions that have been unquestioned in the literature of neutronic uncertainty analysis. This manuscript argues that these assumptions could introduce another source of uncertainty referred to as modeling uncertainties, whose magnitude needs to be quantified in tandem with nuclear data uncertainties. Thus, our primary goal is to explore the interactions between these two uncertainty sources in order to assess whether modeling uncertainties have an impact on parameter uncertainties. To explore this endeavor, the impact of a number of modeling assumptions on core attributes uncertainties is quantified. The study employs a CANDU reactor model, with Serpent and NEWT as lattice physics solvers and NESTLE-C as core simulator. The modeling assumptions investigated include those related with the uncertainty propagation method employed, e.g., deterministic vs. stochastic, the few-group energy structure employed to represent the cross-sections, the resonance treatment in lattice physics calculation, the reference values for the cross-section, and the number of samples employed to render ROM compression. Results indicate that some of the modeling assumptions could have a non-negligible impact on the core responses propagated uncertainties, highlighting the need for a more comprehensive approach to combine parameter and modeling uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Wattile: Probabilistic Deep Learning-based Forecasting of Building Energy Consumption [SWR-20-94]

Accurate energy forecasting is becoming critical due to many reasons: i ) optimal distributed energy resources operations and dispatch, ii) fault detection and diagnostics, and iii) meeting operational energy efficiency targets. Wattile uses deep learning (DL) for the building's short-term load forecasting application. Two specific types of neural networks called, Long Short Term Memory (LSTM) and Sequence-to-Sequence (S2S) models are used to make predictions. Forecasting models are trained using online historical weather and occupancy indicator data streams from the Intelligent Campus Program's data acquisition systems at the National Renewable Energy Laboratory (NREL) for main meters and sub-meters of multiple building types. These models use probabilistic methods to provide quantile-based forecasts in addition to nominal conditional median predictions of electricity consumption.

Frank, Stephen↗