Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Network data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Persistent, “Mysterious” Seismoacoustic Signals Reported in Oklahoma State during 2019

Here, we report on the source of seismoacoustic pulses that were observed across the state of Oklahoma (OK) during summer of 2019, and the subject of national media coverage and speculation. Seismic network data collected across four U.S. states and interviews with witnesses to the pulse’s effect on residential structures demonstrate that they were triggered by routine ammunition disposal operations conducted by McAlester Army Ammunition Plant (McAAP). During these operations, conventional explosives destroy obsolete munitions stored in pits through a controlled sequence of electronically timed shots that occur over tens of minutes. Despite noise-abatement efforts that reduce coupling of acoustic energy with air, some lower frequency, subaudible (infrasonic) sound radiates from these shots as discrete pulses. We use nine months of blast log documents, seismic network records, analyst picks, and physical modeling to demonstrate that seismic stations as far as 640 km from McAAP sample these pulses, which record seasonal patterns in stratospheric and tropospheric winds, as well as the dynamic formation of waveguides and shadow zones. Digital short-term average to long-term average detectors that we augment with dynamic thresholds and time-binning operations identify these pulses with a fair probability, when compared with visual observations. Our analyses thereby provide estimates of observation rates for both partial and full sequences of these pulses, as well as single shots. We suggest that disposal operations can exploit existing, composite seismic networks to predict where residents are likely to witness blasting. Crucially, our data also show that dense seismic networks can record multiscale atmospheric processes in the absence of infrasound arrays.

58 GEOSCIENCES↗

Federated Machine Learning-Based Anomaly Detection System for Synchrophasor Network Using Heterogeneous Data Sets: Preprint

Synchrophasor technology is widely deployed in the energy management system to monitor the grid health at micro level and perform necessary corrective actions in real time; however, integrated phasor devices and data aggregators are exposed to several cybersecurity threats. This paper proposes a federated ML(FML)-based ADS to detect several data integrity attacks in the synchrophasor network. The proposed approach integrates the horizontal FML technique and consists of substation-based local models and a control center-based global model. The proposed methodology includes training local models using heterogeneous data sets that include network and grid information and updating the global model through multiple iterations by sharing model gradients. Finally, the trained global model is applied to identify cyberattacks, normal operation, and physical events. To validate the proof of concept, we used synthetic data sets generated by Mississippi State University and Oak Ridge National Laboratory for training and testing the classification models using the National Renewable Energy Laboratory's high performance computing resources. Our experimental results, computed through several performance measures, reveal that the proposed approach shows consistent performance during the binary, three-class, and multiclass classifications while ensuring privacy of synchrophasor data.

anomaly detection system↗

The 2021 update of the EPA’s adverse outcome pathway database

The EPA developed the Adverse Outcome Pathway Database (AOP-DB) to better characterize adverse outcomes of toxicological interest that are relevant to human health and the environment. Here we present the most recent version of the EPA Adverse Outcome Pathway Database (AOP-DB), version 2. AOP-DB v.2 introduces several substantial updates, which include automated data pulls from the AOP-Wiki 2.0, the integration of tissue-gene network data, and human AOP-gene data by population, semantic mapping and SPARQL endpoint creation, in addition to the presentation of the first publicly available AOP-DB web user interface. Potential users of the data may investigate specific molecular targets of an AOP, the relation of those gene/protein targets to other AOPs, cross-species, pathway, or disease-AOP relationships, or frequencies of AOP-related functional variants in particular populations, for example. Version updates described herein help inform new testable hypotheses about the etiology and mechanisms underlying adverse outcomes of environmental and toxicological concern.

59 BASIC BIOLOGICAL SCIENCES↗

OPFLearn.jl [SWR-21-109]

OPFLearn.jl is a Julia package for creating datasets for machine learning approaches to solving AC optimal power flow (AC OPF). It was developed to provide researchers with a standardized way to efficiently create AC OPF datasets that are representative of more of the AC OPF feasible load space compared to typical dataset creation methods. The OPFLearn dataset creation method uses a relaxed AC OPF formulation to reduce the volume of the unclassified input space throughout the dataset creation process. Over time this input space tightens around the relaxed AC OPF feasible region to increase the percentage of feasible load profiles found while uniformly sampling the input space. Load samples are processed using AC OPF formulations from PowerModels.jl. More information on the dataset creation method can be found in our publication, "OPF-Learn: An Open-Source Framework for Creating Representative AC Optimal Power Flow Datasets". To use OPFLearn.jl a PowerModels network data dictionary is required (can be loaded from Matpower ".m" files) to define the network the dataset is being created for.

Joswig-Jones, Trager↗

OPFLearn.jl v0.1.2 5/18/2023 [SWR-21-109]

OPFLearn.jl is a Julia package for creating datasets for machine learning approaches to solving AC optimal power flow (AC OPF). It was developed to provide researchers with a standardized way to efficiently create AC OPF datasets that are representative of more of the AC OPF feasible load space compared to typical dataset creation methods. The OPFLearn dataset creation method uses a relaxed AC OPF formulation to reduce the volume of the unclassified input space throughout the dataset creation process. Over time this input space tightens around the relaxed AC OPF feasible region to increase the percentage of feasible load profiles found while uniformly sampling the input space. Load samples are processed using AC OPF formulations from PowerModels.jl. More information on the dataset creation method can be found in our publication, "OPF-Learn: An Open-Source Framework for Creating Representative AC Optimal Power Flow Datasets". To use OPFLearn.jl a PowerModels network data dictionary is required (can be loaded from Matpower ".m" files) to define the network the dataset is being created for.

Joswig-Jones, Trager↗

Driver Identification Dataset

The ORNL Driver Identification Dataset was created to collect and analyze driving behavior data from 50 different drivers. Each driver operated a 2014 Kenworth T270 Class 6 truck around Fort Collins, Colorado while various data sources recorded their driving behavior and vehicle performance. The dataset includes CANbus (Controller Area Network) data, GPS data, inertial measurement data, and biometric data from a heart rate monitor. A cyberattack was executed during each drive, which caused multiple dashboard warning lights to illuminate and set the tachometer and speedometer to zero, regardless of actual speed. The attack was stopped either after one minute or if the driver pulled over. By downloading the dataset, you agree to the following: 1) I will not use or disclose the data for any purpose other than Research as that term is defined in 10 CFR 745.102. 2) I will not, under any circumstances, request or accept private or linking identifiers for the data used. 3) I will not attempt to determine the identity of the individuals associated with the data. 4) I will use appropriate safeguards to prevent the use or disclose of the data for any purpose other than Research.

99 GENERAL AND MISCELLANEOUS↗

Automated Detection of Instability-Inducing Channel Geometry Transitions in Saint-Venant Simulation of Large-Scale River Networks

A new sweep-search algorithm (SSA) is developed and tested to identify the channel geometry transitions responsible for numerical convergence failure in a Saint-Venant equation (SVE) simulation of a large-scale open-channel network. Numerical instabilities are known to occur at “sharp” transitions in discrete geometry, but the identification of problem locations has been a matter of modeler’s art and a roadblock to implementing large-scale SVE simulations. The new method implements techniques from graph theory applied to a steady-state 1D shallow-water equation solver to recursively examine the numerical stability of each flowpath through the channel network. The SSA is validated with a short river reach and tested by the simulation of ten complete river systems of the Texas–Gulf Coast region by using the extreme hydrological conditions recorded during hurricane Harvey. The SSA successfully identified the problematic channel sections in all tested river systems. Subsequent modification of the problem sections allowed stable solution by an unsteady SVE numerical solver. The new SSA approach permits automated and consistent identification of problem channel geometry in large open-channel network data sets, which is necessary to effectively apply the fully dynamic Saint-Venant equations to large-scale river networks or for city-wide stormwater networks.

54 ENVIRONMENTAL SCIENCES↗

Elephants Sharing the Highway: Studying TCP Fairness in Large Transfers over High Throughput Links

Escalating bandwidth demand strains high-performance data networks, posing potential performance risks. TCP congestion control algorithms enhance reliability and optimize bandwidth usage. Network performance is influenced by factors such as AQM algorithms and router buffer size. In the context of constrained network resources, understanding how TCP flows share networks and the resulting performance impact is essential. This paper introduces insights into TCP fairness and performance involving a comparison of TCP CUBIC, Reno, Hamilton, and BBR versions 1 and 2 across real-world networks supporting high bandwidths of up to 25 Gbps. The research explores TCP behaviors with AQM algorithms like FIFO, FQ_CODEL, and RED, alongside diverse buffer sizes. Notably, findings reveal that manipulating buffers and queuing methods yields contrasting outcomes based on bandwidth. BBRv2 emerges as a superior fair algorithm, pivotal for swift transfers, particularly in scientific data scenarios. These results provide crucial guidance for future network design, ensuring equitable performance optimization.

Kiran, Mariam↗

Utah FORGE: GES Well 16A(78)-32 and Well 16B(78)-32 Stimulation Seismic Event Catalogs

This dataset contains seismic event catalogs from the hydraulic stimulation of wells 16A(78)-32 and 16B(78)-32 at the Utah FORGE site in April 2024. The data was collected by Geo Energy Suisse (GES) using a variety of seismic monitoring technologies, including 3-component (3C) geophones and distributed acoustic sensing (DAS) systems. These technologies were deployed across several locations, including wells 16A, 16B, and Delano-1, with sensor arrays at multiple depths to capture microseismic activity during the stimulations. The catalogs provide both real-time and manually checked seismic event locations, with detailed parameters such as trigger conditions, velocity models, and data acquisition settings. The dataset includes information on the stimulation stages, event rates, and hydraulic injection conditions for each well, with a report detailing the data acquisition configuration and seismic event location methodologies. Users will need to reference the included report for a complete understanding of the sensor network, data processing techniques, and accuracy considerations.

15 GEOTHERMAL ENERGY↗

Massive Trajectory Data Based on Patterns of Life

Individual human location trajectory and check-in data have been the driving force for human mobility research in recent years. However, existing human mobility datasets are very limited in size and representativeness. For example, one of the largest and most commonly used datasets of individual human location trajectories, GeoLife, captures fewer than two hundred individuals. To help fill this gap, this Data and Resources paper leverages an existing data generator based on fine-grained simulation of individual human patterns of life to produce large-scale trajectory, check-in, and social network data. In this simulation, individual human agents commute between their home and work locations, visit restaurants to eat, and visit recreational sites to meet friends. We provide large datasets of months of simulated trajectories for two example regions in the United States: San Francisco and New Orleans. In addition to making the datasets available, we also provide instructions on how the simulation can be used to re-generate data, thus allowing researchers to generate the data locally without downloading prohibitively large files.

Amiri, Hossein↗

Data reduction through optimized scalar quantization for more compact neural networks

Raw data generation for several existing and planned large physics experiments now exceeds TB/s rates, generating untenable data sets in very little time. Those data often demonstrate high dimensionality while containing limited information. Meanwhile, Machine Learning algorithms are now becoming an essential part of data processing and data analysis. Those algorithms can be used offline for post processing and post data analysis, or they can be used online for real time processing providing ultra low latency experiment monitoring. Both use cases would benefit from data throughput reduction while preserving relevant information: one by reducing the offline storage requirements by several orders of magnitude and the other by allowing ultra fast online inferencing with low complexity Machine Learning models. Moreover, reducing the data source throughput also reduces material cost, power and data management requirements. In this work we demonstrate optimized nonuniform scalar quantization for data source reduction. This data reduction allows lower dimensional representations while preserving the relevant information of the data, thus enabling high accuracy Tiny Machine Learning classifier models for online fast inferences. We demonstrate this approach with an initial proof of concept targeting the CookieBox, an array of electron spectrometers used for angular streaking, that was developed for LCLS-II as an online beam diagnostic tool. We used the Lloyd-Max algorithm with the CookieBox dataset to design an optimized nonuniform scalar quantizer. Optimized quantization lets us reduce input data volume by 69% with no significant impact on inference accuracy. When we tolerate a 2% loss on inference accuracy, we achieved 81% of input data reduction. Finally, the change from a 7-bit to a 3-bit input data quantization reduces our neural network size by 38%.

97 MATHEMATICS AND COMPUTING↗

Radar Wind Profiler at McKinleyville, CA

These data are collected as part of an observational database developed to support the floating offshore wind energy research under the ORACLE project funded by DOE Wind Energy Technologies Office (WETO). The radar wind profiler network data are collected by NOAA (https://psl.noaa.gov/data/obs/datadisplay), and only data within the state of California are part of the database.

17 WIND ENERGY↗

Radar Wind Profiler at Bodega Bay

These data are collected as part of an observational database developed to support the floating offshore wind energy research under the ORACLE project funded by DOE Wind Energy Technologies Office (WETO). The radar wind profiler network data are collected by NOAA (https://psl.noaa.gov/data/obs/datadisplay), and only data within the state of California are part of the database.

17 WIND ENERGY↗

Radar Wind Profiler at Point Sur

These data are collected as part of an observational database developed to support the floating offshore wind energy research under the ORACLE project funded by DOE Wind Energy Technologies Office (WETO). The radar wind profiler network data are collected by NOAA (https://psl.noaa.gov/data/obs/datadisplay), and only data within the state of California are part of the database.

17 WIND ENERGY↗

Radar Wind Profiler at Santa Barbara

These data are collected as part of an observational database developed to support the floating offshore wind energy research under the ORACLE project funded by DOE Wind Energy Technologies Office (WETO). The radar wind profiler network data are collected by NOAA (https://psl.noaa.gov/data/obs/datadisplay), and only data within the state of California are part of the database.

17 WIND ENERGY↗

Radar Wind Profiler at San Nicolas Island

These data are collected as part of an observational database developed to support the floating offshore wind energy research under the ORACLE project funded by DOE Wind Energy Technologies Office (WETO). The radar wind profiler network data are collected by NOAA (https://psl.noaa.gov/data/obs/datadisplay), and only data within the state of California are part of the database.

17 WIND ENERGY↗

Dynamic Boundary Microgrids Under Privatization Considerations

Microgrids have physical, electrical, and logical (data, network, and ownership) boundaries. To power unserved customer loads during an outage, microgrids can extend the traditional operational boundaries. This can become complex when considering microgrid-to-microgrid (M2M) interactions where sensitive information such as competitive microgrid operational data is not shared. This work proposes an optimization method coordinated between microgrid controllers and distribution management systems that limits data sharing. The method involves a competitive bidding strategy that maximizes unserved load coverage while minimizing resource utilization and sensitive operational data sharing among entities. The work is validated on a two-microgrid system with photovoltaic and energy storage systems and curves of load derived from real world residential buildings datasets. Results show that the proposed method, when applied for three distinct use cases of energy storage sufficiency to cover the predefined boundary and/or the expanded boundary, can successfully select and bid the available load coverage.

Starke, Michael [ORNL] (ORCID:0000000221211195)↗

Machine learning-based analysis of COVID-19 pandemic impact on US research networks

Here in this study we explore how fallout from the changing public health policy around COVID-19 has changed how researchers access and process their science experiments. Using a combination of techniques from statistical analysis and machine learning, we conduct a retrospective analysis of historical network data for a period around the stay-at-home orders that took place in March 2020. Our analysis takes data from the entire ESnet infrastructure to explore DOE high-performance computing (HPC) resources at OLCF, ALCF, and NERSC, as well as User sites such as PNNL and JLAB. We look at detecting and quantifying changes in site activity using a combination of t-Distributed Stochastic Neighbor Embedding (t-SNE) and decision tree analysis. Our findings bring insights into the working patterns and impact on data volume movements, particularly during late-night hours and weekends.

97 MATHEMATICS AND COMPUTING↗