Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Network data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Coordinating an operational data distribution network for CMIP6 data

The distribution of data contributed to the Coupled Model Intercomparison Project Phase 6 (CMIP6) is via the Earth System Grid Federation (ESGF). The ESGF is a network of internationally distributed sites that together work as a federated data archive. Data records from climate modelling institutes are published to the ESGF and then shared around the world. It is anticipated that CMIP6 will produce approximately 20 PB of data to be published and distributed via the ESGF. In addition to this large volume of data a number of value-added CMIP6 services are required to interact with the ESGF; for example the citation and errata services both interact with the ESGF but are not a core part of its infrastructure. With a number of interacting services and a large volume of data anticipated for CMIP6, the CMIP Data Node Operations Team (CDNOT) was formed. The CDNOT coordinated and implemented a series of CMIP6 preparation data challenges to test all the interacting components in the ESGF CMIP6 software ecosystem. This ensured that when CMIP6 data were released they could be reliably distributed.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Machine-learning-aided cognitive reconfiguration for flexible-bandwidth HPC and data center networks [Invited]

This paper proposes a machine-learning (ML)-aided cognitive approach for effective bandwidth reconfiguration in optically interconnected datacenter/high-performance computing (HPC) systems. The proposed approach relies on a Hyper-X-like architecture augmented with flexible-bandwidth photonic interconnections at large scales using a hierarchical intra/inter-POD photonic switching layout. We first formulate the problem of the connectivity graph and routing scheme optimization as a mixed-integer linear programming model. A two-phase heuristic algorithm and a joint optimization approach are devised to solve the problem with low time complexity. Then, we propose an ML-based end-to-end performance estimator design to assist the network control plane with intelligent decision making for bandwidth reconfiguration. Numerical simulations using traffic distribution profiles extracted from HPC applications traces as well as random traffic matrices verify the accuracy performance of the ML design estimator ( < <#comment/> 9 % <#comment/> error) and demonstrate up to 5 × <#comment/> throughput gain from the proposed approach compared with the baseline Hyper-X network using fixed all-to-all intra/inter-portable data center interconnects.

Chen, Xiaoliang (ORCID:0000000278056237)↗

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm↗

An Overview of the Usefulness of Machine Learning Techniques on Network Packet Data

Understanding the health and behavior of a computer network allows for better network efficiency and security. We present an overview of various machine learning techniques for classifying network packet data via packet metadata. While some classical machine learning approaches achieve reasonable results, the most accurate classification can be achieved with deep learning. On the four data sets studied herein, a basic deep learning model achieved at or near 100\% classification accuracy. We also propose a method for determining variable importance as a means for potential transfer learning applications to classifying yet unseen network packet data.

97 MATHEMATICS AND COMPUTING↗

AmeriFlux BASE data pipeline to support network growth and data sharing

Abstract AmeriFlux is a network of research sites that measure carbon, water, and energy fluxes between ecosystems and the atmosphere using the eddy covariance technique to study a variety of Earth science questions. AmeriFlux’s diversity of ecosystems, instruments, and data-processing routines create challenges for data standardization, quality assurance, and sharing across the network. To address these challenges, the AmeriFlux Management Project (AMP) designed and implemented the BASE data-processing pipeline. The pipeline begins with data uploaded by the site teams, followed by the AMP team’s quality assurance and quality control (QA/QC), ingestion of site metadata, and publication of the BASE data product. The semi-automated pipeline enables us to keep pace with the rapid growth of the network. As of 2022, the AmeriFlux BASE data product contains 3,130 site years of data from 444 sites, with standardized units and variable names of more than 60 common variables, representing the largest long-term data repository for flux-met data in the world. The standardized, quality-ensured data product facilitates multisite comparisons, model evaluations, and data syntheses.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Anonymization of Network Traces Data through Condensation-based Differential Privacy

Network traces are considered a primary source of information to researchers, who use them to investigate research problems such as identifying user behavior, analyzing network hierarchy, maintaining network security, classifying packet flows, and much more. However, most organizations are reluctant to share their data with a third party or the public due to privacy concerns. Therefore, data anonymization prior to sharing becomes a convenient solution to both organizations and researchers. Although several anonymization algorithms are available, few of them allow sufficient privacy (organization need), acceptable data utility (researcher need), and efficient data analysis at the same time. This article introduces a condensation-based differential privacy anonymization approach that achieves an improved tradeoff between privacy and utility compared to existing techniques and produces anonymized network trace data that can be shared publicly without lowering its utility value. Our solution also does not incur extra computation overhead for the data analyzer. A prototype system has been implemented, and experiments have shown that the proposed approach preserves privacy and allows data analysis without revealing the original data even when injection attacks are launched against it. When anonymized datasets are given as input to graph-based intrusion detection techniques, they yield almost identical intrusion detection rates as the original datasets with only a negligible impact.

97 MATHEMATICS AND COMPUTING↗

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

Toward higher-radix switches with co-packaged optics for improved network locality in data center and HPC networks [Invited]

In this work, we study the network locality improvements that can be achieved by using co-packaged optics in data center and high-performance computing (HPC) networks. The increased escape bandwidth offered by co-packaged optics can enable switches with speeds of 51.2 Tb/s and beyond. From a network architecture perspective, the key advantages of introducing co-packaged optics at the switch points include the implementation of large-scale topologies of >12,000 end points with 4× higher bisection bandwidth and the reduction of the required number of switches by >40% compared with state-of-the-art approaches. From a network operation perspective, improved network locality and faster operation can be achieved since the higher-radix switches can mitigate the impact of network contention. Placing applications under fewer leaf switches reduces the number of packets that cross the spine switches in a leaf-spine topology. The proposed scheme is evaluated via discrete-event simulations: we initially evaluate the network locality properties of the system by using virtual-machine traces from a production data center, and we subsequently quantify the performance improvements by simulating an all-to-all pattern for a variety of message sizes over a number of nodes. The results suggest that co-packaged optics form a promising solution for keeping up with bandwidth scaling in future networks. The virtual-machine analysis shows that large-scale applications can be placed under up to 50% fewer first-level switches, while the network analysis shows speedups of up to 7.1, which translates to execution time reductions of up to 26% and 42.7% for applications with communication ratios of 0.3 and 0.5, respectively.

99 GENERAL AND MISCELLANEOUS↗

Emulation automation and model checking

A method of automating emulations is provided. The method comprising collecting publicly available network data over a predefined time interval, wherein the collected network data might comprise structured and unstructured data. Any unstructured data is converted into structured data. The original and converted structured data is stored in a database and compared to known network vulnerabilities. An emulated network is created according to the collected network data and the comparison of the structured data with known vulnerabilities. Virtual machines are created to run on the emulated network. Director programs and guest actor programs are run on the virtual machines, wherein the actor programs imitate real user behavior on the emulated network. The director programs deliver task commands to the guest actor programs to imitate real user behavior. The imitated behavior is presented to a user via an interface.

Urias, Vincent↗

Scaling of Floods With Geomorphologic Characteristics and Precipitation Variability Across the Conterminous United States

Abstract Accurate flood risk assessment requires a comprehensive understanding of flood sensitivity to regional drivers and climate factors. This paper presents the scaling of floods (duration, peak, volume) with geomorphologic characteristics of the basin (i.e., drainage area, slope, elevation) and precipitation patterns (rainfall accumulation, variability). Long‐term daily streamflow observations over the 20th and early 21st centuries from Hydro‐Climatic Data Network streamgages across the conterminous United States are used to create a flood event database based on their flood stage information. Antecedent daily rainfall accumulation and variability corresponding to these floods are computed using Global Historical Climatology Network daily data set. Two Bayesian scaling models are developed, and the spatial organization of scaling exponents is investigated. The baseline model quantifies the scaling of floods to geomorphologic characteristics. The dynamic model quantifies the scaling of floods to antecedent precipitation distribution which is further conditioned on geomorphologic characteristics. Results show that small and low‐elevation basins have a stronger response to antecedent rainfall distribution in amplifying flood peaks, while high‐elevation steeper basins have a lower response for flood duration and volume. The dynamic models demonstrate that there are significant variations in the flood scaling rates, with the largest rates up to 40% and 4.5% for flood duration, 64% and 44% for peak, and 98% and 40% for volume found across the Northeast, Coastal Southeast, and Northwest with intensifying rainfall accumulation and variability, respectively. This study advances flood predictions by better informing the flood attributes in the context of dynamical land‐atmosphere perturbations.

54 ENVIRONMENTAL SCIENCES↗

AUTOPERF

AutoPerf is a collection of modules for low-overhead performance monitoring on HPC systems, the modules collect monitoring data such as MPI data, network counter data etc. Autoperf modules interoperate with the Open source Darshan software (xgitlab.cels.anl.gov/darshan/) as submodules and leverages the Darshan's log recording and analysis framework.

CHUNDURI, DEVISUDHEER KUMAR↗

Machine Learning Based Network Parameter Estimation Using AMI Data

The expansion of distribution power system and the growing penetration of distributed energy resources present new challenges for situational awareness. Calibrating the extended system model with sensor measurements and maintaining the usability is critical for utilities. This paper presents a distribution network parameter estimation (DNPE) approach using machine learning (ML) and metering data that improve the quality of extended distribution power system modeling. The reliability model can improve the ability of endpoint data to be translated into network-level situational awareness in real time and help distribution system operators (DSOs) solve branch flow and voltage problems. In addition, a data analytic and automate processing scheme is proposed to improve the sensor data quality and prevent misleading information. The effectiveness of the proposed method is verified with actual advanced metering infrastructure (AMI) data on a real utility feeder model, while considering the higher penetration of photovoltaic power generation. The test of DNPE and study results are demonstrated in this paper.

Parameter estimation, machine learning, power dist↗

Utah FORGE: 2024 Discrete Fracture Network Model Data

The Utah FORGE 2024 Discrete Fracture Network (DFN) Model dataset provides a set of files representing discrete fracture network modeling for the FORGE site near Milford, Utah. The dataset includes four distinct DFN model file sets, each corresponding to different time frames and modeling approaches in 2024. These models characterize both natural and induced fractures in the geothermal reservoir, which consists of crystalline granitic and metamorphic rock approximately 8,000 feet below the ground surface. The dataset includes a reference DFN model from February 2024 that incorporates planar fractures and well trajectories, as well as upscaled permeability, porosity, compressibility, and storage values on specified grids. Additionally, there are models based on new microseismic (MEQ) data from May and July 2024, including fracture planes fitted to the latest MEQ catalog datasets, tensile fractures from hydraulic stimulation, and an alternative connected DFN for modeling purposes. Coordinate data is provided in both global and local frames, with detailed instructions on the transformations used to align with principal stress orientations. The dataset also includes notes and calculation files for estimating fracture sizes and differences between various fracture sets. There are subfolders for Global Coordinates and Local Coordinates. To move from the global to the local coordinate frame, fractures and wells were a) rotated 20 degrees counterclockwise looking down about the global point (335376.400482041, 4263189.99998761, 250.093546450195) to better align with the principal stresses; and b) translated by (-335408.68, -4263010.9, 1150). Upscaled permeability values using the _XYZ suffix show directions with respect to the global XYZ coordinate frame, while those using the _IJK suffix are aligned with local coordinate frame.

15 GEOTHERMAL ENERGY↗

Computing Bottleneck Structures at Scale for High-Precision Network Performance Analysis

The Theory of Bottleneck Structures is a recently-developed framework for studying the performance of data networks. It describes how local perturbations in one part of the network propagate and interact with others. This framework is a powerful analytical tool that allows network operators to make accurate predictions about network behavior and thereby optimize performance. Previous work implemented a software package for bottleneck structure analysis, but applied it only to toy examples. In this work, we introduce the first software package capable of scaling bottleneck structure analysis to production-size networks. Here, we benchmark our system using logs from ESnet, the Department of Energy's high-performance data network that connects research institutions in the U.S. Using the previously published tool as a baseline, we demonstrate that our system achieves vastly improved performance, constructing the bottleneck structure graphs in 0.21 s and calculating link derivatives in 0.09 s on average. We also study the asymptotic complexity of our core algorithms, demonstrating good scaling properties and strong agreement with theoretical bounds. These results indicate that our new software package can maintain its fast performance when applied to even larger networks. They also show that our software is efficient enough to analyze rapidly changing networks in real time. Overall, we demonstrate the feasibility of applying bottleneck structure analysis to solve practical problems in large, real-world data networks.

benchmark↗

GNET2: an R package for constructing gene regulatory networks from transcriptomic data

Abstract Motivation The Gene Network Estimation Tool (GNET) is designed to build gene regulatory networks (GRNs) from transcriptomic gene expression data with a probabilistic graphical model. The data preprocessing, model construction and visualization modules of the original GNET software were developed on different programming platforms, which were inconvenient for users to deploy and use. Results Here, we present GNET2, an improved implementation of GNET as an integrated R package. GNET2 provides more flexibility for parameter initialization and regulatory module construction based on the core iterative modeling process of the original algorithm. The data exchange interface of GNET2 is handled within an R session automatically. Given the growing demand for regulatory network reconstruction from transcriptomic data, GNET2 offers a convenient option for GRN inference on large datasets. Availability and implementation The source code of GNET2 is available at https://github.com/jianlin-cheng/GNET2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Toward lower-diameter large-scale HPC and data center networks with co-packaged optics

We investigate the advantages of using co-packaged optics for building low-diameter, large-scale high-performance computing (HPC) and data center networks. The increased escape bandwidth offered by co-packaged optics can enable high-radix switch implementations of more than 150 switch ports, which can be combined with data rates of up to 400 Gb/s per port. From the network architecture perspective, the key benefits of using co-packaged optics in future fat-tree networks include (a) the ability to implement large-scale topologies of > <#comment/> 11 , 000 end points by eliminating the need for a third switching layer and (b) the ability to provide up to 4 × <#comment/> higher bisection bandwidth compared to existing solutions, reducing at the same time the number of required switch application-specific integrated circuits by > <#comment/> 80 % <#comment/> . From the network operation perspective, both reduced energy consumption and lower packet delays can be achieved since fewer hops are required; i.e., packets need to traverse fewer serializer/deserializer lanes and fewer switch buffers, which reduces the probability of contending with other packets and improves the tolerance of network congestion. The performance of the proposed architecture is evaluated via discrete-event simulations for a wide range of representative HPC synthetic-traffic cases that include both hotspot and non-hotspot scenarios. The simulation results suggest that co-packaged optics form a promising solution to keep up with bandwidth scaling in future networks, while the reduced number of switching layers can lead to significant mean packet delay improvements that start from 30% and reach up to 74% for high-load conditions.

Maniotis, Pavlos (ORCID:0000000244905253)↗

A data-driven network optimisation approach to coordinated control of distributed photovoltaic systems and smart buildings in distribution systems

The increasing integration of distributed energy resources, including demand-side resources and distributed photovoltaics (PVs), into distribution systems has resulted in more complicated power system operation. A data-driven network optimisation approach is proposed to coordinate the control of distributed PVs and smart buildings in distribution networks considering the uncertainties of solar power, outdoor temperature and heat gain associated with building thermal dynamics. These uncertain parameters have a significant impact on the operation and control of distributed PVs and smart buildings, bringing challenges to the distribution system operation. In the proposed data-driven distributionally robust optimisation (DRO) approach, the Wasserstein ball is used to construct an ambiguity set for the uncertain parameters, which does not require the probability distributions to be known. Furthermore, a conditional value-at-risk is incorporated into the Wasserstein-based DRO model and converted into a computationally tractable mixed-integer convex optimisation problem. Benchmarked with robust optimisation and chance-constrained programming, the proposed data-driven model can give a less conservative robust solution.

24 POWER TRANSMISSION AND DISTRIBUTION↗

SoDaH: the SOils DAta Harmonization database, an open-source synthesis of soil data from research networks, version 1.0

Data collected from research networks present opportunities to test theories and develop models about factors responsible for the long-term persistence and vulnerability of soil organic matter (SOM). Synthesizing datasets collected by different research networks presents opportunities to expand the ecological gradients and scientific breadth of information available for inquiry. Synthesizing these data is challenging, especially considering the legacy of soil data that have already been collected and an expansion of new network science initiatives. To facilitate this effort, here we present the SOils DAta Harmonization database (SoDaH; https://lter.github.io/som-website, last access: 22 December 2020), a flexible database designed to harmonize diverse SOM datasets from multiple research networks. SoDaH is built on several network science efforts in the United States, but the tools built for SoDaH aim to provide an open-access resource to facilitate synthesis of soil carbon data. Moreover, SoDaH allows for individual locations to contribute results from experimental manipulations, repeated measurements from long-term studies, and local- to regional-scale gradients across ecosystems or landscapes. Finally, we also provide data visualization and analysis tools that can be used to query and analyze the aggregated database. The SoDaH v1.0 dataset is archived and available at https://doi.org/10.6073/pasta/9733f6b6d2ffd12bf126dc36a763e0b4 (Wieder et al., 2020).

54 ENVIRONMENTAL SCIENCES↗