Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “networking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Prediction of Aerodynamic Coefficient using Genetic Algorithm Optimized Neural Network for Sparse Data

Wind tunnels use scale models to characterize aerodynamic coefficients, Wind tunnel testing can be slow and costly due to high personnel overhead and intensive power utilization. Although manual curve fitting can be done, it is highly efficient to use a neural network to define the complex relationship between variables. Numerical simulation of complex vehicles on the wide range of conditions required for flight simulation requires static and dynamic data. Static data at low Mach numbers and angles of attack may be obtained with simpler Euler codes. Static data of stalled vehicles where zones of flow separation are usually present at higher angles of attack require Navier-Stokes simulations which are costly due to the large processing time required to attain convergence. Preliminary dynamic data may be obtained with simpler methods based on correlations and vortex methods; however, accurate prediction of the dynamic coefficients requires complex and costly numerical simulations. A reliable and fast method of predicting complex aerodynamic coefficients for flight simulation I'S presented using a neural network. The training data for the neural network are derived from numerical simulations and wind-tunnel experiments. The aerodynamic coefficients are modeled as functions of the flow characteristics and the control surfaces of the vehicle. The basic coefficients of lift, drag and pitching moment are expressed as functions of angles of attack and Mach number. The modeled and training aerodynamic coefficients show good agreement. This method shows excellent potential for rapid development of aerodynamic models for flight simulation. Genetic Algorithms (GA) are used to optimize a previously built Artificial Neural Network (ANN) that reliably predicts aerodynamic coefficients. Results indicate that the GA provided an efficient method of optimizing the ANN model to predict aerodynamic coefficients. The reliability of the ANN using the GA includes prediction of aerodynamic coefficients to an accuracy of 110% . In our problem, we would like to get an optimized neural network architecture and minimum data set. This has been accomplished within 500 training cycles of a neural network. After removing training pairs (outliers), the GA has produced much better results. The neural network constructed is a feed forward neural network with a back propagation learning mechanism. The main goal has been to free the network design process from constraints of human biases, and to discover better forms of neural network architectures. The automation of the network architecture search by genetic algorithms seems to have been the best way to achieve this goal.

Rajkumar, T.↗

Scale-Free Networks and Commercial Air Carrier Transportation in the United States

Network science, or the art of describing system structure, may be useful for the analysis and control of large, complex systems. For example, networks exhibiting scale-free structure have been found to be particularly well suited to deal with environmental uncertainty and large demand growth. The National Airspace System may be, at least in part, a scalable network. In fact, the hub-and-spoke structure of the commercial segment of the NAS is an often-cited example of an existing scale-free network After reviewing the nature and attributes of scale-free networks, this assertion is put to the test: is commercial air carrier transportation in the United States well explained by this model? If so, are the positive attributes of these networks, e.g. those of efficiency, flexibility and robustness, fully realized, or could we effect substantial improvement? This paper first outlines attributes of various network types, then looks more closely at the common carrier air transportation network from perspectives of the traveler, the airlines, and Air Traffic Control (ATC). Network models are applied within each paradigm, including discussion of implied strengths and weaknesses of each model. Finally, known limitations of scalable networks are discussed. With an eye towards NAS operations, utilizing the strengths and avoiding the weaknesses of scale-free networks are addressed.

Conway, Sheila R.↗

Unified Lunar Control Network 2005 and Topographic Model

There are currently two generally accepted lunar control networks. These are the Unified Lunar Control Network (ULCN) and the Clementine Lunar Control Network (CLCN), both derived by M. Davies and T. Colvin at RAND. We address here our efforts to merge and improve these networks into a new ULCN. The ULCN was described in the last major publication about a lunar control network. The statistics on this and the other networks discussed here. Images for this network are from the Apollo, Mariner 10, and Galileo missions, and Earth-based photographs. The importance of this network is that its accuracy is relatively well quantified and published information on the network is available. The CLCN includes measurements on 43,871 Clementine 750-nm images - the largest planetary control network ever computed. This purpose of this network was to determine the geometry for the Clementine Basemap Mosiac (CBM). The geometry of that mosaic was used to produce the Clementine UVVIS digital image model and the Near-Infrared Global Multispectral Map of the Moon from Clementine. Through the extensive use of these products, they and the underlying CLCN in effect define the generally accepted current coordinate system for reporting and describing the location of lunar coordinates. However, no publication describes the CLCN itself.

Archinal, B. A.↗

Integrated Network Architecture for NASA's Orion Missions

NASA is planning a series of short and long duration human and robotic missions to explore the Moon and then Mars. The series of missions will begin with a new crew exploration vehicle (called Orion) that will initially provide crew exchange and cargo supply support to the International Space Station (ISS) and then become a human conveyance for travel to the Moon. The Orion vehicle will be mounted atop the Ares I launch vehicle for a series of pre-launch tests and then launched and inserted into low Earth orbit (LEO) for crew exchange missions to the ISS. The Orion and Ares I comprise the initial vehicles in the Constellation system of systems that later includes Ares V, Earth departure stage, lunar lander, and other lunar surface systems for the lunar exploration missions. These key systems will enable the lunar surface exploration missions to be initiated in 2018. The complexity of the Constellation system of systems and missions will require a communication and navigation infrastructure to provide low and high rate forward and return communication services, tracking services, and ground network services. The infrastructure must provide robust, reliable, safe, sustainable, and autonomous operations at minimum cost while maximizing the exploration capabilities and science return. The infrastructure will be based on a network of networks architecture that will integrate NASA legacy communication, modified elements, and navigation systems. New networks will be added to extend communication, navigation, and timing services for the Moon missions. Internet protocol (IP) and network management systems within the networks will enable interoperability throughout the Constellation system of systems. An integrated network architecture has developed based on the emerging Constellation requirements for Orion missions. The architecture, as presented in this paper, addresses the early Orion missions to the ISS with communication, navigation, and network services over five phases of a mission: pre-launch, launch from T0 to T+6.5 min, launch from T+6.5 min to 12 min, in LEO for rendezvous and docking with ISS, and return to Earth. The network of networks that supports the mission during each of these phases and the concepts of operations during those phases are developed as a high level operational concepts graphic called OV-1, an architecture diagram type described in the Department of Defense Architecture Framework (DoDAF). Additional operational views on organizational relationships (OV-4), operational activities (OV-5), and operational node connectivity (OV-2) are also discussed. The system interfaces view (SV-1) that provides the communication and navigation services to Orion is also included and described. The challenges of architecting integrated network architecture for the NASA Orion missions are highlighted.

Bhasin, Kul B.↗

Evolution of the Lunar Network

The National Aeronautics and Space Administration (NASA) is planning to upgrade its network Infrastructure to support missions for the 21st century. The first step is to increase the data rate provided to science missions to at least the 100 megabits per second (Mbps) range. This is under way, using Ka-band 26 Gigahertz (GHz), erecting an 18-meter antenna for the Lunar Reconnaissance Orbiter (LRO), and the planned upgrade of the Deep Space Network (DSN) 34-meter network to support the James Webb Space Telescope (JWST). The next step is the support of manned missions to the Moon and beyond. Establishing an outpost with several activities such as rovers, colonization, and observatories, is better achieved by using a network configuration rather than the current method of point-to-point communication. Another challenge associated with the Moon is communication coverage with the Earth. The Moon's South Pole, targeted for human habitat and exploration, is obscured from Earth view for half of the 28-day lunar cycle and requires the use of lunar relay satellites to provide coverage when there is no direct view of the Earth. The future NASA and Constellation network architecture is described in the Space Communications Architecture Working Group (SCAWG) Report. The Space Communications and Navigation (SCAN) Constellation Integration Project (SCIP) is responsible for coordinating Constellation requirements and has assigned the responsibility for implementing these requirements to the existing NASA communication providers: DSN, Space Network (SN), Ground Network (GN) and the NASA Integrated Services Network (NISN). The SCAWG Report provides a future architecture but does not provide implementation details. The architecture calls for a Netcentric system, using hundreds of 12-meter antennas, a ground antenna array, and a relay network around the Moon. The report did not use cost as a variable in determining the feasibility of this approach. As part of the SCIP Mission Concept Review and the second iteration of the Lunar Architecture Team (LAT), the focus is on cost, as well as communication coverage using operational scenarios. This approach maximizes use of existing assets and adds capability in small increments. This paper addresses architecture decisions such as the Radio Frequency (RF) signal and network (Netcentric) decisions that need to be made and the difficulty of implementing them into the existing Space Network and DSN. It discusses the evolution of the lunar system and describes its components: Tracking and Data Relay Satellite System (TDRSS), Earth-based ground stations, Lunar Relay, and surface systems.

Gal-Edd, Jonathan↗

Cascade Back-Propagation Learning in Neural Networks

The cascade back-propagation (CBP) algorithm is the basis of a conceptual design for accelerating learning in artificial neural networks. The neural networks would be implemented as analog very-large-scale integrated (VLSI) circuits, and circuits to implement the CBP algorithm would be fabricated on the same VLSI circuit chips with the neural networks. Heretofore, artificial neural networks have learned slowly because it has been necessary to train them via software, for lack of a good on-chip learning technique. The CBP algorithm is an on-chip technique that provides for continuous learning in real time. Artificial neural networks are trained by example: A network is presented with training inputs for which the correct outputs are known, and the algorithm strives to adjust the weights of synaptic connections in the network to make the actual outputs approach the correct outputs. The input data are generally divided into three parts. Two of the parts, called the "training" and "cross-validation" sets, respectively, must be such that the corresponding input/output pairs are known. During training, the cross-validation set enables verification of the status of the input-to-output transformation learned by the network to avoid over-learning. The third part of the data, termed the "test" set, consists of the inputs that are required to be transformed into outputs; this set may or may not include the training set and/or the cross-validation set. Proposed neural-network circuitry for on-chip learning would be divided into two distinct networks; one for training and one for validation. Both networks would share the same synaptic weights.

Duong, Tuan A.↗

Applying a Space-Based Security Recovery Scheme for Critical Homeland Security Cyberinfrastructure Utilizing the NASA Tracking and Data Relay (TDRS) Based Space Network

Protection of the national infrastructure is a high priority for cybersecurity of the homeland. Critical infrastructure such as the national power grid, commercial financial networks, and communications networks have been successfully invaded and re-invaded from foreign and domestic attackers. The ability to re-establish authentication and confidentiality of the network participants via secure channels that have not been compromised would be an important countermeasure to compromise of our critical network infrastructure. This paper describes a concept of operations by which the NASA Tracking and Data Relay (TDRS) constellation of spacecraft in conjunction with the White Sands Complex (WSC) Ground Station host a security recovery system for re-establishing secure network communications in the event of a national or regional cyberattack. Users would perform security and network restoral functions via a Broadcast Satellite Service (BSS) from the TDRS constellation. The BSS enrollment only requires that each network location have a receive antenna and satellite receiver. This would be no more complex than setting up a DIRECTTV-like receiver at each network location with separate network connectivity. A GEO BSS would allow a mass re-enrollment of network nodes (up to nationwide) simultaneously depending upon downlink characteristics. This paper details the spectrum requirements, link budget, notional assets and communications requirements for the scheme. It describes the architecture of such a system and the manner in which it leverages off of the existing secure infrastructure which is already in place and managed by the NASAGSFC Space Network Project.

Cybersecurity↗

Security-Enhanced Autonomous Network Management

Ensuring reliable communication in next-generation space networks requires a novel network management system to support greater levels of autonomy and greater awareness of the environment and assets. Intelligent Automation, Inc., has developed a security-enhanced autonomous network management (SEANM) approach for space networks through cross-layer negotiation and network monitoring, analysis, and adaptation. The underlying technology is bundle-based delay/disruption-tolerant networking (DTN). The SEANM scheme allows a system to adaptively reconfigure its network elements based on awareness of network conditions, policies, and mission requirements. Although SEANM is generically applicable to any radio network, for validation purposes it has been prototyped and evaluated on two specific networks: a commercial off-the-shelf hardware test-bed using Institute of Electrical Engineers (IEEE) 802.11 Wi-Fi devices and a military hardware test-bed using AN/PRC-154 Rifleman Radio platforms. Testing has demonstrated that SEANM provides autonomous network management resulting in reliable communications in delay/disruptive-prone environments.

Zeng, Hui↗

Ensuring Flexibility and Security in SDN-Based Spacecraft Communication Networks Through Risk Assessment

Software-defined networking (SDN) has enabled elastic networking and resource distribution in cloud computing. The centralization and separation of the Control Plane also offers a high degree of network configurability and management, which can be used to mitigate and manage threats to the network. Space communication networks have historically been restricted and circuit switching in these networks has been a manual process. This study evaluates the potential role of SDN in space communication networks from a networking security standpoint. The evaluation covers the networking security needs of spacecraft missions and their associated assets. The results from the evaluation lead to a risk assessment that identifies vulnerabilities in an SDN-based communications architecture. Security challenges introduced into the network from integrating SDN are also considered. A risk register summarizes the severity of the attack outcomes, as well as occurrence likelihood. The study identifies Denial-of-Service (DoS) attacks as a new threat (presently unmitigated by existing security controls) that would be prevalent in an SDN-based space communication environment. A Mininet-based emulation testbed is built to demonstrate the susceptibility of spacecraft flight software to a flooding DoS attack when on an interconnected SDN-managed network. This type of attack would be highly consequential to mission assets, and therefore SDN-based space communications would need to be resilient to such attacks. Future work will need to be performed to fully characterize DoS attack methods that can apply to the space communication scenario, as well as to devise a comprehensive DoS-resilient solution.

Baker, Dylan Z.↗

Distributed Quantum Processing via Integrated Regional Quantum Networks Using Free-Space Optical Links

The landscape for quantum computing is rapidly evolving with the development of quantum computers and the emergence of future science enabling applications being pursued in many areas, such as distributed quantum computing, hybrid computing scenarios that require cooperatively networked quantum and classical computing assets and more. Quantum networking is also advancing with multiple regional quantum networks and testbeds maturing across the nation. These quantum networking testbeds typically include multiple access points with a heterogeneous mixture of capabilities interconnected by fiber spanning anywhere from tens to hundreds of kilometers; many such networks incorporate a classical networking framework. Alongside these connectivity advancements in the infrastructure, there is quantum sensing device and component technology maturation that is serving to expand the body of quantum networking scenarios that are presently possible and redefining the quantum- enabled future. Bridging the present collective progress in quantum computing and quantum networking will enable the next era of purposeful advancements for national priority use-cases for the benefit of science. One such use-case is enabling future quantum computing applications via integrated regional quantum networking infrastructure in neighboring locations like New York and Maryland, both states are home to robust classical and quantum networking capabilities. To enable near-term distributed quantum computing applications via the practical integration of regional quantum networking infrastructure, free-space optical links via satellite communications will be essential to avoid the exponential loss commonly observed with fiber implementations. There is both an urgent need and a practical justification for pursuing such space-based integration scope now.

Harry Shaw↗

Intelligent network slicing and policy-based routing engine

One or more aspects of the present disclosure are directed to network optimization solutions provided as software agents (applications) executed on network nodes in a heterogenous multi-vendor environment to provide cross-layer network optimization and ensure availability of network resources to meet associated Quality of Experience (QoE) and Quality of Service (QoS). In one aspect, a network slicing engine is configured to receive at least one request from at least one network endpoint for access to the heterogeneous multi-vendor network for data transmission; receive information on state of operation of a plurality of communication links between the plurality of nodes; determine a set of data transmission routes for the request; assign a network slice for serving the request; determine, from the set of data transmission routes, an end-to-end route for the network slice; and send network traffic associated with the request using the network slice and over the end-to-end route.

Mody, Apurva N.↗

On the evaluation and selection of network-level traffic control policies: Perimeter control, TUC, and their combination

Perimeter control (PC) of urban traffic networks can be effective in increasing network-wide efficiency. PC operates on the border of a protected region of a traffic network. Most studies thus far considered fixed-time plans for the inner part of these regions. A few studies have shown that combining PC with locally actuated or decentralized traffic control systems may have positive effects on traffic performance, including better-defined Network Macroscopic Fundamental Diagrams (NMFDs), increased network throughput, and reduced delays. The Traffic-responsive Urban Control (TUC) is a real-time network-wide traffic control system with particular design characteristics, such as the balancing of link's occupancies and an inherent gating feature. These characteristics suggest that TUC may enhance the traffic network performance when combined with PC whilst improving the resulting NMFDs and network throughput and delays. Here, in this work, we investigate the effect of feedback perimeter control (FPC), TUC, and their combination on the NMFD and on the traffic conditions of general traffic and public transport in the microsimulation of a realistic model of the Christchurch Central Business District in New Zealand. We perform a thorough investigation of practical aspects of both control strategies and their combination, including parameter tuning and infrastructure requirements, and how they may affect the control system choice. Results show higher throughput and less hysteresis on the NMFDs, particularly when TUC is involved. PC provides benefits concentrated in the protected region which can greatly benefit public transportation if there is an overlap with the transit network. The combination of TUC and FPC boosts network-wide throughput.

33 ADVANCED PROPULSION SYSTEMS↗

Space Link Extension Protocol Emulation for High-Throughput, High-Latency Network Connections

New space missions require higher data rates and new protocols to meet these requirements. These high data rate space communication links push the limitations of not only the space communication links, but of the ground communication networks and protocols which forward user data to remote ground stations (GS) for transmission. The Consultative Committee for Space Data Systems, (CCSDS) Space Link Extension (SLE) standard protocol is one protocol that has been proposed for use by the NASA Space Network (SN) Ground Segment Sustainment (SGSS) program. New protocol implementations must be carefully tested to ensure that they provide the required functionality, especially because of the remote nature of spacecraft. The SLE protocol standard has been tested in the NASA Glenn Research Center's SCENIC Emulation Lab in order to observe its operation under realistic network delay conditions. More specifically, the delay between then NASA Integrated Services Network (NISN) and spacecraft has been emulated. The round trip time (RTT) delay for the continental NISN network has been shown to be up to 120ms; as such the SLE protocol was tested with network delays ranging from 0ms to 200ms. Both a base network condition and an SLE connection were tested with these RTT delays, and the reaction of both network tests to the delay conditions were recorded. Throughput for both of these links was set at 1.2Gbps. The results will show that, in the presence of realistic network delay, the SLE link throughput is significantly reduced while the base network throughput however remained at the 1.2Gbps specification. The decrease in SLE throughput has been attributed to the implementation's use of blocking calls. The decrease in throughput is not acceptable for high data rate links, as the link requires constant data a flow in order for spacecraft and ground radios to stay synchronized, unless significant data is queued a the ground station. In cases where queuing the data is not an option, such as during real time transmissions, the SLE implementation cannot support high data rate communication.

Computer Networking↗

Feasibility of critical infrastructure protection using network functions for programmable and decoupled ICS policy enforcement over WAN

Industrial control systems (ICS) represent a major component of our critical infrastructure. With the increasing need for more control and monitoring of such systems, ICS have seen an increase in connectivity to wide area networks (WAN) exposing aging equipment to rapidly evolving cybersecurity threats. Furthermore, the ICS data requires a reliability measure from the networks for critical functions for infrastructure monitoring and control. Especially when remote plant sites are involved such as pipelines, energy distribution networks, and transportation, WAN transport impairments most often provide a best effort delivery with no strict reliability guarantees. Network functions can provide a vendor agnostic, programmable critical infrastructure protection with a single maintenance, policy determination, and reliability assurance surface. A network function (NF) can be utilized for policy enforcement over the communication between remote entities and the main control office. This paper presents the research on transparent integration with existing ICS without disrupting communications, resulting in minimal downtime while decoupling the fast paced evolution of defensive security measures from the upgrade cycle of expensive long term hardware. We report our measurements on the resource requirements and overhead in the network for successful NF insertion under a wide variety of network impairments (network packet delay, reordering, and loss). Our paired NF implementation provides a policy enforcement platform extensible to cover myriad cybersecurity-related communication goals, including packet signing for verification, encryption for data privacy, packet filtering and data diode operation (i.e. protecting against eavesdropping, packet injection, and denial-of-service). Furthermore, bundling communication specifications into packet flows allows for tunability in applying policies as coarse- or fine-grained as the needs of the operator. We report on network function resource requirements in the form of required queue depth and network utilization overhead to inform the decision making against hardware cost constraints.

42 ENGINEERING↗

Non-unimodal and non-concave relationships in the network Macroscopic Fundamental Diagram caused by hierarchical streets

Unimodal, concave relationships between average network productivity and accumulation or density aggregated across spatially compact regions of urban networks—so called network Macroscopic Fundamental Diagrams (MFDs)—have recently been shown to exist on homogeneous street networks. When present, MFD relationships facilitate the modeling of traffic congestion at a regional level and have led to the development of various regional traffic control strategies. However, real street networks are not homogeneous—they generally have a hierarchical structure where some streets (e.g., arterials) promote higher mobility than others (e.g., local roads). Here, this paper examines how the presence of hierarchical roadway structures may potentially cause non-unimodal patterns in a network's MFD. These are observed using three types of tools: analytical models of simple network structures, simulations of various idealized roadway networks, and empirical data. The impacts of street hierarchy depend on how vehicles use different roadway types to move within the network; i.e., their routing strategy. The findings suggest that the presence of roadway hierarchies may lead to MFDs that have non-unimodal or non-concave patterns on the free-flow branch when vehicles route themselves according to user equilibrium principles, which is closest to what would be observed in realistic situations. Such patterns are contrary to what is traditionally assumed in most MFD-based modeling frameworks. However, the unimodal and concave MFD should be expected under system optimal routing conditions that maximize network productivity for a given traffic state.

42 ENGINEERING↗

Evolved Gas Analysis–Mass Spectrometry Exposes Polymer Network Structures

Polymer network structures in epoxy thermosets play an important role in the final thermoset material properties. However, analytical characterization of these network structures is difficult due to their amorphous nature. In this work, the application of evolved gas analysis–mass spectrometry (EGA-MS) to characterize the polymer network structures of bisphenol A (BPA)-based thermosets is demonstrated. Analytical characterization of the polymer network structures is accomplished by monitoring the Product-Specific Kinetics (PSK) of BPA monomer formation during thermal degradation investigations. We relate observed differences in the activation energy (E a ) of BPA monomer formation to the local packing environment around the BPA monomer units within the polymer network. Variations in the local environment related to the polymer networks manifest qualitatively as broadening in the thermal profile of the BPA monomer evolution and quantitatively as changes in the activation energy (E a ). Three BPA thermoset formulations were investigated; two amine-cured thermoset with 4,4′-diaminodiphenylmethane (DDM) or poly(propylene glycol) bis(2-amino-propyl ether) (PPG400) and a homopolymerized thermoset via curing with Epikure 3253 catalyst (3253). Results revealed that the 3253 thermoset contained two distinct packing densities in the polymer network, while DDM and PPG400 thermosets had uniform distributions of packing densities. Results from the DDM thermoset revealed a gradually decreasing E a , while the apparent E a of PPG400 was consistent over the entire degradation. Furthermore, these differences in E a were concluded to stem from the flexibility of the corresponding polymer networks and the ability of the network components to rearrange and occupy formed voids. Due to the minimal sample required for analysis (100–200 μg), this EGA-MS technique has great potential for postproduction evaluation of composite parts to identify changes in the polymer networks from use and aging, which could signal compromised performance.

Degradation↗

Throughput Estimation of Data Transport Networks From Digital Twin Measurements

Digital twins of networked infrastructures, known as Virtual Infrastructure Twins (VITs), are increasingly used for software development, pre-deployment testing, and design space exploration. While VITs avoid the costs and potential disruptions associated with experiments on operational networks, their throughput measurements are typically not sufficiently accurate for performance profiling of wide-area networks that they emulate. Here, machine learning (ML) methods are developed to transform these inaccurate VIT network throughput measurements to closely match in peak and overall profile of those from a physical testbed or production network. First, a micro kernel network reflecting a physical network is utilized to collect one-time measurements on a host to support this ML transformation. Then, a generic multi-modal ML method is developed to learn a map that transforms measurements from subsequent VITs on the same host to match past, current and follow-on testbed and cloud networks. ML generalization equations are derived to establish its correctness and probabilistically guarantee its generalization accuracy. Experimental results are presented for a variety of VIT hosts with target testbed and cloud networks; they include a case study of a four-site science ecosystem wherein inaccurate convex VIT measurement profiles are transformed into accurate concave profiles of target networks.

97 MATHEMATICS AND COMPUTING↗

Hybrid PDES Simulation of HPC Networks Using Zombie Packets

Although high-fidelity network simulations have proven to be reliable and cost-effective tools to peer into architectural questions for high-performance computing (HPC) networks, they incur a high resource cost. The time spent in simulating a single millisecond of network traffic in the highest detail can take hours, even for static, well-behaved traffic patterns such as uniform random. Surrogate models offer a significant reduction in runtime, yet they cannot serve as complete replacements and should only be used when appropriate. Thus, there is a need for hybrid modeling, where high-fidelity simulation and surrogates run side-by-side. Here, we present a surrogate model for HPC networks in which: packets bypass the network, while the network state is left untouched, i.e., suspended. To bypass the network, we use historical data to estimate the arrival time at which every packet should be scheduled at; to suspend the network, all in-flight packets are scheduled to arrive at their destinations, and are kept in the system to awaken as zombies when switching back to high-fidelity. Speedup for a hybrid model is relative to the proportion of surrogate to high-fidelity. This light-weight surrogate obtained up to 76× speedup. Keeping the zombies in the network showed an increase in the accuracy of the high-fidelity simulation on restart when compared to restarting the network from an empty state.

HPC networks↗