Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Performance of DSRC V2V Communication Networks in an Autonomous Semi-Truck Platoon Application

Autonomy for multiple trucks to drive in a fixedheadway platoon formation is achieved by adding precision GPS and V2V communications to a conventional adaptive cruise control (ACC) system. The performance of the Cooperative ACC (CACC) system depends heavily on the reliability of the underlying V2V communications network. Using data recorded on precision-instrumented trucks at both ACM and NCAT test tracks, we provide an understanding of various effects on V2V network performance: Occlusions - non-line-of-sight (NLOS) between the Tx and Rx antenna may cause network signal loss. Rain - water droplets in the air may cause network signal degradation. Antenna position - antennas at higher elevation may have less ground clutter to deal with. RF interference - interference may cause network packet loss. GPS outage - outages caused by tree cover, tunnels, etc. may result in degraded performance. Road curvature - curves may affect antenna diversity. Road grade - antenna may have limited vertical coverage. Our results, which include multiple plots and graphs and their analysis, build on those reported by others and could be of interest to researchers, practitioners, and end-users in the broader autonomous-connected vehicle community.

V2V communications, GPS, truck platoon, semi-auton↗

Quartz Dissolution Effects on Flow Channelization and Transport Behavior in Three‐Dimensional Fracture Networks

We perform a set of reactive transport simulations in three-dimensional fracture networks to characterize the impact of geochemical reactions on flow channelization. Flow channelization, a frequently observed phenomenon in porous and fractured subsurface rock formations, results from the spatially variable hydraulic resistance offered by a geological structure. In addition to geo-structural features such as network connectivity, geometry, and hydraulic resistance, geochemical reactions, for example, dissolution and precipitation, can dynamically inhibit or enhance flow channelization. These geochemical processes can change the fracture permeability leading to increased flow channelization, which are localized connected regions of high volumetric flow rates that are seemingly ubiquitous in the subsurface. In our simulations, fractures partially filled with quartz are gradually dissolved until quasi-steady state conditions are obtained. We compare the flow field's initial unreacted and final dissolved states in terms of flow and transport observations. We observe that the dissolved fracture networks provide less resistance to flow and exhibit increased flow channelization when compared to their unreacted counterparts. However, there is substantial variability in the magnitude of these changes which implies that the channelization strongly depends on the network structure. In turn, we identify the interplay between the particular network structure and the impact of geochemical dissolution on flow channelization. The presented results indicate that geological systems that have been weathering or reactive for longer times in older landscapes are likely to have increased flow channelization compared to their equivalent but younger counterparts, which implies a time dependence on flow channelization in fractured media.

Hyman, Jeffrey D.↗

Updated Application of Frequency of Detection Methods for the INL Site Ambient Air Monitoring Network

This report presents a quantitative assessment of the current INL Site air monitoring network using frequency of detection (FD) methods. The first assessment of the INL network was performed in 2015 and made recommendations for improving the network. As a result of changes made in response to the recommendations and the addition of new source locations, the network was modified and reassessed in 2017. Since 2017, administration of the air sampling program has been consolidated under one contractor, which resulted in additional changes to the network (sampler numbers and locations) and changes in radionuclide detection levels. As a result of these changes and others, an updated assessment of the INL Site ambient air monitoring network was performed. The same two exposure scenarios used in previous assessments were used for this assessment: a resident scenario and a shepherd/rancher scenario. The resident was assumed to be continuously present at their residence/business/farm operation outside the INL Site boundary while the shepherd/rancher was assumed to be present 24-hours at the nearest INL grazing allotment boundary in each of the 22.5-degree sectors along the sector centerline from each source. Updates to both the resident and shepherd/rancher receptor locations were included. Other changes include updates to flow rates for stack sources, expansion of the list of important radionuclides based on the most recent National Emission Standards for Hazardous Pollutants (NESHAPs) analysis, and updated dose coefficients. The assessment was conducted to determine whether the current INL monitoring network is capable of detecting releases of important radionuclides from INL Site sources that have the potential to exceed a conservative dose threshold for the two exposure scenarios. The assessment revealed that for the resident scenario, the current network meets the desired performance objective (FD = 95%) for all radionuclides and sources except for Cl-36 from the TRA-770 stack (94.4%). For the shepherd/rancher scenario, the FD performance objective is met for all radionuclides and sources except tritium from MFC-774 and TAN 679 (91% for both). An investigation of reported emissions for the past three years revealed that Cl 36 is not emitted from TRA-770, and routine tritium emissions from MFC-774 and TAN-679 are very small and the sources are likely incapable of emitting enough tritium to cause a release that should be detectable by the monitoring network. This assessment is based on a conservative dose threshold. This coupled with fact that the FD for Cl 36 is only slightly less than the performance objective and Cl-36 is not emitted from TRA-770, modifying the network (i.e. adding another sampler, increasing sampler flow rate, moving samplers) to meet the 95% performance objective for this radionuclide/source/receptor scenario is not warranted. Similarly, because tritium emissions from MFC-774 and TAN-679 are very small and these two sources are likely incapable of causing a dose due to tritium release that should be detectable by the network, modifications to increase tritium detection for these sources is also unwarranted at this time. However, if it is required to meet the performance objective for tritium for all sources and receptor scenarios, additional analysis determined the FD could be raised from 91% to > 99% for both sources by adding two tritium samplers to the network.

61 RADIATION PROTECTION AND DOSIMETRY↗

Path-synchronous performance monitoring of interconnection networks based on source code attribution

Examples disclosed herein relate to path-synchronous performance monitoring of an interconnection network based on source code attribution. A processing node in the interconnection network has a profiler module to select a network transaction to be monitored, determine a source code attribution associated with the network transaction to be monitored, and issue a network command to execute the network transaction to be monitored. A logger module creates, in a buffer, a node temporal log associated with the network transaction and the network command. A drainer module periodically captures the node temporal log. The processing node has a network interface controller to receive the network command and mark a packet generated for the network command to be temporally tracked and attributed back to the source code attribution at each hop of the interconnection network traversed by the marked packet.

Chabbi, Milind M.↗

Inference-Engine v0.1.0

Given a pre-trained neural network, Inference-Engine performs maps network inputs to outputs by executing the forward pass through the provided network. Although the predominant programming language for machine-learning is Python, most high-performance computing (HPC) applications are written in Fortran, C, or C++. Inference-Engine aims to support HPC programs and is written in Fortran, a language with a large feature set supporting interoperability with C. This software exposes concurrency in a portable way by using standard language features that some modern Fortran compilers can exploit with various optimizations, including offloading computation to a Graphics Processing Unit (GPU). In particular, this software makes extensive use of Fortran's "do concurrent" parallel loop construct, implicitly parallel array statements, and pure procedures that can be invoked inside "do concurrent" blocks. Inference-Engine also supports dynamic choice of inference methods at runtime. Two current options include one method that uses Fortran's "dot_product" intrinsic function inside "do concurrent" blocks and another method that instead uses Fortran' "matmul" array intrinsic function. We plan to investigate automatic compiler offloading of "do concurrent" calculations to GPUs and compile-time substitution of optimized libraries such as the Basic Linear Algebra Library (BLAS) for "matmul" invocations. We also envision the potential for the choice of which method to use could happen at program launch based on in situ performance measurements on any given platform.

Rouson, Damian↗

Ergodicity, lack thereof, and the performance of reservoir computing with memristive networks and nanowire

Networks composed of nanoscale memristive components, such as nanowire and nanoparticle networks, have recently received considerable attention because of their potential use as neuromorphic devices. In this study, we explore ergodicity in memristive networks, showing that the performance on machine leaning tasks improves when these networks are tuned to operate at the edge between two global stability points. We find this lack of ergodicity is associated with the emergence of memory in the system. We measure the level of ergodicity using the Thirumalai-Mountain metric, and we show that in the absence of ergodicity, two different memristive network systems show improved performance when utilized as reservoir computers (RC). We highlight that it is also important to let the system synchronize to the input signal in order for the performance of the RC to exhibit improvements over the baseline.

97 MATHEMATICS AND COMPUTING↗

Ergodicity, lack thereof, and the performance of reservoir computing with memristive networks

Abstract Networks composed of nanoscale memristive components, such as nanowire and nanoparticle networks, have recently received considerable attention because of their potential use as neuromorphic devices. In this study, we explore ergodicity in memristive networks, showing that the performance on machine leaning tasks improves when these networks are tuned to operate at the edge between two global stability points. We find this lack of ergodicity is associated with the emergence of memory in the system. We measure the level of ergodicity using the Thirumalai-Mountain metric, and we show that in the absence of ergodicity, two different memristive network systems show improved performance when utilized as reservoir computers (RC). We highlight that it is also important to let the system synchronize to the input signal in order for the performance of the RC to exhibit improvements over the baseline.

97 MATHEMATICS AND COMPUTING↗

The optimal use of segmentation for sampling calorimeters

One of the key design choices of any sampling calorimeter is how fine to make the longitudinal and transverse segmentation. Here, to inform this choice, we study the impact of calorimeter segmentation on energy reconstruction. To ensure that the trends are due entirely to hardware and not to a sub-optimal use of segmentation, we deploy deep neural networks to perform the reconstruction. These networks make use of all available information by representing the calorimeter as a point cloud. To demonstrate our approach, we simulate a detector similar to the forward calorimeter system intended for use in the ePIC detector, which will operate at the upcoming Electron Ion Collider. We find that for the energy estimation of isolated charged pion showers, relatively fine longitudinal segmentation is key to achieving an energy resolution that is better than 10% across the full phase space. These results provide a valuable benchmark for ongoing EIC detector optimizations and may also inform future studies involving high-granularity calorimeters in other experiments at various facilities.

47 OTHER INSTRUMENTATION↗

Two-fermion negativity and confinement in the Schwinger model

We consider the fermionic (logarithmic) negativity between two fermionic modes in the Schwinger model. Recent results pointed out that fermionic systems can exhibit stronger entanglement than bosonic systems, exhibiting a negativity that decays only algebraically. The Schwinger model is described by fermionic excitations at short distances, while its asymptotic spectrum is the one of a bosonic theory. We show that the two-mode negativity detects this confining, fermion-to-boson transition, shifting from an algebraic decay to an exponential decay at distances of the order of the de Broglie wavelength of the first excited state. We derive analytical expressions in the massless Schwinger model and confront them with tensor network simulations. We also perform tensor network simulations in the massive model, which is not solvable analytically, and close to the Ising quantum critical point of the Schwinger model, where we show that the negativity behaves as its bosonic counterpart. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Inference of three-dimensional hot-spot and shell morphology in inertial confinement fusion experiments using a convolutional neural network

The performance of inertial confinement fusion (ICF) implosions is sensitive to the three-dimensional (3D) morphology of the hot-spot and shell configurations. The ability to infer shell-mass uniformity and reconstruct 3D hot spots is crucial for quantifying the degradation of ignition criteria and improving symmetry in ICF implosion experiments. In this work, we present a deep-learning convolutional neural network (CNN) for reconstructing 3D hot-spot and shell structures for ICF capsules. The 3D geometry of the hot spot is reconstructed from x-ray images measured from multiple lines of sight on OMEGA. The shell configuration is inferred indirectly through machine learning using a convolutional neural network extensively trained on a dec3d simulation database. This simulation-dependent approach yields consistent agreement between reconstructed 3D shell densities and machine-learning optimized dec3d simulation results. This work demonstrates a CNN framework that successfully reconstructs 3D capsule structures from two-dimensional images in ICF implosions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Assessing parallel path cooling tower performance via artificial neural networks

Real-time monitoring of a research nuclear reactor, a system in which all generated power is dissipated to the environment, can be performed via analysis of the heat rejection from the cooling system. Given an inlet water temperature and flow rate, the reactor power can be well-approximated from the outlet water temperature; however, the instrumentation to measure outlet conditions may not be robust or accurate. If we know how a cooling tower performs from historical data, but cannot measure the outlet temperature, a mathematical representation of the system can be inverted to obtain the outlet water temperature that describes the cooling capacity. Unfortunately, model inversion processes are computationally expensive. To address this, an artificial neural network (ANN) is implemented to assess the performance of a multi-cell cooling tower for a nuclear reactor. This approach leverages the Merkel model to obtain an extensive data set describing performance of the cooling tower cells throughout a wide array of potential operating conditions. The Merkel model is expressed as a function of four parameters: the inlet and outlet water temperatures, inlet air wet bulb temperature, and ratio of liquid-to-gas mass flow rates (L/G), which together provide a non-dimensional number indicative of cooling tower performance, called the Merkel integral. Computing a 4-dimensional data structure that describes finite combinations of the Merkel integral, an inverse model is then generated using an ANN to determine the cell outlet water temperature from the other three model parameters along with the computed Merkel integral. Compared to traditional model inversion methods, the ANN reduces the computational time by approximately 4 orders of magnitude, with effectively no sacrifice to solution accuracy, and could be applied for different cooling towers in the event the performance curve is known. Finally, three use cases of the ANN are then reviewed: (1) determining the cell outlet water temperatures when gas flow at rated conditions (GFRC) is known, (2) performing the prior case without knowledge of the GRFC, and (3) assessing performance differences between the individual tower cells.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

The two main thrusts of computational science are increasingly accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g. neural network potentials, and novel hardware architectures, e.g. GPUs. Current implementations of neural network potentials are orders of magnitude slower than traditional interatomic models and while looming exascale computing offers the ability to run large, accurate simulations with these models, achieving portable performance for MD with new and varied exascale hardware requires rethinking traditional algorithms, using novel data structures, and library solutions. We re-implement a neural network interatomic model in CabanaMD, an MD proxy application, built on libraries developed for performance portability. Our implementation shows significantly improved thread scaling in this complex kernel as compared to a current LAMMPS implementation, across both strong and weak scaling. Our single-source solution enables simulations up to 20 million atoms on a single CPU node and 4 million atoms with improved performance on a single GPU. Furthermore, we also explore parallelism and data layout choices (using flexible data structures called AoSoAs) and their effect on performance, seeing up to ~50% and ~5% improvements in performance on a GPU by choosing the right level of parallelism and data layout respectively.

97 MATHEMATICS AND COMPUTING↗

How to Design and Implement an Equitable Building Performance Standard: Lessons from the Building Performance Standards Technical Assistance Network

States and local governments seek to accomplish the intersecting goals of reducing greenhouse gas emissions, improving building operations, and bettering the daily lives of their communities. Building Performance Standards (BPS) have emerged as a critical policy lever to reach these intertwined climate and societal goals. These policies, if shaped and implemented well, have the chance to not only help reach our nation's climate, energy, and livability goals, but to do so with the active participation of those traditionally excluded from policy processes. This nascent policy movement provides jurisdictions across the country the opportunity to shape these policies from the outset to deliver comfort, health, and safety in our built environment for all. The US Department of Energy in partnership with our National Laboratories have been providing technical assistance for jurisdictions interested in, adopting, and implementing Building Performance Standards. Through this work, the BPS Technical Assistance Network (TA Network) has tracked and documented the innovative and equitable approaches to BPS across the country. The TA Network has crafted foundational technical analysis, such as building stock and emissions impacts, equity prioritization and peak load impacts, and aggregated cost benefit analysis. And by combining powerful technical analysis with dissemination of best practices resources to support equitable implementation, the TA Network provides jurisdictions with the tools and support necessary to embark on their ambitious policy goals.

building performance standards↗

Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs

Despite foreseeing tremendous speedups over conventional deep neural networks, the performance advantage of binarized neural networks (BNNs) has merely been showcased on general-purpose processors such as CPUs and GPUs. In fact, due to being unable to leverage bit-level-parallelism with a word-based architecture, GPUs have been criticized for extremely low utilization (1%) when executing BNNs. Consequently, the latest tensorcores in NVIDIA Turing GPUs start to experimentally support bit computation. In this work, we look into this brand new bit computation capability and characterize its unique features. We show that the stride of memory access can significantly affect performance delivery and a data-format co-design is highly desired to support the tensorcores for achieving superior performance than existing software solutions without tensorcores. We realize the tensorcore-accelerated BNN design, particularly the major functions for fully-connect and convolution layers — bit matrix multiplication and bit convolution. Evaluations on two NVIDIA Turing GPUs show that, with ResNet-18, our BTC-BNN design can process ImageNet at a rate of 5.6K images per second, 77% faster than state-of-the-art. Our BNN approach is released on https://github.com/pnnl/TCBNN.

Li, Ang↗

Performance trade-offs in reconfigurable networks for HPC

Designing efficient interconnects to support high-bandwidth and low-latency communication is critical toward realizing high performance computing (HPC) and data center (DC) systems in the exascale era. At extreme computing scales, providing the requisite bandwidth through overprovisioning becomes impractical. These challenges have motivated studies exploring reconfigurable network architectures that can adapt to traffic patterns at runtime using optical circuit switching. Despite the plethora of proposed architectures, surprisingly little is known about the relative performances and trade-offs among different reconfigurable network designs. We aim to bridge this gap by tackling two key issues in reconfigurable network design. First, we study how cost, power consumption, network performance, and scalability vary based on optical circuit switch (OCS) placement in the physical topology. Specifically, we consider two classes of reconfigurable architectures: one that places OCSs between top-of-rack (ToR) switches—ToR-reconfigurable networks (TRNs)—and one that places OCSs between pods of racks—pod-reconfigurable networks (PRNs). Second, we tackle the effects of reconfiguration frequency on network performance. Our results, based on network simulations driven by real HPC and DC workloads, show that while TRNs are optimized for low fan-out communication patterns, they are less suited for carrying high fan-out workloads. PRNs exhibit better overall trade-off, capable of performing comparably to a fully non-blocking fat tree for low fan-out workloads, and significantly outperform TRNs for high fan-out communication patterns.

Teh, Min Yee↗

Low Size, Weight, and Power Neuromorphic Computing to Improve Combustion Engine Efficiency

Neuromorphic computing offers one path forward for AI at the edge. However, accessing and effectively utilizing a neuromorphic hardware platform is non-trivial. In this work, we present a complete pipeline for neuromorphic computing at the edge, including a small, inexpensive, low-power, FPGA-based neuromorphic hardware platform, a training algorithm for designing spiking neural networks for neuromorphic hardware, and a software framework for connecting those components. We demonstrate this pipeline on a real-world application, engine control for a spark-ignition internal combustion engine. We illustrate how we connect engine simulations with neuromorphic hardware simulations and training software to produce hardware-compatible spiking neural networks that perform engine control to improve fuel efficiency. We present initial results on the performance of these spiking neural networks and illustrate that they outperform open-loop engine control. We also give size, weight, and power estimates for a deployed solution of this type.

Schuman, Catherine↗