SEARCH · Engineering Papers
Results for “network performance”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Systems and methods for rapid processing and storage of data
Systems and methods of building massively parallel computing systems using low power computing complexes in accordance with embodiments of the invention are disclosed. A massively parallel computing system in accordance with one embodiment of the invention includes at least one Solid State Blade configured to communicate via a high performance network fabric. In addition, each Solid State Blade includes a processor configured to communicate with a plurality of low power computing complexes interconnected by a router, and each low power computing complex includes at least one general processing core, an accelerator, an I/O interface, and cache memory and is configured to communicate with non-volatile solid state memory.
Sensitivity Study of Mini-Batch Size on a Long Short-Term Memory Network for In-situ Sensing of Core-to-shell Ratio of Microencapsulated Phase Change Materials
Microencapsulated phase change materials are being studied for applications for thermal energy storage in concentrated solar fields. During fabrication, the thickness of the encapsulation cannot be readily measured for real-time control. Therefore, a machine learning network, specifically a Long Short-Term Memory network, is being developed to estimate the ratio of the shell radius to core radius based on a one second temperature history. The mini-batch size determines how often the algorithm weights are updated during network training, and shuffle indicates whether the training data is shuffled during training. A general factorial design is used to analyze the effects of varying mini-batch size and shuffle, along with the core-to-shell ratio, on the RMSE of the response from the Long Short-Term Memory network. It was found that the network performed better for smaller core to shell ratios (less than 0.6) and had the lowest RMSE when the minibatch size was 128. The minimum RMSE found was 0.00501.
Toward designing effective exascale scientific computing workflows: experiences and best practices
Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.
BitGNN: Unlocking the Performance Potential of Binary Graph Neural Networks on GPUs
Graph Neural Networks (GNNs) have shown compelling results in many graph-based learning tasks. They are, however, time-consuming. Recent work has shown a promising direction in improving GNN speed and shrinking the size — network binarization, which binarizes network values and operations. Prior work, however, mainly focused on algorithm designs, leaving it open on how to fully materialize the performance potential. This work fills the gap by proposing techniques to best map binary GNNs and their computations to fit the nature of bit manipulations, optimizations and algorithms to maximize BSpMM kernel efficiency, and solutions to other factors influencing the end-to-end time on GPUs. Results on real-world graphs show that the proposed techniques outperform state of-the-art binary GNN implementations by 21-67× with little accuracy loss.
Predicting PEMFC performance from a volumetric image of catalyst layer structure using pore network modeling
Not Available
Quantifying Negative Effects of Carbon-Binder Networks from Electrochemical Performance of Porous Li-Ion Electrodes
Porous Li-ion electrodes contain active particles, ion transporting electrolyte, and carbon-binder networks. While macrohomogeneous models are often used to predict electrode behavior, accurate predictions remain challenging, owing to the incomplete understanding of the critical role of carbon-binder networks and how they affect the electrochemical response. The present study systematically characterizes these effects in terms of effective properties by utilizing macrohomogeneous models to analyze the measured responses for electrodes with different carbon-binder content, electrode thickness, and porosity but with identical materials. We find that the impact of the carbon-binder network is more severe than previously thought. Even for low carbon-binder content (5 %wt. dry electrode), the presence of the network decreases the reaction area and increases the ion transport resistance, negatively impacting electrode performance. These effects scale with not just porosity or active material volume but also with carbon-binder content. The findings underscore the importance of connecting all effective properties to electrode specifications in a full factorial sense to transform the electrode design paradigm.
Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL
SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.
Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL
SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.
Evaluating factors influencing infrasonic signal detection and automatic processing performance utilizing a regional network
Physical and deployment factors that influence infrasound signal detection and assess automatic detection performance for a regional infrasound network of arrays in the Western U.S. are explored using signatures of ground truth (GT) explosions (yields). Despite these repeated known sources, published infrasound event bulletins contain few GT events. Arrays are primarily distributed toward the south-southeast and south-southwest at distances between 84 and 458 km of the source with one array offering azimuthal resolution toward the northeast. Events occurred throughout the spring, summer, and fall of 2012 with the majority occurring during the summer months. Depending upon the array, automatic detection, which utilizes the adaptive F-detector successfully, identifies between 14% and 80% of the GT events, whereas a subsequent analyst review increases successful detection to 24%–90%. Combined background noise quantification, atmospheric propagation analyses, and comparison of spectral amplitudes determine the mechanisms that contribute to missed detections across the network. This analysis provides an estimate of detector performance across the network, as well as a qualitative assessment of conditions that impact infrasound monitoring capabilities. Finally, the mechanisms that lead to missed detections at individual arrays contribute to network-level estimates of detection capabilities and provide a basis for deployment decisions for regional infrasound arrays in areas of interest.
Self-Sacrificial Template Synthesis of Fe-N-C Catalysts with Dense Active Sites Deposited on A Porous Carbon Network for High Performance in PEMFC
In this study, iron-nitrogen-carbon (Fe-N-C) single-atom catalysts are promising sustainable alternatives to the costly and scarce platinum (Pt) to catalyze the oxygen reduction reactions (ORR) at the cathode of proton exchange membrane fuel cells (PEMFCs). However, Fe-N-C cathodes for PEMFC are made thicker than Pt/C ones, in order to compensate for the lower intrinsic ORR activity and site density of Fe-N-C materials. The thick electrodes are bound with mass transport issues that limit their performance at high current densities, especially in H 2 /air PEMFCs. Practical Fe-N-C electrodes must combine high intrinsic ORR activity, high site density, and fast mass transport. Herein, it has achieved an improved combination of these properties with a Fe-N-C catalyst prepared via a two-step synthesis approach, constructing first a porous zinc-nitrogen-carbon (Zn-N-C) substrate, followed by transmetallating Zn by Fe via chemical vapor deposition. A cathode comprising this Fe-N-C catalyst has exhibited a maximum power density of 0.53 W cm -2 in H 2 /air PEMFC at 80 °C. The improved power density is associated with the hierarchical porosity of the Zn-N-C substrate of this work, which is achieved by epitaxial growth of ZIF-8 onto g-C 3 N 4 , leading to a micro-mesoporous substrate.
Surface-modified and oven-dried microfibrillated cellulose reinforced biocomposites: Cellulose network enabled high performance
Microfibrillated cellulose (MFC) is widely used as a reinforcement filler for biocomposites due to its unique properties. However, the challenge of drying MFC and the incompatibility between nanocellulose and polymer matrix still limits the mechanical performance of MFC-reinforced biocomposites. In this study, we used a water-based transesterification reaction to functionalize MFC and explored the capability of oven-dried MFC as a reinforcement filler for polylactic acid (PLA). Remarkably, this oven-dried, vinyl laurate–modified MFC improved the tensile strength by 38 % and Young’s modulus by 71 % compared with neat PLA. Our results suggested improved compatibility and dispersion of the fibrils in PLA after modification. This study demonstrated that scalable water-based surface modification and subsequent straightforward oven drying could be a facile method for effectively drying cellulose nanomaterials. We find that the method helps significantly disperse fibrils in polymers and enhances the mechanical properties of microfibrillar cellulose-reinforced biocomposites.
Protocols for estimating multiple functions with quantum sensor networks: Geometry and performance
Not Available
Using Monitoring Data to Improve HPC Performance via Network-Data-Driven Allocation.
Abstract not provided.
Using Monitoring Data to Improve HPC Performance via Network-Data-Driven Allocation.
Abstract not provided.
Performance Evaluation of Network Flow and Device Classification using Network Features and Device Embeddings
Explore the source record for details and available documents.
Intelligent Networks for High Performance Computing.
Abstract not provided.