Engineering Papers⌕ Search

Engineering topics

Shi, Runbin

Publications and source records attributed to Shi, Runbin.

O3BNN-R: An Out-Of-Order Architecture for High-Performance and Regularized BNN Inference

Binarized Neural Networks (BNN) have drawn tremendous attention due to significantly reduced computational complexity and memory demand. They have especially shown great potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching a certain accuracy bar is often sufficient, and real-time is highly desired.In this work, we demonstrate that the highly-condensed BNN model can be shrunk significantly further by dynamically pruning irregular redundant edges. Based on two new observations on BNN-specific properties, an out-of-order (OoO) architecture – O3BNN-R, can curtail edge evaluation in cases where the binary output of a neuron can be determined early. Similar to Instruction-Level-Parallelism(ILP), these fine-grained, irregular, runtime pruning opportunities are traditionally presumed to be difficult to exploit. In order to increase the pruning opportunities, we also optimize the training process by adding 2 regularization items in the loss function (1) for pooling pruning and (2) for threshold pruning. We evaluate our design on an FPGA platform using three well-known networks, including VggNet-16, AlexNet for ImageNet, and a VGG-like network for Cifar-10.

Geng, Tong↗

AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload Rebalancing

The recent development of deep learning has been mostly focusing on Euclidean data, such as images, videos, audios, etc. However, most real-world information and relation are often expressed as graphs. To efficiently learn from graph data, graph convolutional networks (GCNs) emerge as a promising approach, showing advantages in several practical applications such as social network analysis, knowledge discovery, 3D modeling, motion capturing, etc. Real-world graphs are usually extremely large and imbalanced, posting significant performance demand and design challenges on the hardware dedicated for GCN inference. In this paper, we propose an architecture design called UW-GCN to accelerate graph convolutional network inference. To tackle the major performance bottleneck from workload imbalance, we propose dynamic neighborhood stealing and remote chunk shuffling techniques, relying on hardware flexibility to achieve hardware auto-tuning under negligible area or delay overhead. Specifically, UW-GCN is able to smartly profile the sparse graph pattern while continuously adjusting the workload distribution via routing reconfiguration among parallel processing elements (PEs). The ideal configuration is then reused in the remaining iterations. To the best of our knowledge, this is the first accelerator design particularly for GCN and the first work relying on hardware auto-tuning, which is normally based on software, to achieve near-optimal workload balance in processing sparse structures.

Geng, Tong↗

CSB-RNN: A Faster-Than-Realtime RNN Acceleration Framework with Compressed Structured Blocks

Recurrent Neural Networks (RNN) is widely applied to temporal sequence analysis, where real-time performance is usually in demand. However, RNN suffers a heavy computational workload as the model comes with a large weight matrix. To alleviate the pain, model compression (pruning) schemes have been proposed for RNN that pruning the redundant (near-zero) weight-values. On the one hand, the non-structured pruning methods achieve a considerable pruning rate while bringing the computational irregularity, which is un-friendly to parallel-hardware. On the other hand, the existing structured pruning methods consider the hardware parallelism; However, they suffer a poor pruning rate due to the restrict constraints on pruning structure. This paper presents CSB-RNN, an optimized full-stack RNN framework with the novel compressed structured block (CSB) technique. The CSB-pruned RNN model comes with both fine-granularity that benefits the pruning rate and regular structure that facilitates the hardware-parallelism. Further, we propose a novel hardware architecture for inferencing the CSB-pruned model. Different from conventional parallel hardware, this architecture solves the block-workload imbalance issue and achieves an over 95% hardware utilization. With the experiments on 10 RNN models in 5 application domains, the CSB-RNN realizes 7×-20× lossless compression and up to 50× acceptable lossy-compression, which is 2×-7× to the prior art. With the addition of the novel hardware, the compressed-RNN inference reaches a super real-time latency of 10-400µs with FPGA implementation.

Shi, Runbin↗