DOE OSTI · 2580087
Tailor : Altering Skip Connections for Resource-Efficient Inference
Abstract
Deep neural networks use skip connections to improve training convergence. However, these skip connections are costly in hardware, requiring extra buffers and increasing on- and off-chip memory utilization and bandwidth requirements. In this article, we show that skip connections can be optimized for hardware when tackled with a hardware-software codesign approach. We argue that while a network’s skip connections are needed for the network to learn, they can later be removed or shortened to provide a more hardware-efficient implementation with minimal to no accuracy loss. We introduceTailor, a codesign tool whose hardware-aware training algorithm gradually removes or shortens a fully trained network’s skip connections to lower the hardware cost.Tailorimproves resource utilization by up to 34% for block random access memories (BRAMs), 13% for flip-flops (FFs), and 16% for look-up tables (LUTs) for on-chip, dataflow-style architectures.Tailorincreases performance by 30% and reduces memory bandwidth by 45% for a two-dimensional processing element array architecture.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Weng, Olivia (ORCID:000000031213421X), Marcano, Gabriel (ORCID:0000000248047305), Loncar, Vladimir (ORCID:0000000336510232), Khodamoradi, Alireza (ORCID:0000000188112258), G, Abarajithan (ORCID:0000000197685349), Sheybani, Nojan (ORCID:0000000243290197), Meza, Andres (ORCID:0000000242830833), Koushanfar, Farinaz (ORCID:0000000307983794), Denolf, Kristof (ORCID:000000022087865X), Duarte, Javier (ORCID:0000000250767096), Kastner, Ryan (ORCID:0000000190625570). 2024-01-27. Tailor : Altering Skip Connections for Resource-Efficient Inference. https://doi.org/10.1145/3624990
Cite the original work for its findings. Save a collection to share your selection of sources.