DOE OSTI · 2283346
Vision Transformers Explained Series
Abstract
Since their introduction in 2017 with Attention is All You Need¹, transformers have established themselves as the state of the art for natural language processing (NLP). In 2021, An Image is Worth 16x16 Words successfully adapted transformers for computer vision tasks. Since then, numerous transformer-based architectures have been proposed for computer vision. This article walks through the Vision Transformer (ViT) as laid out in An Image is Worth 16x16 Words.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Callis, Skylar Jean. 2024-01-26. Vision Transformers Explained Series. https://doi.org/10.2172/2283346
Cite the original work for its findings. Save a collection to share your selection of sources.