WishWiki

Attention Is All You Need

"Attention Is All You Need" is a 2017 research paper that quietly detonated a revolution in machine learning design. It is the 2017 NeurIPS paper from Google Brain / Google Research that introduced the Transformer, a sequence-to-sequence architecture built entirely on attention, eliminating recurrence and convolutions. The paper was authored by Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin — eight researchers whose names would become, improbably, some of the most cited in computing history.

Before this paper, sequence modeling — translating a sentence, summarizing a document — leaned on recurrent networks that marched through text one token at a time, or on convolutional stacks that peered at only a local window. Both approaches struggled with long-range dependencies: a pronoun on page three referring back to a name on page one. The Transformer's radical proposal was that recurrence and convolution were unnecessary scaffolding. All you needed was attention — a mechanism for letting every element of a sequence look directly at every other element, weighing relevance regardless of distance.

The Mechanism

At the heart of the architecture is scaled dot-product attention, in which queries are compared against keys to produce weights that blend values together, with the dot products scaled down to keep gradients well-behaved. Rather than computing this once, the model runs it many times in parallel through multi-head attention, letting different heads specialize in different kinds of relationships — one might track syntax, another coreference, another rhythm. Because attention alone has no innate sense of order, the model injects a positional encoding built from sine and cosine waves of different frequencies, so the network can distinguish "the dog bit the man" from its unfortunate reversal. Encoder and decoder stacks are wired together with residual connections and layer normalization.

Why It Mattered

The design's genius was parallelizability: unlike recurrent models, every token's attention can be computed simultaneously, making the Transformer a natural fit for modern GPU clusters — a leap in training efficiency. That scalability, paired with the model's uncanny ability to capture context, became the seed for an entire ecosystem: BERT, GPT, and countless descendants that now underpin translation, search, code generation, and conversational agents. It shifted how models learn meaning — away from hand-built grammar rules and toward patterns learned across enormous text corpora.

Legacy

The Transformer reshaped not just Natural language processing but bled into vision, biology, and beyond, treating images, protein sequences, and molecules alike as "tokens" to attend over. Its title — a playful nod to the sufficiency of a single idea — proved almost prophetic: today it is the foundation of nearly every Large language model.

Related

Sources

  1. Attention Is All You Need, explained like you’re smart and busy | by Adnan Masood, PhD. | Medium
  2. Attention Is All You Need — Grokipedia
  3. The Paper That Built Modern AI: How 'Attention Is All You Need' (2017) Gave Us the Transformer | Jerry Cards
  4. Attention Is All You Need Vaswani et al. NeurIPS 2017 Presented by Luke Song
  5. Attention Is All You Need - Wikipedia
  6. ALL-IN-ONE: Multi-Task Learning BERT models for Evaluating Peer Assessments
  7. ATTENTION IS ALL YOU NEED. Original Paper by Ashish Vaswani, Noam… | by Vishal dhawal | Medium
  8. [1706.03762] Attention Is All You Need
  9. Non-Autoregressive Sentence Ordering
  10. Dynamically Relative Position Encoding-Based Transformer for Automatic Code Edit
Wishing…
your wish is being written

✨ Wish for a new page

👁 Wish for another view of this page

Sign in to WishWiki

Keep your wishes together, see your activity — and later, get your own private wiki space.

⏱ Page history