Skip to main content
Skip to main content
Back to Blog
AI IndustryTorontoCommunity

Attention Is All You Need Complete Guide: The 2017 Paper That Launched the AI Revolution—And Canada's Role In It

Complete guide to 'Attention Is All You Need'—the 2017 Google paper that introduced the Transformer architecture and changed AI forever. Featuring Canadian co-author Aidan Gomez, now Cohere's CEO.

A

AGI House Canada

Community Team

January 13, 202614 min read
Attention Is All You Need Complete Guide: The 2017 Paper That Launched the AI Revolution—And Canada's Role In It

On June 12, 2017, eight researchers at Google submitted a paper to arXiv with a title borrowed from the Beatles: "Attention Is All You Need." Seven years later, that paper has accumulated over 173,000 citations—making it the seventh most-cited paper of the 21st century—and its core innovation, the Transformer architecture, underpins virtually every AI system making headlines today: ChatGPT, Claude, Gemini, Llama, and thousands more.

Among those eight authors was a 20-year-old University of Toronto undergraduate named Aidan Gomez, interning at Google Brain. He would go on to co-found Cohere, now valued at 7 billion USD. This is the complete guide to the paper that changed everything—and Canada's surprising role in the AI revolution it sparked.

What Is "Attention Is All You Need"?

"Attention Is All You Need" is a research paper published in 2017 that introduced the Transformer, a new neural network architecture that replaced recurrent and convolutional approaches with self-attention mechanisms.

Paper at a Glance

DetailInformation
TitleAttention Is All You Need
PublishedJune 12, 2017 (arXiv)
ConferenceNeurIPS 2017 (31st Conference on Neural Information Processing Systems)
Authors8 (all at Google at publication)
Citations173,000+ (as of January 2026)
Ranking7th most-cited paper of the 21st century
Primary InnovationTransformer architecture with self-attention

The Title's Origin

The paper's title is a deliberate reference to "All You Need Is Love" by the Beatles—a playful nod that has since become one of the most recognized paper titles in computer science history.


Who Wrote "Attention Is All You Need"?

The paper was authored by eight researchers, all working at Google at the time of publication. In the years since, every author has left Google, and six have founded or co-founded AI startups collectively worth tens of billions of dollars.

The "Transformer 8" Authors

AuthorRole in PaperCurrent Status (2026)
Ashish VaswaniLead author, designed first models with IlliaCEO, Essential AI (56.5M USD raised)
Noam ShazeerScaled dot-product attention, multi-head attentionReturned to Google (2.7B USD deal), Gemini technical lead
Niki ParmarDesigned and tuned model variantsCo-founder, Essential AI
Jakob UszkoreitProposed replacing RNNs with self-attentionCo-founder, Inceptive
Llion JonesInitial codebase, efficient inferenceCTO, Sakana AI (200M USD raised)
Aidan N. GomezImplemented tensor2tensor with KaiserCEO, Cohere (7B USD valuation)
Łukasz KaiserImplemented tensor2tensor with GomezOpenAI researcher
Illia PolosukhinDesigned first models with AshishCEO, NEAR Foundation (550M+ raised)

Individual Contributions

According to the paper's acknowledgments section:

  • Jakob Uszkoreit proposed the idea of replacing RNNs with self-attention and started the research effort
  • Ashish Vaswani and Illia Polosukhin designed and implemented the first Transformer models
  • Noam Shazeer proposed scaled dot-product attention, multi-head attention, and parameter-free position representation
  • Niki Parmar designed, implemented, tuned, and evaluated countless model variants
  • Llion Jones built the initial codebase and developed efficient inference and visualizations
  • Łukasz Kaiser and Aidan Gomez spent "countless long days" implementing tensor2tensor

The Canadian Connection: Aidan Gomez

The most direct Canadian connection to the Transformer paper is Aidan Gomez, who was a 20-year-old University of Toronto undergraduate when he co-authored "Attention Is All You Need" as a Google Brain intern.

Aidan Gomez Profile

DetailInformation
Birth~1996/1997
OriginBrighton, Ontario, Canada
EducationBSc Computer Science & Mathematics, University of Toronto (2013-2018)
PhDUniversity of Oxford (completed 2024, paused to found Cohere)
Google RoleIntern under Łukasz Kaiser and Geoffrey Hinton
Current RoleCEO and Co-founder, Cohere

From Intern to Founder

Gomez and his seven colleagues famously raced to finish the paper for the NeurIPS deadline, even sleeping in the office to complete it. After the internship, Gomez returned to U of T to finish his undergraduate degree, then began a PhD at Oxford.

In 2019, Gomez paused his doctoral studies to co-found Cohere with Ivan Zhang (a former Google AI researcher) and Nick Frosst (a Hinton student). Cohere has since raised over 1 billion USD and achieved a 7 billion USD valuation, becoming one of Canada's most valuable private AI companies.

Recognition

  • TIME 100 AI (2023): Named among the most influential people in artificial intelligence
  • Maclean's AI Trailblazers (2023): Ranked #1 alongside Cohere co-founders
  • Rivian Board (2025): Elected to the board of the electric vehicle company

What Problem Did the Transformer Solve?

Before the Transformer, most natural language processing relied on Recurrent Neural Networks (RNNs) and their variants like Long Short-Term Memory (LSTM) networks.

The RNN Problem

LimitationDescription
Sequential ProcessingRNNs process tokens one at a time in order
Slow TrainingCannot parallelize across sequence positions
Long-Range DependenciesStruggle to connect distant words in text
Vanishing GradientsInformation degrades over long sequences

The Self-Attention Solution

The Transformer's key insight was that attention mechanisms alone—without recurrence or convolution—could model relationships between all positions in a sequence simultaneously.

InnovationBenefit
Parallel ProcessingAll positions computed simultaneously
Constant Path LengthAny two positions connected in O(1)
Scalable TrainingTraining time dramatically reduced
Better Long-RangeDirect connections between distant tokens

How Does the Transformer Architecture Work?

The Transformer consists of an encoder (for understanding input) and a decoder (for generating output), each built from stacked layers of attention and feed-forward networks.

Architecture Components

ComponentFunction
Input EmbeddingConvert tokens to vectors
Positional EncodingAdd position information (since no recurrence)
Multi-Head AttentionAttend to different representation subspaces
Feed-Forward NetworkProcess each position independently
Layer NormalizationStabilize training
Residual ConnectionsEnable deep networks

The Self-Attention Mechanism

Self-attention computes relationships between all positions in a sequence using three learned projections:

ProjectionRole
Query (Q)What the current position is looking for
Key (K)What each position offers to be matched
Value (V)What information to retrieve if matched

The famous "scaled dot-product attention" formula:

Attention(Q, K, V) = softmax(QK^T / √d_k) × V

The scaling factor (√d_k) was Noam Shazeer's contribution—it prevents dot products from growing too large and pushing softmax into regions with tiny gradients.

Multi-Head Attention

Rather than performing attention once, the Transformer uses multiple "heads" that learn different types of relationships:

AspectDetails
Original Paper8 attention heads
Head Dimension64 (512 total / 8 heads)
PurposeDifferent heads learn different relationship types

Original Model Size

ParameterBase ModelBig Model
Layers66
Model Dimension5121024
Feed-Forward20484096
Attention Heads816
Total Parameters65M213M

By modern standards, these are tiny—GPT-4 is estimated at over 1 trillion parameters—but the architecture's scalability was key to its success.


What Were the Paper's Results?

The paper demonstrated state-of-the-art performance on machine translation tasks while requiring far less training time than previous approaches.

Translation Benchmarks

TaskBLEU ScorePrevious BestImprovement
WMT 2014 En-De28.426.2+2.2 BLEU
WMT 2014 En-Fr41.840.4+1.4 BLEU

Training Efficiency

ComparisonTraining Cost
Transformer (Big)3.5 days on 8 GPUs
Previous SOTAWeeks to months

The paper emphasized that "Transformers can be trained significantly faster than architectures based on recurrent or convolutional layers."


What Did the Transformer Enable?

The Transformer architecture became the foundation for virtually all major AI advances since 2017.

Direct Descendants

ModelYearOrganizationArchitecture
BERT2018GoogleEncoder-only Transformer
GPT2018OpenAIDecoder-only Transformer
GPT-22019OpenAI1.5B parameter decoder
T52019GoogleEncoder-decoder
GPT-32020OpenAI175B parameter decoder
ChatGPT2022OpenAIGPT-3.5/4 + RLHF
Claude2023AnthropicConstitutional AI
Gemini2023GoogleMultimodal Transformer
Llama2023-24MetaOpen-source decoder

Beyond Text

DomainApplication
VisionVision Transformer (ViT), DALL-E, Stable Diffusion
AudioWhisper, MusicLM
ProteinAlphaFold 2
RoboticsRT-2, PaLM-E
MultimodalGPT-4V, Gemini

Where Are the Authors Now?

The eight authors have collectively founded or joined companies worth over 100 billion USD in combined valuation.

Career Trajectories

Ashish Vaswani (Lead Author)

  • Left Google 2021 for Adept AI Labs
  • Co-founded Essential AI with Niki Parmar in 2023
  • Raised 56.5M USD (investors: Google, NVIDIA, AMD)
  • Bloomberg profile: "The AI Pioneer Trying to Save AI From Big Tech"

Noam Shazeer (Multi-Head Attention)

  • Built Meena chatbot at Google
  • Left 2021 when Google refused to release it
  • Co-founded Character.AI (valued at 1B USD)
  • Returned to Google 2024 for 2.7B USD (estimated personal gain: 750M-1B USD)
  • Now technical lead on Gemini

Illia Polosukhin (First Models)

  • Founded NEAR Protocol in 2017 (blockchain for AI)
  • NEAR launched 2020, raised 550M+ USD
  • CEO of NEAR Foundation since November 2023
  • Only Web3 founder invited to speak at NVIDIA GTC 2025

Llion Jones (Initial Codebase)

  • Left Google 2023
  • Co-founded Sakana AI in Tokyo
  • Raised 200M USD Series A (2024)
  • Publicly says he's "absolutely sick of transformers"

Aidan Gomez (tensor2tensor)

  • Co-founded Cohere 2019
  • 7B USD valuation (2024)
  • TIME 100 AI (2023)
  • PhD completed at Oxford (2024)
  • Elected to Rivian board (2025)

Łukasz Kaiser (tensor2tensor)

  • Remained at Google until ~2022
  • Joined OpenAI
  • Continues research on language models

Niki Parmar (Model Variants)

  • Co-founded Adept AI Labs
  • Co-founded Essential AI with Vaswani
  • Focus on enterprise AI automation

Jakob Uszkoreit (Original Proposal)

  • Co-founded Inceptive
  • Focus on RNA medicine with AI
  • The person who first proposed replacing RNNs with self-attention

What Is the Paper's Legacy?

"Attention Is All You Need" is considered one of the most influential papers in the history of artificial intelligence.

Citation Trajectory

YearCumulative Citations
2017~100
2018~2,000
2020~20,000
2022~60,000
2024~140,000
2026~173,000+

Recognition

  • 7th most-cited paper of the 21st century (Nature analysis)
  • NeurIPS Test of Time Award consideration
  • GTC 2024: Jensen Huang presented Ashish Vaswani with a signed DGX-1 cover

Semantic Scholar Classification

According to Semantic Scholar:

  • 173,000+ total citations
  • 18,954 "highly influential citations"
  • Papers citing it span NLP, computer vision, robotics, biology, and more

Frequently Asked Questions

What is "Attention Is All You Need"?

"Attention Is All You Need" is a 2017 research paper by eight Google researchers that introduced the Transformer neural network architecture. The paper proposed replacing recurrent neural networks (RNNs) with self-attention mechanisms for sequence modeling, enabling parallel processing and better handling of long-range dependencies. With 173,000+ citations, it is the seventh most-cited paper of the 21st century and the foundation for all modern large language models including ChatGPT, Claude, and Gemini.

Who wrote the Transformer paper?

The eight authors of "Attention Is All You Need" are Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin—all working at Google at the time. Since publication, all eight have left Google. Six have founded AI startups collectively worth over 100 billion USD, including Cohere (Gomez), Character.AI (Shazeer), NEAR Protocol (Polosukhin), Sakana AI (Jones), and Essential AI (Vaswani and Parmar).

What is the Canadian connection to the Transformer paper?

Aidan Gomez, a Canadian computer scientist from Brighton, Ontario, was a co-author of "Attention Is All You Need" while interning at Google Brain as a 20-year-old University of Toronto undergraduate. He interned under Łukasz Kaiser and Geoffrey Hinton. Gomez went on to co-found Cohere in 2019, which has become one of Canada's most valuable private AI companies at 7 billion USD valuation. He was named to TIME's 100 Most Influential People in AI (2023) and completed his PhD at Oxford in 2024.

What is self-attention in Transformers?

Self-attention is the mechanism that allows Transformers to compute relationships between all positions in a sequence simultaneously, regardless of distance. Unlike RNNs that process tokens sequentially, self-attention uses Query, Key, and Value projections to determine which parts of the input are relevant to each position. The formula softmax(QK^T/√d_k)×V computes weighted combinations of values based on query-key similarity. This enables parallel processing and direct connections between any two positions in O(1) path length.

Why is the Transformer important?

The Transformer is important because it became the foundation for all modern AI systems. Before 2017, language models used slow sequential processing (RNNs). The Transformer's parallel architecture enabled training much larger models much faster. This led directly to BERT (2018), GPT (2018), and eventually ChatGPT (2022). Beyond text, Transformers now power computer vision (ViT), protein folding (AlphaFold 2), speech recognition (Whisper), and multimodal AI. The "T" in ChatGPT stands for Transformer.

How many citations does "Attention Is All You Need" have?

As of January 2026, "Attention Is All You Need" has over 173,000 citations, making it the seventh most-cited paper of the 21st century according to a Nature analysis. The paper accumulates approximately 30,000-40,000 new citations per year. On Semantic Scholar, over 18,954 citations are classified as "highly influential"—meaning the citing papers significantly build upon the Transformer architecture rather than merely referencing it.

What startups did the Transformer authors found?

The eight authors have founded or co-founded multiple high-value companies: Cohere (Aidan Gomez, 7B USD), Character.AI (Noam Shazeer, 2.7B USD Google deal), NEAR Protocol (Illia Polosukhin, 550M+ raised), Sakana AI (Llion Jones, 200M USD Series A), Essential AI (Ashish Vaswani and Niki Parmar, 56.5M USD), and Inceptive (Jakob Uszkoreit). Only Łukasz Kaiser remained in a research role, joining OpenAI after leaving Google.

Why is the paper called "Attention Is All You Need"?

The title "Attention Is All You Need" is a reference to the Beatles song "All You Need Is Love"—a playful acknowledgment by the authors. Technically, the title reflects the paper's central claim: that attention mechanisms alone, without recurrence or convolution, are sufficient for state-of-the-art sequence modeling. The paper demonstrated this by achieving the best machine translation results while being significantly faster to train than RNN-based approaches.


Related Reading

Canadian AI Leaders

AI History

Canadian AI Ecosystem


Join AGI House Canada to connect with AI researchers, engineers, and founders building on Transformer technology. Scan the QR code to join our WhatsApp group or subscribe to our newsletter.

Information current as of January 2026. Paper citation counts from Google Scholar and Semantic Scholar.

Share this post
A

AGI House Canada

Community Team