On June 12, 2017, eight researchers at Google submitted a paper to arXiv with a title borrowed from the Beatles: "Attention Is All You Need." Seven years later, that paper has accumulated over 173,000 citations—making it the seventh most-cited paper of the 21st century—and its core innovation, the Transformer architecture, underpins virtually every AI system making headlines today: ChatGPT, Claude, Gemini, Llama, and thousands more.
Among those eight authors was a 20-year-old University of Toronto undergraduate named Aidan Gomez, interning at Google Brain. He would go on to co-found Cohere, now valued at 7 billion USD. This is the complete guide to the paper that changed everything—and Canada's surprising role in the AI revolution it sparked.
What Is "Attention Is All You Need"?
"Attention Is All You Need" is a research paper published in 2017 that introduced the Transformer, a new neural network architecture that replaced recurrent and convolutional approaches with self-attention mechanisms.
Paper at a Glance
The Title's Origin
The paper's title is a deliberate reference to "All You Need Is Love" by the Beatles—a playful nod that has since become one of the most recognized paper titles in computer science history.
Who Wrote "Attention Is All You Need"?
The paper was authored by eight researchers, all working at Google at the time of publication. In the years since, every author has left Google, and six have founded or co-founded AI startups collectively worth tens of billions of dollars.
The "Transformer 8" Authors
Individual Contributions
According to the paper's acknowledgments section:
- Jakob Uszkoreit proposed the idea of replacing RNNs with self-attention and started the research effort
- Ashish Vaswani and Illia Polosukhin designed and implemented the first Transformer models
- Noam Shazeer proposed scaled dot-product attention, multi-head attention, and parameter-free position representation
- Niki Parmar designed, implemented, tuned, and evaluated countless model variants
- Llion Jones built the initial codebase and developed efficient inference and visualizations
- Łukasz Kaiser and Aidan Gomez spent "countless long days" implementing tensor2tensor
The Canadian Connection: Aidan Gomez
The most direct Canadian connection to the Transformer paper is Aidan Gomez, who was a 20-year-old University of Toronto undergraduate when he co-authored "Attention Is All You Need" as a Google Brain intern.
Aidan Gomez Profile
From Intern to Founder
Gomez and his seven colleagues famously raced to finish the paper for the NeurIPS deadline, even sleeping in the office to complete it. After the internship, Gomez returned to U of T to finish his undergraduate degree, then began a PhD at Oxford.
In 2019, Gomez paused his doctoral studies to co-found Cohere with Ivan Zhang (a former Google AI researcher) and Nick Frosst (a Hinton student). Cohere has since raised over 1 billion USD and achieved a 7 billion USD valuation, becoming one of Canada's most valuable private AI companies.
Recognition
- TIME 100 AI (2023): Named among the most influential people in artificial intelligence
- Maclean's AI Trailblazers (2023): Ranked #1 alongside Cohere co-founders
- Rivian Board (2025): Elected to the board of the electric vehicle company
What Problem Did the Transformer Solve?
Before the Transformer, most natural language processing relied on Recurrent Neural Networks (RNNs) and their variants like Long Short-Term Memory (LSTM) networks.
The RNN Problem
The Self-Attention Solution
The Transformer's key insight was that attention mechanisms alone—without recurrence or convolution—could model relationships between all positions in a sequence simultaneously.
How Does the Transformer Architecture Work?
The Transformer consists of an encoder (for understanding input) and a decoder (for generating output), each built from stacked layers of attention and feed-forward networks.
Architecture Components
The Self-Attention Mechanism
Self-attention computes relationships between all positions in a sequence using three learned projections:
The famous "scaled dot-product attention" formula:
Attention(Q, K, V) = softmax(QK^T / √d_k) × V
The scaling factor (√d_k) was Noam Shazeer's contribution—it prevents dot products from growing too large and pushing softmax into regions with tiny gradients.
Multi-Head Attention
Rather than performing attention once, the Transformer uses multiple "heads" that learn different types of relationships:
Original Model Size
By modern standards, these are tiny—GPT-4 is estimated at over 1 trillion parameters—but the architecture's scalability was key to its success.
What Were the Paper's Results?
The paper demonstrated state-of-the-art performance on machine translation tasks while requiring far less training time than previous approaches.
Translation Benchmarks
Training Efficiency
The paper emphasized that "Transformers can be trained significantly faster than architectures based on recurrent or convolutional layers."
What Did the Transformer Enable?
The Transformer architecture became the foundation for virtually all major AI advances since 2017.
Direct Descendants
Beyond Text
Where Are the Authors Now?
The eight authors have collectively founded or joined companies worth over 100 billion USD in combined valuation.
Career Trajectories
Ashish Vaswani (Lead Author)
- Left Google 2021 for Adept AI Labs
- Co-founded Essential AI with Niki Parmar in 2023
- Raised 56.5M USD (investors: Google, NVIDIA, AMD)
- Bloomberg profile: "The AI Pioneer Trying to Save AI From Big Tech"
Noam Shazeer (Multi-Head Attention)
- Built Meena chatbot at Google
- Left 2021 when Google refused to release it
- Co-founded Character.AI (valued at 1B USD)
- Returned to Google 2024 for 2.7B USD (estimated personal gain: 750M-1B USD)
- Now technical lead on Gemini
Illia Polosukhin (First Models)
- Founded NEAR Protocol in 2017 (blockchain for AI)
- NEAR launched 2020, raised 550M+ USD
- CEO of NEAR Foundation since November 2023
- Only Web3 founder invited to speak at NVIDIA GTC 2025
Llion Jones (Initial Codebase)
- Left Google 2023
- Co-founded Sakana AI in Tokyo
- Raised 200M USD Series A (2024)
- Publicly says he's "absolutely sick of transformers"
Aidan Gomez (tensor2tensor)
- Co-founded Cohere 2019
- 7B USD valuation (2024)
- TIME 100 AI (2023)
- PhD completed at Oxford (2024)
- Elected to Rivian board (2025)
Łukasz Kaiser (tensor2tensor)
- Remained at Google until ~2022
- Joined OpenAI
- Continues research on language models
Niki Parmar (Model Variants)
- Co-founded Adept AI Labs
- Co-founded Essential AI with Vaswani
- Focus on enterprise AI automation
Jakob Uszkoreit (Original Proposal)
- Co-founded Inceptive
- Focus on RNA medicine with AI
- The person who first proposed replacing RNNs with self-attention
What Is the Paper's Legacy?
"Attention Is All You Need" is considered one of the most influential papers in the history of artificial intelligence.
Citation Trajectory
Recognition
- 7th most-cited paper of the 21st century (Nature analysis)
- NeurIPS Test of Time Award consideration
- GTC 2024: Jensen Huang presented Ashish Vaswani with a signed DGX-1 cover
Semantic Scholar Classification
According to Semantic Scholar:
- 173,000+ total citations
- 18,954 "highly influential citations"
- Papers citing it span NLP, computer vision, robotics, biology, and more
Frequently Asked Questions
What is "Attention Is All You Need"?
"Attention Is All You Need" is a 2017 research paper by eight Google researchers that introduced the Transformer neural network architecture. The paper proposed replacing recurrent neural networks (RNNs) with self-attention mechanisms for sequence modeling, enabling parallel processing and better handling of long-range dependencies. With 173,000+ citations, it is the seventh most-cited paper of the 21st century and the foundation for all modern large language models including ChatGPT, Claude, and Gemini.
Who wrote the Transformer paper?
The eight authors of "Attention Is All You Need" are Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin—all working at Google at the time. Since publication, all eight have left Google. Six have founded AI startups collectively worth over 100 billion USD, including Cohere (Gomez), Character.AI (Shazeer), NEAR Protocol (Polosukhin), Sakana AI (Jones), and Essential AI (Vaswani and Parmar).
What is the Canadian connection to the Transformer paper?
Aidan Gomez, a Canadian computer scientist from Brighton, Ontario, was a co-author of "Attention Is All You Need" while interning at Google Brain as a 20-year-old University of Toronto undergraduate. He interned under Łukasz Kaiser and Geoffrey Hinton. Gomez went on to co-found Cohere in 2019, which has become one of Canada's most valuable private AI companies at 7 billion USD valuation. He was named to TIME's 100 Most Influential People in AI (2023) and completed his PhD at Oxford in 2024.
What is self-attention in Transformers?
Self-attention is the mechanism that allows Transformers to compute relationships between all positions in a sequence simultaneously, regardless of distance. Unlike RNNs that process tokens sequentially, self-attention uses Query, Key, and Value projections to determine which parts of the input are relevant to each position. The formula softmax(QK^T/√d_k)×V computes weighted combinations of values based on query-key similarity. This enables parallel processing and direct connections between any two positions in O(1) path length.
Why is the Transformer important?
The Transformer is important because it became the foundation for all modern AI systems. Before 2017, language models used slow sequential processing (RNNs). The Transformer's parallel architecture enabled training much larger models much faster. This led directly to BERT (2018), GPT (2018), and eventually ChatGPT (2022). Beyond text, Transformers now power computer vision (ViT), protein folding (AlphaFold 2), speech recognition (Whisper), and multimodal AI. The "T" in ChatGPT stands for Transformer.
How many citations does "Attention Is All You Need" have?
As of January 2026, "Attention Is All You Need" has over 173,000 citations, making it the seventh most-cited paper of the 21st century according to a Nature analysis. The paper accumulates approximately 30,000-40,000 new citations per year. On Semantic Scholar, over 18,954 citations are classified as "highly influential"—meaning the citing papers significantly build upon the Transformer architecture rather than merely referencing it.
What startups did the Transformer authors found?
The eight authors have founded or co-founded multiple high-value companies: Cohere (Aidan Gomez, 7B USD), Character.AI (Noam Shazeer, 2.7B USD Google deal), NEAR Protocol (Illia Polosukhin, 550M+ raised), Sakana AI (Llion Jones, 200M USD Series A), Essential AI (Ashish Vaswani and Niki Parmar, 56.5M USD), and Inceptive (Jakob Uszkoreit). Only Łukasz Kaiser remained in a research role, joining OpenAI after leaving Google.
Why is the paper called "Attention Is All You Need"?
The title "Attention Is All You Need" is a reference to the Beatles song "All You Need Is Love"—a playful acknowledgment by the authors. Technically, the title reflects the paper's central claim: that attention mechanisms alone, without recurrence or convolution, are sufficient for state-of-the-art sequence modeling. The paper demonstrated this by achieving the best machine translation results while being significantly faster to train than RNN-based approaches.
Related Reading
Canadian AI Leaders
- Aidan Gomez Complete Guide — Transformer co-author
- Cohere Enterprise LLM Guide — Gomez's company
- Geoffrey Hinton Complete Guide — Gomez's intern supervisor
AI History
- ImageNet 2012 Moment Guide — The breakthrough that preceded Transformers
- History of AI in Canada 1980-2000 — Early foundations
- CIFAR Complete Guide — Organization that funded deep learning
Canadian AI Ecosystem
- Canadian AI Unicorns Complete Guide — Billion-dollar companies
- Toronto AI Ecosystem Guide — Where Gomez studied
- AI Jobs in Canada Guide — Career opportunities
Join AGI House Canada to connect with AI researchers, engineers, and founders building on Transformer technology. Scan the QR code to join our WhatsApp group or subscribe to our newsletter.
Information current as of January 2026. Paper citation counts from Google Scholar and Semantic Scholar.