Adolescence: 2012–2019
Deep learning broke through when neural networks, large datasets and GPUs finally came together. Image recognition leapt forward, machines learned to play games from raw pixels, and the transformer architecture arrived in 2017.
Most of what is called AI today traces back to ideas that matured in these years, even though the public barely noticed until the end of them.
It learned to see
Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton published AlexNet, a deep convolutional neural network that crushed the competition at the ImageNet Large Scale Visual Recognition Challenge. By training on two GPUs, they proved that deep neural networks could recognize objects with unprecedented accuracy.
Why it mattered. This single result shattered the computer vision establishment's reliance on hand-coded features and ignited the deep learning boom.
papers.nips.ccIt mapped meaning
A team at Google led by Tomas Mikolov introduced Word2Vec, a technique that represented words as continuous vectors in a high-dimensional space. The model captured semantic relationships so well that simple vector arithmetic could solve analogies, famously calculating that 'King - Man + Woman = Queen'.
Why it mattered. Word embeddings became the foundational layer for natural language processing, allowing models to compute meaning rather than just counting words.
papers.nips.ccIt learned to play from pixels
Researchers at the newly formed startup DeepMind published a system that learned to play Atari 2600 games directly from raw pixels. Using a convolutional neural network combined with Q-learning, the agent figured out winning strategies without being taught the rules.
Why it mattered. This was the first successful demonstration of deep reinforcement learning, proving that a single architecture could master complex control tasks from scratch.
arxiv.orgIt learned to imagine
Ian Goodfellow and colleagues introduced Generative Adversarial Networks (GANs), pitting two neural networks against each other. A generator tried to create realistic fake images, while a discriminator tried to tell the fakes from the real ones, forcing both to improve.
Why it mattered. GANs gave AI the ability to generate highly realistic synthetic data, opening the door to deepfakes, AI art, and a new era of generative models.
papers.nips.ccIt learned to translate
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le published 'Sequence to Sequence Learning with Neural Networks'. They used a multilayer Long Short-Term Memory (LSTM) network to map an input sequence to a fixed-length vector, and another to decode it into an output sequence.
Why it mattered. This architecture became the standard for machine translation, allowing models to process entire sentences at once rather than translating word by word.
papers.nips.ccIt learned to go deep
Kaiming He and his team at Microsoft Research introduced Deep Residual Learning (ResNet). By adding 'skip connections' that allowed signals to bypass layers, they successfully trained a 152-layer neural network, which was vastly deeper than previous models.
Why it mattered. ResNet solved the vanishing gradient problem that had prevented networks from scaling, setting a new standard for computer vision architectures that remains in use today.
arxiv.orgIt learned to play Go
DeepMind's AlphaGo defeated 18-time world champion Lee Sedol 4-1 in a five-game match in Seoul. The system combined deep neural networks with Monte Carlo tree search, evaluating board positions and selecting moves in a game long thought too complex for brute-force computation.
Why it mattered. It shattered the timeline for AI progress, achieving a milestone experts believed was still a decade away and demonstrating the power of deep reinforcement learning.
deepmind.googleIt learned to pay attention
Researchers at Google Brain and the University of Toronto published 'Attention Is All You Need', introducing the Transformer architecture. It discarded recurrent neural networks entirely, relying instead on a self-attention mechanism to process sequences of data in parallel.
Why it mattered. It removed the sequential bottleneck of previous models, allowing for massive parallelization during training and setting the foundation for the large language models that followed.
arxiv.orgIt learned the rules of the game
DeepMind introduced AlphaZero, a generalized version of AlphaGo that could master chess, shogi, and Go from scratch. Given only the rules of the games, it trained entirely through self-play, defeating world-champion programs Stockfish, elmo, and AlphaGo Zero within 24 hours.
Why it mattered. It demonstrated that a single reinforcement learning algorithm could achieve superhuman performance across multiple complex domains without relying on human data or domain-specific heuristics.
arxiv.orgIt learned to read in both directions
Google introduced BERT, a model that pre-trained deep bidirectional representations from unlabeled text. By masking words and forcing the model to predict them from surrounding context, it learned deep contextual relations across entire sentences.
Why it mattered. It established pre-training and fine-tuning as the standard paradigm for natural language processing, immediately breaking records on multiple language understanding benchmarks.
arxiv.orgIt learned to write too well
OpenAI announced GPT-2, a 1.5 billion parameter transformer trained to predict the next word on 40GB of internet text. The model generated highly coherent, multi-paragraph text, prompting OpenAI to initially withhold the full model due to concerns about malicious applications like fake news.
Why it mattered. It proved that simply scaling up language models predictably improved their zero-shot performance, while introducing the modern debate over the safe release of powerful AI systems.
openai.comFeed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.