Coming of age: 2020–2023
Scale changed what models could do. Systems began learning tasks from a handful of examples, predicting protein structures, drawing images from text and holding conversations that millions of people tried for themselves.
This is the stretch in which AI stopped being a research topic and became a product category.
It learned from a handful of examples
OpenAI published GPT-3, a 175-billion parameter language model that demonstrated 'few-shot learning'—the ability to perform new tasks with just a few examples in its prompt. By scaling up the model size and training data, researchers showed it could translate languages, answer questions, and write coherent articles without task-specific fine-tuning.
Why it mattered. It proved that simply making language models larger unlocked emergent capabilities, setting off a race to build massive foundation models and establishing prompting as a new way to program AI.
arxiv.orgIt folded proteins
DeepMind's AlphaFold 2 achieved unprecedented accuracy at the CASP14 protein structure prediction competition, effectively solving a 50-year-old grand challenge in biology. The system used an attention-based neural network to predict the 3D shapes of proteins from their amino acid sequences with atomic-level precision.
Why it mattered. It demonstrated that deep learning could solve fundamental scientific problems, accelerating biological research by providing structural data that would have taken decades to map experimentally.
deepmind.googleIt learned to draw
OpenAI introduced DALL-E, a 12-billion parameter version of GPT-3 trained to generate images from text descriptions. It could combine disparate concepts, attributes, and styles to create surreal but coherent pictures, like an 'illustration of a baby daikon radish in a tutu walking a dog'.
Why it mattered. It was the first public demonstration of high-quality, open-domain text-to-image generation, proving that language models could bridge the gap between text and vision.
openai.comIt learned to code
OpenAI released Codex, a GPT model fine-tuned on publicly available code from GitHub, which powered the newly launched GitHub Copilot. Evaluated on a new benchmark called HumanEval, Codex demonstrated the ability to synthesize functional Python programs directly from natural language docstrings.
Why it mattered. It transformed programming into a collaborative process with AI, becoming the first generative AI tool to achieve widespread daily use among software developers.
arxiv.orgIt learned to think aloud
Google researchers published 'Chain-of-Thought Prompting Elicits Reasoning in Large Language Models', showing that asking models to generate intermediate reasoning steps significantly improved their performance on complex tasks. By simply providing a few examples of step-by-step logic, models could suddenly solve math word problems and logic puzzles.
Why it mattered. It revealed that language models possessed latent reasoning capabilities that could be unlocked through prompting, changing how developers interacted with and evaluated AI.
arxiv.orgIt learned to chat
OpenAI launched ChatGPT, a conversational interface for a fine-tuned GPT-3.5 model optimized for dialogue using reinforcement learning from human feedback (RLHF). It could answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests.
Why it mattered. It became the fastest-growing consumer application in history, bringing generative AI to the general public and triggering an industry-wide race to deploy conversational agents.
openai.comIt learned to use tools
Meta researchers introduced Toolformer, a language model trained to decide which APIs to call, when to call them, and how to incorporate the results into its text. It taught itself to use calculators, search engines, and translation systems in a self-supervised way.
Why it mattered. It broke language models out of their text-only isolation, paving the way for AI agents that could take actions in the real world and fetch up-to-date information.
arxiv.orgIt took the bar exam
OpenAI released GPT-4, a large-scale multimodal model capable of accepting both image and text inputs. It exhibited human-level performance on various professional and academic benchmarks, including passing a simulated Uniform Bar Examination with a score in the top 10% of test takers.
Why it mattered. It established a new state-of-the-art for AI capabilities, proving that scaling up models with multimodal training could yield reliable systems for complex professional tasks.
openai.comFeed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.