Opening
Prologue: The Older Intelligence
There is a thing that happens before thought. Before language. Before the word arrives to name what the body has already recognized.
A cook walks into a kitchen at six in the morning and knows, before anyone speaks, that the day is going to break wrong. The walk-in is one degree warm. The prep list is two items short. The dishwasher has called out and nobody has said so yet, but the silence where his voice should be is already shaping the day. The cook has not been told any of this. The cook knows.
A doctor walks into a room and knows, before the chart is open, that the patient is not going to recover the way the labs suggest. Something in the breath. Something in the eyes. Something the labs cannot see.
A captain looks at a sea that is technically calm and orders the boat home.
The cook. The doctor. The captain. Same skill. They see problems before they happen. They work from the future and implement before the present.
This is intelligence. Not the part that wins games or solves equations. The older part. Pattern recognition under load. Threshold detection in a field of noise. The conscious mind arrives later and narrates what the body already decided. The cook says "I just had a feeling." The feeling was the whole room read at once, below the level of speech.
Knowledge is a different thing. Knowledge is what intelligence produces when it is given time and a place to write things down. Intelligence sees the pattern. Knowledge is the record after the pattern resolves. The chart. The recipe. The playbook. Knowledge compounds. It can be taught. It can be lost.
The two are partners, and the partnership is the species. The fire one cook learned to keep alive becomes the hearth the village builds around. The hearth becomes the kiln. The kiln becomes the furnace. The furnace becomes the foundry. Each generation inherits what the last one learned and reads new patterns against it. The compounding is the civilization.
At certain moments the compounding accelerates. Writing was one. A worker in one generation could suddenly read workers a thousand years dead. The printing press was another. The scientific method was another. It gave us a discipline for telling patterns that hold from patterns that only seemed to hold.
We are standing inside the next one. For the first time in the species' history, intelligence is not running only in human bodies. It runs in machines we built and trained on our own accumulated knowledge.
This module is the story of how that happened. Real names. Real dates. Real receipts. It's a shorter story than you think. And stranger.
Section I
Thought Becomes Machinery

George Boole, 1815 to 1864. A shoemaker's son with no degree who wrote that the operations of human reasoning could be done as algebra. In 1854 it looked like philosophy. It was a parts list.
Public domain · Wikimedia Commons
The machines did not start with machines. They started with a claim about logic that sounded, at the time, like philosophy.
George Boole. Son of a shoemaker. No university degree. Self-taught, and taught so well that Queen's College, Cork made him a professor anyway. In 1854 he published An Investigation of the Laws of Thought, and the claim in the title was the claim in the book: the operations of human reasoning could be written as algebra. Statements are true or false. One or zero. Combine them with AND, OR, NOT, and you can calculate with thought the way you calculate with numbers. In 1854 this looked like an eccentric's philosophy. It turned out to be a parts list.
Gottlob Frege, 1879, built the rest of the toolkit. His Begriffsschrift gave logic a full working notation: predicates, quantifiers, "all" and "some" rendered as machinery you could crank. Modern logic starts on that page. Almost nobody read it at the time.
By the 1920s David Hilbert, the most powerful mathematician in Europe, gathered all of this into a program with three demands. Prove mathematics is complete: every true statement provable. Prove it is consistent: no contradictions derivable. Prove it is decidable: a mechanical procedure that can settle any mathematical statement by rote. The third demand had a formal name, posed in 1928. The Entscheidungsproblem. The decision problem. Hilbert expected yes on all three. He had his confidence carved on his tombstone: "We must know. We will know."
He said those words in a radio address at Königsberg in September 1930. One day earlier, at a conference in the same city, a 24-year-old Viennese logician named Kurt Gödel had quietly announced a result that killed the first demand on the spot. Any consistent formal system rich enough to contain arithmetic contains true statements it cannot prove. Within months, his full paper killed the second as well: no such system can prove its own consistency from inside. Completeness: dead. Provable consistency: dead. The timing was not cruelty. It was just the field moving faster than its founder. Hold that image. It recurs.

Alan Turing at Princeton, 1936, the year he described a machine that did not exist yet, precisely enough that it could be built. He answered Hilbert with a pencil.
Public domain · Wikimedia Commons
The third demand, decidability, stood for six more years.
In 1936, before there was a computer, before the word computer meant a machine and not a person, a twenty-four-year-old mathematician at Cambridge published a paper that should not have been possible to write. The paper was called On Computable Numbers, with an Application to the Entscheidungsproblem. The title was forbidding. The paper was not. What Alan Turing did, in thirty-six pages of mathematics that almost no one alive at the time could fully follow, was describe a machine that did not exist. He described it precisely enough that it could be built. He described what it could do. He described what it could not do. He did this without a circuit, without a transistor, without an electrical engineer in the room. He did it with a pencil.
The machine he described came to be called the Turing Machine. It was not a physical thing. It was an idea. The idea. Write the thinking down as unambiguous steps. Any thinking that can be written that way, a dumb apparatus reading and writing symbols on a tape can perform. That is computing. All of it. The mathematics was austere. The implication was civilizational. Turing had defined what it means to compute. Which defined what a computer could one day be. And he had answered Hilbert. The answer was no. There is no mechanical procedure that settles every mathematical question, and Turing could point at a specific hole: no machine can determine, in general, whether another machine will ever finish running.
He was not yet thirty when he did this. Alonzo Church at Princeton reached the same verdict months earlier by a different road, a formalism called the lambda calculus. Church got the priority. Turing got the century. His version was the one you could build.
Three years later the war came. Turing reported to Bletchley Park the day after Britain declared. He sat in a hut and broke the German Enigma traffic with mathematics and a machine he helped design called the Bombe. An electromechanical device the size of three wardrobes. It weighed a ton. It ran around the clock, working through the rotor settings implied by a guessed scrap of plaintext until the day's German naval configuration fell out. Harry Hinsley, the official historian of British intelligence, put the resulting shortening of the war at not less than two years. Some estimates run to four years and fourteen million lives. Estimates, and labeled as such. The order of magnitude is not seriously disputed.

The Bombe. Three wardrobes of electromechanical logic that broke the German naval Enigma. The official history put the war shortened by not less than two years.
Public domain · Wikimedia Commons
Turing did not call any of this artificial intelligence. The vocabulary did not exist yet. But the people who watched him work knew what they were watching.
He survived the war. He did not survive the peace. In 1952 Britain prosecuted him for being homosexual and offered him a choice: prison, or chemical castration. He took the injections. He died in 1954 of cyanide poisoning, ruled a suicide. He was forty-one. A year before the proposal that would name the field he founded. Britain pardoned him in 2013. Fifty-nine years late.
In 1950, four years before he died, he wrote one more paper. Computing Machinery and Intelligence. It opened with a question. Can machines think? Then it dismissed its own question as ill-posed and replaced it with a better one. Can a machine, in conversation, persuade a human that it is human? The test he proposed has been called the Turing Test ever since. It has been argued about for three quarters of a century. Every time a new model ships, someone runs a version of the test, someone declares it passed, and someone else declares the test was never the right test. The argument is the point. Turing knew the argument was the point. He defined the question everyone would still be asking three generations after his death.
“The compounding is the civilization.”
Section II
The Machine and the Message

John von Neumann. His 1945 draft described a computer that stored its instructions in the same memory as its data. Every machine you have ever touched is built on it.
Public domain, US Government work · Los Alamos
While Turing was working in mathematics that did not yet have a machine, John von Neumann was working in mathematics that already had one. Hungarian-born. American-naturalized. By the 1940s he was the most respected mind in Western science. Game theory. Quantum mechanics. Shock-wave physics for the Manhattan Project. Weather prediction. He worked, when he could, on whatever unsolved problem he found interesting enough to walk into.
In 1945, while the war was ending, von Neumann wrote a document called First Draft of a Report on the EDVAC. It was a draft. It was never formally published. It circulated among a small group of engineers building the successors to the wartime ENIAC. The draft described a computer that, instead of being rewired by hand for each new problem, would store its instructions in the same memory as its data. Instructions could be loaded, modified, replaced. The machine could change what it was doing without being physically rebuilt. The typescript ran 101 pages. One honest asterisk belongs here: the stored-program idea was in the room before the draft was, and the ENIAC engineers J. Presper Eckert and John Mauchly spent the rest of their lives pointing out that the credit landed on the one name the document happened to carry. The dispute is real. The architecture is still called von Neumann's.
That architecture is the one every computer you have ever touched is built on. Every laptop. Every phone. Every server in every data center on every continent. The chip in your refrigerator. The chip in the satellite passing overhead. The foundation of the entire digital age was sketched in a wartime draft by a man who did not live to see what it became. He died in 1957, of cancer, at fifty-three.
In his last years he took up the question Turing did not live to reach. If a machine can compute, can a machine grow? Reproduce? Evolve? He worked on it until the cancer made work impossible, and left notes describing a theory of self-reproducing automata, published posthumously in 1966. Half the later century's thinking about artificial life and machine learning sits on those notes.
Turing asked whether a machine could think. Von Neumann asked whether a machine could become. They overlapped at Princeton for two years before the war. Von Neumann tried to hire him as an assistant; Turing went back to Cambridge instead. They never collaborated. They respected each other. Two men, working in separate hemispheres of the same problem, laying the foundation the next eighty years would build on.
Turing defined what computation was. Von Neumann defined what it would run on. Claude Shannon defined what it was for.

ENIAC. The room-sized machine the stored-program idea was drawn to replace. Rewired by hand for every new problem, until it did not have to be.
Public domain · Wikimedia Commons
Shannon had already made one legendary contribution before turning thirty. His 1937 MIT master's thesis showed that Boole's eighty-year-old algebra of thought mapped exactly onto electrical switching circuits. AND, OR, NOT, rendered in relays. It has been called the most important master's thesis of the twentieth century, because it is the reason logic could become hardware at all. Boole wrote the parts list. Shannon read it. Shannon was twenty-one.
Then, in 1948, at Bell Labs, he published A Mathematical Theory of Communication, in two installments, and invented information theory. The paper defined what information is. Not what a message says. What it measures. It gave us a mathematics for signal, noise, redundancy, and channel capacity that every form of digital communication has run on since. Every text you have sent. Every video call. Every encrypted transaction. Every satellite uplink. One paper. He was thirty-two.
He also wrote the first serious paper on programming a computer to play chess, built a mechanical mouse named Theseus that solved mazes, and juggled while riding a unicycle through the halls of Bell Labs. During the war, Turing spent two months at Bell Labs, and the two took tea together most days, forbidden by classification from discussing the cryptographic work both were doing. So they talked about whether machines could think. By every account, Shannon was the most playful serious mind of his generation. He treated mathematics the way a great cook treats a kitchen. The work and the joy were the same activity.
Section III
The Room at Dartmouth
Dartmouth College. In the summer of 1956, a proposal here coined the term artificial intelligence and named a field that did not exist yet.
Public domain · Wikimedia Commons
In the summer of 1956, John McCarthy, a young mathematician at Dartmouth College, convened the workshop he had proposed the year before. The proposal was co-signed by Marvin Minsky, Claude Shannon, and Nathaniel Rochester of IBM. It asked the Rockefeller Foundation for $13,500 and contained a sentence that named a field which did not yet exist:
"We propose that a 2 month, 10 man study of artificial intelligence be carried out during the summer of 1956 at Dartmouth College in Hanover, New Hampshire."
That sentence is the founding document. McCarthy coined the term artificial intelligence in that proposal partly to distinguish the new work from cybernetics, which was Norbert Wiener's field, and McCarthy did not want to spend the summer arguing with Norbert Wiener about whose field the work belonged to. The naming was political as much as conceptual. He named the field so the field would have a room of its own.
The proposal also carried the era's confidence in one line: every aspect of learning, or any other feature of intelligence, could in principle be described so precisely that a machine could be made to simulate it. In principle. That phrase would spend the next seventy years doing heavy lifting.
Rockefeller gave them $7,500. About half the ask. The "10 man study" turned out to be a rotating cast that arrived, argued, and left all summer, rarely all in the same room at once. They wrote some programs. They got very little done, by their own admission. McCarthy said later that the goal, significant progress on machine reasoning in two months, had been wildly optimistic. They had not understood how hard the problem was. Nobody had. The next seventy years were everyone re-learning it. Repeatedly.
But they had named the field, and they had assembled the people who would carry it for two decades. McCarthy went on to invent the LISP programming language in 1958. It is still in use. Minsky had already built the SNARC in 1951, a vacuum-tube contraption that was one of the first machines to learn anything at all, and he would spend a career arguing about whether neural networks or symbolic logic was the right road. Rochester had built the IBM 701, the company's first commercial scientific computer. Shannon came, contributed, and went back to Bell Labs. For him, Dartmouth was a side trip.
“Eleven pages.”
Section IV
The Winters
What the Dartmouth generation did not know was that the field they had just named was entering a nearly forty-year cycle of manias and collapses. The collapses came to be called the AI winters.
The pattern, both times, was identical. Demonstrations on toy problems. Promises scaled to the demonstrations plus imagination. Funding scaled to the promises. Reality arriving on schedule. Funding leaving on schedule.
Round one was symbolic. McCarthy, Minsky, Allen Newell, Herbert Simon, and their students built systems that manipulated logical symbols, with Arthur Samuel's checkers player at IBM running alongside. The systems solved algebra word problems. They played strong checkers. They proved theorems in propositional logic. Real results. Then the promises detached from the results. Herbert Simon, 1957: within ten years a computer would be world chess champion. It took forty. Simon again, 1965: machines would be capable, within twenty years, of doing any work a man can do. Minsky told Life magazine in 1970 that a machine with the general intelligence of an average human being was three to eight years out.
The press did its part. In 1958 Frank Rosenblatt demonstrated the perceptron, a machine that learned to recognize simple patterns by adjusting its own connection weights. Learning. Actual learning, in 1958. The demo ran as a simulation on an IBM 704. The custom Mark I hardware followed in 1960. The New York Times relayed the Navy's expectation of a device that would "walk, talk, see, write, reproduce itself and be conscious of its existence." The perceptron could, at the time, sort shapes.
Then the wall. Machine translation failed in public; a 1966 government report called ALPAC concluded it was slower, worse, and more expensive than human translators, and that funding died. In 1969 Minsky and Seymour Papert published Perceptrons, a book with a proof in it: the single-layer networks of the era could not represent something as basic as "exclusive or." The book was right about the narrow claim. It was read as a verdict on neural networks in general. Funding for the entire approach collapsed for fifteen years. Rosenblatt died in a boating accident in 1971, on his forty-third birthday, decades short of his vindication.
In 1973 the British government asked the mathematician James Lighthill to survey the field. His report was a demolition. In every category he examined, AI had failed to deliver what it promised, and its methods broke outside toy problems. British AI funding was cut to a handful of universities. The American agencies drew the same conclusion on their own timeline. The first winter ran from roughly 1974 to 1980. People left the field. The ones who stayed worked in obscurity, often without funding, often without students, holding the knowledge alive.
Round two was expert systems. The idea: capture a human expert's knowledge, a doctor's, an engineer's, a chemist's, as a structured set of if-then rules, and deploy it at scale. It worked. Commercially, at first, it genuinely worked. Digital Equipment Corporation's XCON configured computer orders and the company credited it with saving tens of millions of dollars a year. By the mid-1980s a billion-dollar industry existed to build these systems and the specialized hardware they ran on. Japan launched its Fifth Generation project in 1982 and put roughly $400 million behind a national push for intelligent machines.
Then the same wall, wearing new paint. The systems were brittle. They could not learn. They failed in ways their builders could not predict, and every rule had to be extracted from a human and maintained by hand, forever. Maintenance ate the margins. In 1987 the market for specialized AI hardware collapsed when ordinary workstations got cheap enough to run the same software. Symbolics, the flagship of that industry, held the first .com domain ever registered. March 15, 1985. The company cratered anyway. The Fifth Generation project wound down in 1992 with none of its headline goals met. The second winter ran from roughly 1987 to 1993, and by the end of it "artificial intelligence" was a phrase researchers scrubbed from their own grant applications.
Two winters. Same autopsy both times: the field kept promising general intelligence and shipping brittle special cases. Hold that autopsy. You will want it every time you read a press release.
Section V
The Ones Who Stayed

The perceptron beside the neuron it was drawn from. Rosenblatt's machine learned from examples in 1958. A 1969 book proved what one layer could not do, and the funding left for fifteen years.
Public domain · Wikimedia Commons
Through both winters, a small number of people kept working on the unfashionable thing.
The unfashionable thing was learning. The symbolic program assumed intelligence could be written down as rules. The alternative assumed intelligence had to be learned from data, by networks of simple units adjusting their connections. Rosenblatt's road, the one the 1969 book closed. Through the 1970s and 1980s, taking that road was a career liability. A few people took it anyway.
John Hopfield, a physicist, published a paper in 1982 that treated a neural network as a physical system with an energy landscape. That one move handed the field back the respectability it lost in 1969. Physicists started paying attention. The community the Perceptrons book had nearly killed re-formed around the paper. Hopfield was 91 when the Nobel Prize in Physics arrived in 2024. The paper it honored was 42 years old. That is what this field's timescale looks like from inside one career.
Geoffrey Hinton spent those same decades on one question: how do you train a network with many layers? A single layer was provably limited. Minsky and Papert had done the proving. But a deep network needed a way to work out which internal connection deserved the blame for which error. In 1986, David Rumelhart, Hinton, and Ronald Williams published the answer in Nature. Backpropagation. Pass the error backwards through the network, layer by layer, and adjust every weight in proportion to its share of the fault. The core idea had precursors reaching back to the early 1970s; the 1986 paper is the one that made it land. And then almost nothing happened for twenty-five years, because the computers of 1986 were too slow to train networks deep enough to matter. Hinton moved to Toronto and kept working. He trained the students who would build the 2010s. He shared the 2024 Nobel with Hopfield.
Yann LeCun, at Bell Labs in the late 1980s, built the first convolutional neural networks trained with backpropagation, wired with the same trick a visual cortex uses: small filters scanning across an image, reused everywhere. He aimed them at handwritten digits, ZIP codes supplied by the Postal Service, and by the late 1990s systems descended from that work were reading millions of checks deposited in the United States. Real deployment. Real money. The field's response was to keep ignoring neural networks for another decade. LeCun, Hinton, and Yoshua Bengio, who spent the same lean years on the same problems in a Montreal lab, shared the 2018 Turing Award, computing's highest honor, for the work everyone had spent twenty years walking past.
And Minsky? He lived to watch the networks he buried in 1969 return and take the field. He died in 2016, at 88, having acknowledged the return without exactly apologizing for the book.
“Capability was arriving as a side effect of scale.”
Section VI
The Bet on Data
What changed was not the algorithm. Backpropagation was 1986. Convolutional networks were 1989. What changed was that the two missing ingredients showed up. Data. And compute.
The data came from a bet almost nobody endorsed. In 2007 a young computer-vision professor named Fei-Fei Li looked at a field obsessed with cleverer algorithms and concluded it was starving for something else entirely: examples. Her lab set out to label a photograph of everything. The project was called ImageNet. Close to 50,000 workers in 167 countries, hired through Amazon's Mechanical Turk, hand-labeled millions of images. Senior colleagues advised her the project would sink her career. The dataset landed in 2009 to modest interest. In 2010 her team turned it into a public competition: here are 1.2 million labeled images across 1,000 categories; post the lowest error rate.
The 2010 winner posted 28.2% error, using conventional computer-vision methods. The 2011 winner posted 25.8%. A couple of points a year. That was the expected pace of the field.
Then 2012. Two of Hinton's students, Alex Krizhevsky and Ilya Sutskever, entered a deep convolutional network. 60 million parameters, trained for about a week on two consumer gaming GPUs. The gaming part matters. Those chips existed because teenagers wanted better explosions, and the same arithmetic that renders explosions turns out to train neural networks. The network, later called AlexNet, posted 15.3% error. The runner-up posted 26.2%. Competitions do not produce eleven-point gaps. Instruments get rechecked before fields accept a jump that size. Everyone checked. The instrument was fine.
That number ended the argument that had run since 1969. Within three years, every serious entry in the competition was a deep network. By 2015 the winning system posted 3.57%, past the measured human baseline of 5.1%. The field re-tooled around exactly the approach it had spent four decades declining to fund, and the people who carried that approach through the winters went from marginal to canonical in about 36 months. Hinton has been dry about this in public. He earned the dryness.
Section VII
Eleven Pages
In June 2017, eight researchers at Google published a paper called Attention Is All You Need. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, Illia Polosukhin. Gomez was an intern. The paper introduced an architecture called the Transformer. It ran eleven pages. It did not claim, in its own text, to be revolutionary. The authors thought they were proposing a better machine-translation model.
The idea. Stop reading one word at a time. Throw the recurrence out. Let every word in the sequence weigh every other word at once. They called it attention. It runs in parallel. Parallel means fast on GPUs. Fast on GPUs means you can build them big. Big, it turned out, means things nobody predicted.
Every frontier language model since runs on this architecture. GPT is a Transformer. Claude is a Transformer. The systems currently writing code, drafting contracts, and answering the world's questions run on the design in those eleven pages. The paper has been cited more than a hundred thousand times. All eight eventually left Google, mostly to found or join companies built on top of their own paper. Google has since paid billions to bring one of them back.
Eleven pages.
“That ratio is the whole story.”
Section VIII
The Loud Years
The Transformer shipped in June 2017. The public noticed in November 2022. The five years in between were the quiet years, and the quiet years are where the present was actually built.
The recipe was simple to state and expensive to run: take a Transformer, feed it a large slice of the written internet, and make it do one thing, predict the next word. Then make it bigger and do it again. In 2019, OpenAI's GPT-2, at 1.5 billion parameters, wrote paragraphs coherent enough that the company staged its release across nine months, citing misuse risk. The internet survived. In 2020, GPT-3 arrived at 175 billion parameters, a hundredfold jump in one generation, and something new showed up with it. Give it a few examples of a task inside the prompt, nearly any task, translation, arithmetic, bad poetry, and it would do the task. Nobody trained it for those tasks specifically. Capability was arriving as a side effect of scale.
The same year, researchers published the scaling laws: model performance improves as a smooth, predictable function of size, data, and compute. Draw the line, spend the money, collect the capability. The labs drew the line. And yet some abilities still arrived off-schedule, appearing abruptly at scales where smaller models had nothing. The field calls these emergent capabilities, which is a precise-sounding name for "we did not predict this and cannot fully explain it." Nobody has a complete read on why scale does what it does. That sentence is still true as this module goes to press.
On November 30, 2022, OpenAI put a chat interface on one of these models and released it as a research preview. ChatGPT reached one million users in five days. An estimated 100 million in two months, the fastest consumer-product adoption recorded to that point. The news cycle. The boardroom. The classroom. The group chat. All of it got, in one winter, the news the labs had been sitting on for five years.
The years since belong to the frontier labs: OpenAI, Anthropic, Google DeepMind, Meta, and a short list of others, shipping successively more capable models at a cadence no prior field has sustained. And coverage louder than the shipping. Every day another story about what GPT did or what Claude did. You already have the autopsy from the winters. Keep it in reach. The state of that race, who leads, what the models can actually do, and what they still cannot, is a moving target. It gets its own module, with a dateline on it.
Section IX
From Talking to Acting
One more shift. It is the one you are living inside.
A chat model talks. You type; it types back. Useful, novel, sometimes startling. Also, in the end, a surface. The model lives on one side of a text box and your actual work lives on the other, and everything crosses by copy and paste.
Starting in 2023 and accelerating hard through 2025, the models got hands. The technical word is tools. A model can now search the web, read your files, run code it just wrote, call other software, fill the form, file the ticket. Chain those steps and you have an agent: a system you hand a goal instead of a prompt. Find the bug and fix it. Reconcile these invoices. Watch this inbox and draft the replies. The model plans, acts, checks its own work, and comes back when it is done or stuck.
This is the moment the whole seventy-year story was pointing at, and it moves the question. Turing asked whether a machine could think. For working purposes, that question is retired. The working question is what the machine is allowed to do, and who signs off before it does it. A model that talks can embarrass you. A model that acts can spend money, email customers, delete a database. Every serious deployment of an agent is a design decision about where the human sits. What fires on its own. What waits for sign-off. What never fires at all. The industry is settling that question in real time, unevenly, with receipts accumulating on every side.
That present tense, what these systems can do this month, what they cost, where they fail, and how to read the announcements without getting played, is Module 2.
The history you just read runs 172 years, Boole to now. Module 2 covers about eighteen months. That ratio is the whole story.