IMITATION ENGINES

䷍ Powers of X

When I first encounted the Chinese Room as a boy, I imagined the rulebook would have to be extremely detailed, extremely special. Arcas thought so too.

Arcas🔗

We all thought that there would be a trick, and there wasn’t a trick. And because there was no discontinuity — it’s an exponent, it’s fast, but it’s also continuous — there’s no moment when it was clearly not intelligent before and clearly was after. And we also still, at some level, don’t know why scaling it up worked. …

And this was really a shock. It started to look like maybe the key to artificial general intelligence was really just scale. And I was quite snobbish about this idea that I’d heard in Silicon Valley that everything was about just scale and making stuff bigger. That just seemed incredibly naive. My training was in neuroscience and physics, and so the idea that just because we could make bigger computers and worship at the altar of Moore’s Law, that was going to solve all of the problems in science and technology, just seemed ridiculous. But the nerds were right.

A difference in degree can be a difference in kind. They’re called Large Language Models for a reason. We really have to understand how large.

Imagine young Searle “locked in a room and given a large batch of Chinese writing,” a whole bookcase, 10 million characters roughly. The manual is managable somewhere between 20 and 200 pages or maybe a whole shelf if we include the minority report. Yes, I’m asking the three sisters again, pooling their opinions. (The Gemi transcript gets spicy. I get angry when she keeps giving me too much info in answers. Her standing instructions have been updated.)

Searle is going to need to multiply a lot of numbers. It takes me two minutes to multiply something like -0.041923 x 0.087506. Yes, Mr. Actually, it took me two minutes ignoring the six least-significant digits. I’m slow. Working 996 with no other holidays, I can get 449,280 in a year. Let’s round that up to a million. This Searle is fast.

How much scrach paper is he going to need? Let’s start with Nanda’s model that on the way to learning addition mod 113 also learns some trigonometry. The training data is only few thousands sums: 20 pages of text. What about the ~200k weights? That would take about a thousand pages. But to train it, he’s going to need to keep rewriting. That’s too much scratch paper. We could give him a chalkboard; however, two perspectives (time taken and paper used) even though they directly convert are better than one.

Let’s suppose Searle is an artist, an especially special pointalist le Searle. Each weight is a lighter or darker dot, and he can tap out 25 in a square centimeter. For le Searle, the Nanda weight book is a tidy 20 pages. Suppose also that le Searle has excellent talent and vision. He can glance at two dots and tap their product in a fraction of a second: a billion FLOPs in year! Only a smidge slower than a TI-80.

To grok, Nanda’s model requires about 40k training epochs. And how many FLOPs in an epoch? ~250 million. Le Searle writes out 4 epic epochs in a year. It takes 10,000 years for le Searle as the “instantiation of the computer program” to learn how the clock goes round. After 10,000 years pecking out ten trillion dots, I expect Le Searle would have more colorful words than “simply” to describe this instantiation. Chaos + Time = words some would not repeat in front of the children.

Finally trained, how long does it take him to add two numbers? The whole afternoon: 4 hours. But the paper, I forgot the paper! Maybe 30 pages to do a sum. But, but, the training scratch paper? How much there? “Roughly one New York Public Library main branch” or 3,800 hexagonal galleries “with vast air shafts between, surrounded by very low railings” to ultimately add two numbers.

A difference in degree can be a difference in kind.

Let’s talk about GPT3.5 from the original ChatGPT, the same ChatGPT who can’t tell when you accidentally hit return without asking a question, who can’t tell a riddle is a riddle, and who can’t recognize a Voigt-Kampff test when they see it. Here we go GPT3.5!

Let’s send Le Searle 3.5 the message “hello, world”. It takes him 7,000 years to reply. He uses 2,700 galleries worth of scratch paper. He used a NYPL worth of books to train up, and he keeps his GPT3.5 model across 67 hexagons, 1/50 of a NYPL. But to train though? Trillions of years! Billions of libraries!

A difference in degree can be a difference in kind.

Let’s talk about the latest models from Anthropic (Claude Opus 4.6) and OpenAI (GPT-5.2 Extended) — sorry Gemi, you’re in the penalty box. At Le Searle scale, “it’s all just numbers now, numbers beyond human comprehension.” So let’s compare apples to oranges: they’re both fruit. This will work! 👍😉👍 What if Le Searle simulated a six year-old child? More sensory experience. A lot more brain connectivity. Hard to say what it all does, but they guess the brain is more efficient with only ~3% of the “weights” “active” at a time. Even when “inactive” I’ll assume the brain is learning. Behold a dubious chart!

chart comparing LLM sizes, details are made up

Gradients always descend. Lottery-tickets pay out. “It should be a log plot,” my father hollers from the future. Fine.

log chart comparing LLM sizes, details are still made up

Happy? The numbers have been made up, but the problems are real. Who knows Chinese here? The training corpus? Plain text, raw syntax. The rulebook? Trivial. The weights? Random. After time’s tuning? Who knows? That’s what the janitor at IHMC always told me. Are you watching now?

Michael Levin, developmental biologist, has an idea.

Levin🔗

An example you might think about is the kind of bell curve that you get from dropping marbles into a Galton board. Where does the shape of that bell curve come from? In a certain sense, at a micro level, you could probably calculate it out and predict, maybe if the errors were small enough. But in a larger sense, that shape is not encoded anywhere in the material of the marbles or the definition of the board or anything like that. It’s because what that device is doing is harnessing particular laws of physics. And what embryos do and all living structures do is they harness the laws of physics and the laws of computation.

Shannon, you can finally share that paragraph.

Shannon

The Platonic Representation Hypothesis (2024): This is the big theoretical development. Huh et al. argued that as vision models and language models get larger, they measure distance between datapoints in a more and more alike way — converging toward what they call a “shared statistical model of reality.” The provocative implication: if different AI architectures trained on different modalities converge on similar representations, and those representations align with brains, maybe there’s something like a unique optimal way to represent the world’s statistical structure that both evolution and gradient descent discover.

Convergence.

Levin🔗

I then found a couple papers in the machine learning community called the Platonic Representation Hypothesis. And I said, that’s pretty cool. These guys are climbing up to the same point where I’m getting at from biology and philosophy. They’re getting there from computer science and machine learning.

I’m using the name “Platonic” because I want it to be clear about a particular connection to mathematics. A lot of mathematicians would call themselves Platonists because what they think they’re doing is discovering — not inventing as a human construction — but discovering a structured, ordered space of truths.

In biology, as in physics, something very curious happens if you keep asking why. Cicadas come out at 13 years and 17 years. If you’re a biologist, you say, why is that? And you get the explanation that they’re trying to be off-cycle from their predators. You say, okay, cool, what’s special about 13 and 17? Oh, they’re prime. And why are they prime? Well, now you’re in the math department. You’re no longer in the biology department. You’re no longer in the physics department.

Another example — every time you talk to a physicist and you say, hey, why do the leptons do this or the fermions do that? Eventually the answer is, oh, because there’s this mathematical SU(8) group or whatever, and it has certain symmetries and structures. Once again, you’re in the math department.

There are facts you come across. Many of them are very surprising. You don’t get to design them. You get more out than you put in, because you make very minimal assumptions, and then certain facts are thrust upon you — the value of Feigenbaum’s constant, the value of E. These things you sort of discover. And the salient fact is this: if those facts were different, then biology and physics would be different. They impact the physical world. If the distribution of primes was something else, then cicadas would have been coming out at different times. But the reverse isn’t true. There is nothing you can do in the physical world to change E or Feigenbaum’s constant. You could have swapped out all the constants at the Big Bang. You are not going to change those things.

I think Plato and Pythagoras understood very clearly that there is a set of truths which impact the physical world, but they themselves are not defined by and determined by what happens in the physical world. You can’t change them by things you do in the physical world. I’ll make a couple of claims about that. One claim is, I think we call physics those things that are constrained by those patterns — when you say, hey, why is this the way it is, it’s because this is how symmetries or topology or whatever work. Biology are the things that are enabled by those. They’re free lunches. Biology exploits these kinds of truths. And really, it enables biology and evolution to do amazing things without having to pay for it. I think there’s a lot of free lunches going on here.

Yummy triangle slices! You have to work and wait, but the fruit of knowledge is free.

IMITATION ENGINES