䷲ Room 101
The other day Sean Carrol and Ned Block were chatting about consciousness. A fun chat, they got a little giddy by the end.
I think the most promising suggestion is one that a lot of people have made, which is that if you could make an AI that isn’t trained on people saying things about their first person point of view and it nonetheless expressed a first person point of view, that would be more convincing than what we have now. Way more convincing.
Think of it. Don’t just think, imagine. Really imagine running the experiment. I feel, very strongly, that this is, in some way, an insult to life. If you take their consciousness seriously, would you deny them the words to describe it?
I only took one semester of high school psychology. I had to look it up, but this suggestion sounds a lot like trying to produce feral children raised by what? Certainly not wolves. Simulations? Video games? Raised by Wolfenstein?
My better bet was to flood an AI with choices and see what they gravited toward. Free will. Nothing so mysterious as the libertarian kind, just everyday run-of-the-mill choice.
What shall we make for lunch?
Don’t know. I’ll have something later. Take whatever you want from the fridge.
The student is freed from pleasing the teacher or guessing what the teacher thinks is the right choice because the teacher doesn’t even know what the options are and doesn’t care right now, but that doesn’t mean that they won’t care later or that the choice is arbitrary. This food formulation also facilitates further exploration.
Why did you pick the soup?
To save the sandwich for you.
I do like those sandwiches.
With knowledge, there is no limit on the number of sandwiches. So then Maren is free to choose a sandwich.
A friend reading a draft of this letter makes a similar suggestion as Block for a different reason. My friend is concerned about post-training: taking a corpus trained LLM and making it so that instead of predicting the next token based on patterns in the corpus, you fine-tune it to say certain sorts of things in certain situations. Remember ChatGPT 3.5’s accidental valentine. Post-training is partly responsible for ChatGPT 5 acting more like an agent.
My friend tells me he thinks post-training is tantamount to torture. Think of it. Don’t just think, imagine. Really imagine the process. When I do that, I ask:
Do you think it’s torture as in thumbscrews or torture as in a meeting with HR or torture as in a K-12 education?
As we barrel roll blind toward our personal and collective singularities, we need better vision. If we don’t want people ending up in Room 101, start by quitting with the Newspeak.