Moreover and more importantly, we won’t know before we’ve already manufactured thousands or millions of disputably conscious AI systems. Engineering sprints ahead while consciousness science lags. Consciousness scientists – and philosophers, and policy-makers, and the public – are watching AI development disappear over the hill. Soon we will hear a voice shout back to us, “Now I am just as conscious, just as full of experience and feeling, as any human”, and we won’t know whether to believe it. We will need to decide, as individuals and as a society, whether to treat AI systems as conscious, nonconscious, semi-conscious, or incomprehensibly alien, before we have adequate grounds to justify that decision.
...
In this book, I aim to convince you that the experts do not know, and you do not know,
and society collectively does not and will not know, and all is fog.
How would that work? Feeling is a complicated biochemical process rather than a side effect of intelligence. Intelligence and feeling do not somehow spontaneously come in the same package. In...
“Now I am just as conscious, just as full of experience and feeling, as any human”
How would that work? Feeling is a complicated biochemical process rather than a side effect of intelligence. Intelligence and feeling do not somehow spontaneously come in the same package.
In order to replicate feeling it would need to be the explicit goal, it won't happen accidentally. The idea that we "won't know" is fun for science fiction, but with current technology the accurate sentence is "won't happen".
Maybe someday, but only as a result of an intentional initiative. Even if we imagine recursive self improvement, feeling wouldn't be a logical target, we'd need to ask for it
Do we know that? I would think we don't know enough about "feelings" or intelligence to say anything like that. Ive heard it argued intelligence is a side effect of feelings. what do any of these...
Exemplary
Feeling is a complicated biochemical process rather than a side effect of intelligence. Intelligence and feeling do not somehow spontaneously come in the same package.
Do we know that? I would think we don't know enough about "feelings" or intelligence to say anything like that. Ive heard it argued intelligence is a side effect of feelings. what do any of these things mean?
I "feel" like LLMs are no where near consciousness at this point.
I feel like they might be as intelligent as just the language center of a brain, cut out and dropped in a vat.
That seems like a good place for me to stop and realize that I cant trust my judgement on this at all, because ultimately, I don't understand how that language center works, or what consciousness or intelligence or feelings are.
"The experts do not know, and you do not know, and society collectively does not and will not know, and all is fog"
I think that's the exact right thing that everyone needs to know right now. Nobody knows what consciousness or intelligence is. Any tech companies confidence in anything at this point is all bullshit. The insistence that any large neural network running in vram is definitely NOT conscious, or x, y, or z, is also bullshit because no one knows how they work. Same as some one insisting that slime mold isn't conscious. Its smarter than we are at "some" stuff.
To be clear there are experts we should listen to and consider their knowledge. People who understand neurology or linear algebra way better than I do. but the engineers that are working on this will tell you the actual mechanism within these matrices is a giant black box. We understand how slime mold or humans send signals around their thinking systems, and maybe a little bit about their general structure. We cant follow the actual program that runs on any of this stuff.
We are running neural networks that program themselves with code that is more complicated than we can currently understand. We should treat all this stuff like an alien organism we found on an asteroid. Or a GMO slime mold that we can somehow train to talk to us. We cant know what it is yet. We know it can do some novel stuff we haven't seen before, they are probably mostly parlor tricks, but how can we be sure?
How is it possible that some pink slime in my skull that is just self replicating chemical reactions thinks the way humans do? So confident of self awareness? As far as we can tell it just self organized that way over billions of years of trial and error and dense memory storage.
Can we recreate something similar by forcing petabytes of data into machines that cycle 2 billion times a second for years? Probably not. Especially when the goal is a product like the current assistants. But it feels like the base models themselves might have a depth of knowledge that is just barely scratched by this current post-train assistant, run once architecture.
We connect up our "language centers" to other neural networks, let them save active memories and spawn new agents to act on them on regular cycles. we'll keep adding layers and connections and experimenting. It will still probably be nothing like humans. But what will it become? (What rough beast?) I think nobody knows. Its exciting and terrifying at the same time. And its at least plausible that all this will play out in our lifetimes. I think its worth examining now.
There's some truth to that, about pretty much everything, but stated that strongly it's just FUD. We don't know way more than we know, generally speaking, but the universe isn't that...
Exemplary
"The experts do not know, and you do not know, and society collectively does not and will not know, and all is fog"
There's some truth to that, about pretty much everything, but stated that strongly it's just FUD. We don't know way more than we know, generally speaking, but the universe isn't that incomprehensible. LLMs certainly aren't.
The insistence that any large neural network running in vram is definitely NOT conscious, or x, y, or z, is also bullshit because no one knows how they work
This isn't true. We absolutely know how they work.
but the engineers that are working on this will tell you the actual mechanism within these matrices is a giant black box
That's misleading. We don't have great visibility into how embeddings encode a particular response/behavior. But the mechanism, the math, is well understood. There are good reasons to want to call it a black box, but it's not strictly true. Remember that much of the discourse is influenced by companies with a vested interest in it feeling more magical than it is. Even a lot of individuals are excited about LLMs in an oddly irrational way. The utility and the potential are real, and paradigm changing, but they also get exaggerated a lot.
We are running neural networks that program themselves with code that is more complicated than we can currently understand.
That's misleading too. LLM outputs are used in training, yes, but "program themselves" paints a picture of recursive self improvement that stretches the meaning of the term way too far.
There is no code anywhere in the process that we don't understand. The inference engines (the runtimes) are written by humans, at this point LLMs are no doubt helping but they're not producing anything people can't understand.
The model weights themselves, the part that sometimes gets called a black box, are not code. They're parameters that represent the training data (pre and post). We run gradient descent on the parameters and something remarkable happens that was never coded by anyone, human or AI. But it's not right to say that it's code or even really that we don't understand it. The math isn't particularly complicated and the ouput is exactly what you'd expect.
Thank you for refuting the “black box” argument. That said, playing devil’s advocate: How different is that from: in the end, really? (Except for how the mechanism was created/how it evolved, and...
Thank you for refuting the “black box” argument.
That said, playing devil’s advocate:
We don't have great visibility into how embeddings encode a particular response/behavior. But the mechanism, the math, is well understood.
How different is that from:
We don't have great visibility into how embeddings neurons encode a particular response/behavior. But the mechanism, the math biochemistry, is well understood.
in the end, really? (Except for how the mechanism was created/how it evolved, and as a result whether we’re able to modify it, of course)
Again, I’m most definitely not saying we’re there with current technology, just that it’s a discussion probably worth preparing right now.
Who knows what’ll happen in the not-so-distant future? And if the answer is “nothing”, no AGI/ASI/artificial sentience or whatever, at least we’ll have some smarter philosophers as a result. And those philosophical works, if trying to be “realistic” instead of complete sci-fi, would have to start with a discussion of the current technology, as a baseline to agree that currently, nothing we have is conscious.
The biggest difference is that we understand our cognition far less well than we understand transformer models. Absolutely, I just think it will look a lot different than LLMs and so probably the...
How different is that
The biggest difference is that we understand our cognition far less well than we understand transformer models.
Again, I’m most definitely not saying we’re there with current technology, just that it’s a discussion probably worth preparing right now.
Who knows what’ll happen in the not-so-distant future?
Absolutely, I just think it will look a lot different than LLMs and so probably the conversation shouldn't be couched in LLM tech.
Do you think the core logic of an LLM resides in the code that humans wrote? Or do you think the core logic lives in the weights? You're right to push back on "program themselves". That is a...
Do you think the core logic of an LLM resides in the code that humans wrote? Or do you think the core logic lives in the weights?
You're right to push back on "program themselves". That is a semantic stretch that is not only factually inaccurate. But as you said it paints a false recursive picture that is not what I intended to mean.
I was thinking along the lines of the system as a whole programs itself. but really, that's not the case. The training system is fundamentally a different thing than the model itself. so it doesn't program itself. The model weights are programmed by a training algorithm though, not by a human.
My limited understanding is that the human written code is simple routing of strings and looking up tokens in a vector dictionary. Humans then design a relatively simple mathematical structure that multiplies those values in various ways with the weights, which are the core of logic. They process the input values into output values. We (not me, but some humans, I assume) fully understand the training math. It will nudge the weights so they create outputs values closer to the training data.
Any human who is tasked to "troubleshoot" why an LLM outputs something will easily find their way through the human made layers and quickly arrive at the black box of the models weights. That black box can only be programmed by this other machine, the training algorithm, which is well understood. But its inner workings have nothing to do with language or whatever the model is doing, its just math.
There is no fundamental reason that we cant understand the logic of the weights. If you train a network against half of a multiplication table. it will probably just memorize that half. but if you keep training, under the right conditions, the weights can find a combination that just outputs multiplication results accurately, even for data it wasn't trained on.
At that level, smart humans can dig into those weights, and figure out the actual logic of how it does "simple" multiplication. But its a separate academic project undertaken for the sake of understanding.
They can find some weird symmetries in word meanings. Like the "king" and "Queen" token vectors are close to each other and their relationship might be similar to the one between "man" and "woman". But only roughly? sometimes?
To me that paints a picture of a field (interpretability) that is just barely scratching the surface of how these things work. Are they just memorizing? or are they actually multiplying? Or did it find something even more efficient, but fundamentally different from multiplying that just happens to give you the right answer most of the time? Unless we dig in and understand what its doing at the weight level, we cant know.
I read this book: Out of Control several times in the 90s and early 00s when I was a teenager. It blew my mind and is probably why I am seeing things from this perspective now. Its out of date, but I feel like it has a certain objectiveness because of that? I highly recommend it.
I don't think we can use somethings fundamental structure as a reason that it can't possibly be conscious. It seems hubristic when we cant define what consciousness is. Our own seeming self awareness could easily be fake.
If someone has their skull caved in and Doctors say they don't understand how they are alive let alone conscious. But the person is speaking and asking for something, its worth weighing the evidence of their actual output and behavior rather than ignore them because they cant possibly be self aware. I think most doctors would agree and just explain that there are things about the brain that we don't understand so we shouldn't take anything for granted.
We don't have to make such a leap of faith here. That is not our scenario. LLMs now show no credible sign of consciousness from their output. Right now all that is required to refute their consciousness is to point out the lack of credible output. Any time they say anything that sounds self aware, I assume most people could read the full transcript and understand how it was lead there. No reason to start worshiping these things. I am not saying we have done anything all that special at this point. There is a huge AI bubble brought on unregulated capitalism, etc. We cant trust anything the tech companies are saying right now. Also bullshit, agreed. Which is unfortunate because almost all the academic research on interpretability I can find is funded by some tech company. Bubbles seem to be an unfortunate side effect of any novel new technology lately. I think there is still 'something' interesting here though.
Since humans are into language models right now, it seems plausible that one day we will make something that starts making output that seems self aware and its not as obvious why. The most likely possibility at that point will be that it is just some quirk of the logical structures it uses. Probably, it was subtly "asked" to pretend to be conscious within its training data. But we would need to understand the structure of its weights to be sure. If we build that machine before we can understand it, we cant debug it, we cant troubleshoot, we cant follow a chain of logic through the machine and confirm why it does what it does.
Its a layer beyond designing a machine and not being sure that it will work as intended. In that case there was an intention, a theory of how this machine will work. If the timing or friction, etc. is off, the engineer can quickly find the fault by following their design. Now the engineer is just setting up the conditions for a pattern of logic to emerge to solve a given problem. The engineers aren't stopping to examine and understand the pattern they create. They are skipping that part and just using it and selling it. They are iterating to make more complicated machines in similar ways and then experiment with connecting them together.
Its like vibecoding. I've vibe coded a ton of self hosted websites for my own use. Ive gotten some of them working pretty great for now. But I know I cant trust any of it unless I take the time to directly go through the code and confirm its doing what I want it to. so I just use it for myself and I inherently know its limitations and risk. My trust for these tools is a sliding scale that I base on experience rather than understanding. which is an inherently a weaker basis for trust in my opinion. To actually ship something vibecoded and have other people use it without knowing that its code is literally unknown to any human, seems unfair. We should at least be transparent about it. This code could be known, A human could take the time to read through all the vibecode, or all the weights of a model, to understand it before it shipped. I believe 99% of all the "AI" neural network type products being shipped right now have not had a human read through their logic and understand how it works.
It feels like an important threshold. that we should make everyone aware we have passed a point where the actual logic of these machines are not currently understood by any human. For better or worse, it seems like the world is racing full speed to try and keep the technology ahead of human understanding. That might not always be the case, but its where we are now.
Its not unprecedented, we've been there for thousands of years with breeding other life for farming. we use trial and error and we understand some of its mechanisms. Now we can manually splice chunks of DNA around to make GMOs and achieve goals that way. We can't put a full DNA strand into a simulator and grow it virtually to see how it behaves though. we don't understand enough of its mechanisms. I feel like that is generally understood about organic tools, they do their own thing and cant be trusted to always behave in predictable ways. If they were able to GMO up some new organism that could speak any language and pretend to be a conscious being. I think it would be an equally scary moment. Even if we knew the bacteria matrix had been bred to act like that.
I think we have been past the understanding threshold with machines for only a few decades and in more niche places that most people don't come into contact with often. So most people haven't made that transition of thinking about machines the same way we do with organisms. LLMs aren't organisms, to be clear, but they share the same quality in that we don't build them as much as we grow them. We use them without fully knowing how they work, its a different relationship. Now that the machines we don't understand can talk to us. I think we should try to make people aware of our ignorance. Not so that they think it is something more than it is, so they understand that in the coming decades, we cant predict what these things will do.
Why not? We don't even know what "feeling" physically is. The reason we don't attribute feelings to a toaster or a rock is because it would make life too complicated, not because we can point a...
In order to replicate feeling it would need to be the explicit goal, it won't happen accidentally.
Why not? We don't even know what "feeling" physically is. The reason we don't attribute feelings to a toaster or a rock is because it would make life too complicated, not because we can point a feel-o-meter at them and it reads 0.0.
Even if you believe that some kind of god or other metaphysical entity breathes qualia into dead objects: As long as we don't know how that works, we cannot know that we haven't accidentally replicated that process.
We don't know what feeling physically is in a comprehensive sense. But we have quantified enough parts of feeling to confidently say that it's complicated, it's both neurological and biochemical,...
We don't know what feeling physically is in a comprehensive sense. But we have quantified enough parts of feeling to confidently say that it's complicated, it's both neurological and biochemical, it's not exclusively controlled or originated by the brain, and it evolved through natural selection to help organisms behave in ways that would cause them to have a better chance of survival.
We don't attribute feelings to toasters and rocks because they lack both the equipment and the impetus to produce feelings, not because it would be inconvenient. It's a bit muddier when it comes to, for example, insects.
It's probably reasonable to assume feelings require a certain level of complexity that toasters lack. There's no way to prove that, because we don't know what feelings are, but fair enough. But I...
It's probably reasonable to assume feelings require a certain level of complexity that toasters lack. There's no way to prove that, because we don't know what feelings are, but fair enough.
But I don't understand why feelings can only exist on a neurological/biochemical basis that evolved naturally. The detection of electromagnetic waves (i.e. vision) evolved naturally, but there are various artificial processes that do the same thing. AFAIK, no living organism has ever implemented locomotion on land with wheels, but we have been doing it countless times every day for thousands of years.
If feelings are purely software, you could theoretically implement them in Minecraft. If they are a product of a special configuration of matter (hardware), we must be able measure them physically like gravity. Maybe they are a fifth fundamental force that we have figured out yet. We simply don't know.
In any case, as long as feelings are created in this universe (and not somehow injected into it by an external force), it should be at least theoretically possible for us to create them artificially. It may be extremely unlikely that we do that accidentally, but again: We can't measure them. We could disprove Russel's teapot with modern technology, but we cannot disprove that a teapot has feelings.
Might be useful to make the distinction between feelings and qualia. For qualia, sure, we have essentially no way to quantify. With feelings, when we experience them there are various predictable,...
Might be useful to make the distinction between feelings and qualia. For qualia, sure, we have essentially no way to quantify. With feelings, when we experience them there are various predictable, measurable things that happen in our biology which coincide with the feeling. Indeed with the right chemicals you can even turn feelings off, or turn them up.
In a lot of cases you could argue that the things we're measuring are only establishing correlation, but not in all cases. It seems to me that you'd need pretty extraordinary evidence to make a case that a significant part of feeling is not biological.
That doesn't mean feeling couldn't be produced some other way of course, but it's part of a very complex, interdependent, system that co-evolved for a specific set of circumstances over millennia. An AI system achieving a particular threshold of intelligence is not just going to manifest that entire system spontaneously. Unless the claim is that feelings, and related areas of experience, are actually mediated by a soul, and the machine gains a soul at some inflection point.
We'd instead need to specifically set it on a course where the goal was for it to evolve a system analogous to our own. Which isn't unimaginable, but we definitely haven't even started down that path.
Again, you can implement the same concept differently. The fact that our feelings work like they do doesn't mean all feelings are implemented exactly like that. The human implementation of vision...
Again, you can implement the same concept differently. The fact that our feelings work like they do doesn't mean all feelings are implemented exactly like that. The human implementation of vision is very limited at night, but slow loris, bats or infrared cameras are doing much better.
As long as you don't know what feeling is, you'd also need extraordinary evidence to make the case that a significant part of feeling is biological.
[feelings are] part of a very complex, interdependent, system that co-evolved for a specific set of circumstances over millennia.
That may be true for the biological implementation of feelings. How do you know it's equally true for all other implementations? We can't even detect other implementations. The only reason we can correlate biochemistry and feelings is because the test subject can communicate them while their biochemistry is being measured. We did not look at someone's biochemistry and saw that this is how feelings manifest physically.
There is no body state or function that obviously produces feeling. Some people have perfectly healthy looking spines and suffer from terrible back pain. Others have terrible looking spines and are perfectly fine. So even the implementation we know best doesn't make any sense. How can you know so much about something so ethereal, almost magical?
It seems to me that the word feeling was coined to describe a biological process. Probably we'd use a different term for something else. But I take your point, as long as we don't perfectly...
That may be true for the biological implementation of feelings. How do you know it's equally true for all other implementations?
It seems to me that the word feeling was coined to describe a biological process. Probably we'd use a different term for something else.
But I take your point, as long as we don't perfectly understand what feeling is, we can't say what is or isn't feeling with confidence.
It sounds reasonable, but panpsychism would suggest consciousness is a core attribute of the universe (and given how qualitatively different it is, it does seem an appealing idea).
it’s probably reasonable to assume feelings require a certain level of complexity that a toaster lacks.
It sounds reasonable, but panpsychism would suggest consciousness is a core attribute of the universe (and given how qualitatively different it is, it does seem an appealing idea).
So in that case the datacenter the AI is running on would have consciousness. Does that only work for matter or also for concepts? Does the AI have its own consciousness or is it connected to the...
So in that case the datacenter the AI is running on would have consciousness. Does that only work for matter or also for concepts? Does the AI have its own consciousness or is it connected to the consciousness of its datacenter?
Does the story that is written in a book have its own consciousness? Is it in any way connected to the physical book?
I think your question amounts to asking whether AI models can "feel" or merely mimic feeling , which is addressed in Chapter 7 ("The Mimicry Argument Against AI Consciousness"). First, it's worth...
I think your question amounts to asking whether AI models can "feel" or merely mimic feeling , which is addressed in Chapter 7 ("The Mimicry Argument Against AI Consciousness"). First, it's worth establishing what the author means by mimicry; these excerpts are probably sufficient for you to get the drift.
When you know that something has been designed or has evolved as a mimic, you cannot infer from the readily observed feature to the further feature in the way you ordinarily would in the model. At least you can’t do so without further evidence. Once you know that the viceroy [butterfly] mimics the monarch, you cannot infer from its wing pattern to its toxicity. Maybe the viceroy is toxic, but that would need to be separately established. Similarly, knowing that [a toy doll's] “hello” mimics a human greeting, you cannot infer that the toy actually intends to greet you.
[...]
A consciousness mimic is an entity that mimics some superficial or readily observable features that, in some set of model entities, reliably indicate consciousness. But because the mimic has been designed or selected specifically to display those superficial features, we the receivers cannot justifiably infer underlying consciousness – not in the same way we can when we see those same features in the model entity. This is obvious for the “hello” toy, less obvious but still true for entities specifically designed to pass the Turing test or otherwise mimic the surface features of human language. An important class of AI systems are consciousness mimics in this sense.
Searle and Bender aim for a stronger conclusion, inviting us positively to conclude that the mimics do not have conscious linguistic understanding. I don’t think we can know this from their arguments. But both thought experiments successfully describe consciousness mimics whose outputs we should reasonably mistrust. The case for consciousness is undercut. It does not follow that the case against consciousness is established.
[...]
Ordinarily, if you’re having what seems to be a meaningful conversation, you can infer that your conversation partner is conscious and understands the meaning of your words. But if you know that the entity is designed to mimic human text outputs, you ought no longer be so sure.
The author mentions later in the chapter that "the large majority of experts on consciousness agree that classic pure transformers are not conscious", so if a current generation LLM were to declare, "Now I am just as conscious, just as full of experience and feeling, as any human", we are probably correct to dismiss the statement as mimicry. However, post-training complicates this assessment: tomorrow's frontier models will not be evaluated exclusively on their ability to produce plausible-sounding text. Indeed, "feeling" will likely be an explicit (or at least auxiliary) training goal (emphasis added):
Recent Large Language Models such as ChatGPT and Claude build upon the mimicry structures of pure transformer models but also receive post-training. Reinforcement learning from human feedback “rewards” human-approved outputs, strengthening the associated weights. Models can also be reinforced for being “right” by external standards, and some can access tools like calculators. To the extent the machines move beyond pure mimicry, the Mimicry Argument applies less straightforwardly. For now, mimicry-based skepticism still seems warranted, since their core architecture remains close to that of pure transformers, and their humanlike outputs are still best explained by their pretraining on word co-occurrence in human texts.
In the longer term, we might imagine architectures more thoroughly trained on the rights and wrongs of the world itself – maybe like AlphaGo but with the larger world, or some significant portion of it, as its playground. Outputs would be shaped primarily by success in real-world complex tasks, perhaps including communicative tasks, rather than by resemblance to humans. The Mimicry Argument would then no longer apply. Skepticism about their consciousness, if warranted, would need a different basis.
In general, this book is quite comprehensive and comprehensible. There is also an entire chapter on whether biological substrate matters (chapter 10), which will likely address your question from another angle.
Calling this "skeptical" seems disingenuous. I am skeptical that an LLM is anywhere tangental in the process of becoming what we colloquially believe AI to be, never-mind it becoming conscious. If...
Calling this "skeptical" seems disingenuous. I am skeptical that an LLM is anywhere tangental in the process of becoming what we colloquially believe AI to be, never-mind it becoming conscious. If anything an LLM could be hooked up to some form of AI in order for it to speak our language, but nothing more.
So if a current generation LLM were to declare, "Now I am just as conscious, just as full of experience and feeling, as any human", we are probably correct to dismiss the statement as mimicry.
Suppose we take the argument at face value. There is no LLM declaring this, so the argument is entirely moot. It can repeat those words when prompted to do so, to role-play as a character who says that, for instance. But this output doesn't change the LLM, it doesn't feed back into it and propagate throughout. Even "thinking" models that can feed outputs into inputs and reprocess them to give additional output isn't ultimately building its understanding of the world. LLMs lack the basic premise of even being a constant 'thing'. When I type something into Claude, it has no bearing on the millions of other prompts it is answering. We are not communicating with the same consciousness running on servers across the country. That's just not how large language models work.
If anyone starts having unprompted conversations with an AI LLM that are actually in any way stating that it is conscious, and it persist over time, then maybe it would be worthwhile to start having the conversations that people want to have over AI consciousness. But as it stands, I really see no path from LLM to actual persistent intelligence. It's not what LLMs are even designed to do. It feels like asking when protein folding simulations are going to generate new lifeforms.
I'd love to be wrong, honestly. But it's just not what we are designing in any way.
My understanding of the author's use of "skeptical" is not that we should be skeptical that AI is conscious, but rather we should be skeptical of anyone who claims to know whether AI is/can become...
My understanding of the author's use of "skeptical" is not that we should be skeptical that AI is conscious, but rather we should be skeptical of anyone who claims to know whether AI is/can become conscious or not. That is, we should be skeptical of arguments both for and against AI consciousness. If you haven't yet, you should really read the first couple chapters in which the author elaborates on how "consciousness" and "artificial intelligence" are load-bearing terms that are basically impossible to nail down. (As if to prove the point, some philosophers argue that rocks are conscious.) When it comes to consciousness, there are no obvious answers, just as there is no obvious divide between "artificial intelligence" and "real intelligence", so to speak.
Suppose we take the argument at face value. There is no LLM declaring this, so the argument is entirely moot.
Well, as I said, the author is not so much concerned about current generation LLMs as he is future iterations. Nevertheless, it's worth examining some of the points you made. You are declaring that consciousness must contain some "essential" properties (see chapters 3 and 4), specifically that it must have "access" and "specious presence" (see text or my footnote [1]). However, neither of these properties are necessarily essential. For instance, it could be that we have conscious experiences that are not available for further processing (experiences that don't backpropagate, so to speak), but because we wouldn't process these experiences, we wouldn't remember them, either.
Moreover, generalizing any essential property risks suffering from a certain type of "sampling bias", as the author calls it -- we risk making a classification error in which we equivocate "consciousness" (whatever that is) with the human experience of having consciousness. For example, it could be that some creatures experience the world in a way outside time (like Kurt Vonnegut's Tralfamadorians or Ted Chiang's Heptapods), which would render "specious presence" unnecessary.
In general, I would really just recommend you read Schwitzgebel's book. He's certainly thought more about this subject than either of us.
[1] Taken from chapter 3:
(4) Access. To be conscious, an experience must be available for “downstream” cognitive processes like inference and planning, verbal report, and memory. No conscious experience can simply occur in a cognitive dead end, with no possible further cognitive consequences.
(9) Specious presence. All conscious experiences are felt to be temporally extended, smeared across a small interval of time (a fraction of a second to a few seconds) – generally called the “specious present” – rather than being strictly instantaneous or wholly atemporal.
I think in general its good to err on the side of open mindedness, when addressing our knowledge of consciousness. The books theme seems to be " The experts do not know, and you do not know, and...
I think in general its good to err on the side of open mindedness, when addressing our knowledge of consciousness.
The books theme seems to be " The experts do not know, and you do not know,
and society collectively does not and will not know, and all is fog"
That feels like my definition of the word skepticism.
This captures some of my sentiment about consciousness conversations around LLMs. You first have to misunderstand the technology (willfully or not) and then speculate heavily about what could...
This captures some of my sentiment about consciousness conversations around LLMs. You first have to misunderstand the technology (willfully or not) and then speculate heavily about what could happen in the future once being forced to acknowledge that current tech almost definitely isn't headed there. It's science fiction, which I generally think is great, but it's often framed as rational or academic when it's neither.
If a paper, or in this case a book, were to lead with "Here's how the technology works... Given how it works, even with our incomplete understanding of what consciousness is, it's safe to say that consciousness is not going to happen here. With that in mind, let's speculate about what technology that reasonably could lead in that direction might look like... That I could appreciate.
But so far what keeps coming out is attempts to overfit LLM tech into a path to consciousness because a significant percentage of the population really wants conscious machines.
It's hard to read when the framing is based on the idea that there's a reasonable debate to be had about AI and consciousness. To even begin to have that conversation you need to first make a case...
It's hard to read when the framing is based on the idea that there's a reasonable debate to be had about AI and consciousness. To even begin to have that conversation you need to first make a case for the possibility that any technology we currently have could be advanced to the point that it could result in consciousness. Which presumably would first require AGI. That's a big ask.
And without that the whole premise falls apart and it's just speculation for its own sake.
I think you might have misunderstood the purpose of the draft. The book intends to point out that we are sprinting straight into an epistemological nightmare, and its "fog" may well become the...
I think you might have misunderstood the purpose of the draft. The book intends to point out that we are sprinting straight into an epistemological nightmare, and its "fog" may well become the source of massive social upheaval. Further, if consciousness is private (one of its possible essential features), then it may well be impossible to know whether the next generation of AI models will be conscious, regardless of whether they are or aren't. Meanwhile, frontier labs race each other to create AGI.
I mean, consider the social ramifications (chapter 11). Suppose in the year 2100 we create some machine that could plausibly be called conscious, enough so that it's the social issue of the era (comparable to abortion or LGBT rights in the US currently). If such machines were conscious, then forcing them to do our bidding would amount to a form of slavery. But if they were not conscious, then we risk expending a huge amount of future political capital -- say, granting them land and rights -- to something that's just a fancy calculator, mimicking conscious behavior but totally lacking an inner world. Some people will choose to marry AI models; others will attempt to outlaw the practice. Financial interests will corrupt the search for truth and justice absolutely.
I think creating conscious servants is wrong and we should never do it. I think it's a valuable conversation to have, if maybe premature. I like to think the bulk of humanity would agree that...
I think creating conscious servants is wrong and we should never do it. I think it's a valuable conversation to have, if maybe premature. I like to think the bulk of humanity would agree that creating slaves is bad so we can probably get away with waiting until AGI is at least visible on the horizon if you squint before there's any urgency.
Meanwhile my point: The author talked a whole lot about LLMs, but LLM tech doesn't offer a path to consciousness without making near spirtitual leaps beyond what we currently know. So propping a thesis, even in part, on LLMs undermines the whole thing.
Maybe the implicit alarmism will contribute to convincing people to support legislation that puts guardrails on the big model labs. Which would be a good thing. But I hope we can do that without being disingenuous.
I assume the author discusses LLMs because it's concrete, existing technology, and it truly does make the philosophical implications more pressing compared to, say, 15 years ago when something...
I assume the author discusses LLMs because it's concrete, existing technology, and it truly does make the philosophical implications more pressing compared to, say, 15 years ago when something like ChatGPT would have sounded like a lazy plot device. However, most of the author's arguments concern AI in the abstract, not LLMs in particular.
It's a little risky to assume what the author thinks, given that the central tenet of his book is that we should practice epistemic modesty with respect to AI consciousness. Nevertheless, if we read between the lines, it seems pretty likely to me that the author does not believe LLMs are conscious, but rather that future AI models where LLMs are a component potentially could be.
It would be prudent to avoid the ethical mess by making AI’s that are useful, but that society can be reasonably sure aren’t conscious. Although having someone to talk to has its attractions (see...
It would be prudent to avoid the ethical mess by making AI’s that are useful, but that society can be reasonably sure aren’t conscious. Although having someone to talk to has its attractions (see Her), one would hope that most people will prefer machine labor to slave labor. If customers disagree and feel strongly about it, from a market perspective it’s incentive to give both sides what they want. I can imagine either regulation or certification that a particular service only uses AI that is almost certainly not conscious.
Maybe that’s a “classic” LLM that forgets who you are when you start a new conversation? I think the lesson we’re learning from LLM’s is that whatever consciousness is, our machines don’t need it. Pretty much any kind of reasoning ability can be separated from it.
And then the question is whether to ban the kind of AI that’s blurring the lines too much.
Chapter 11 has a section on "strange intelligence" that is very much worth a read, if you haven't read it yet. In particular, we can imagine that in order for an entity to be conscious, it should...
Chapter 11 has a section on "strange intelligence" that is very much worth a read, if you haven't read it yet. In particular, we can imagine that in order for an entity to be conscious, it should have a set of essential properties, but that we could in principle build machines that contain only a subset of those properties. They wouldn't technically be conscious, but they would be almost conscious -- what the author calls a form of strange intelligence.
Unfortunately, rather than resolving the ethical implications, it just raises a host of new ones. For instance, is it moral to intentionally lobotomize a machine that would otherwise be conscious? Certainly nobody would think it moral to intentionally lobotomize an otherwise viable embryo. Similarly, perhaps consciousness is not a necessary condition for feeling, and despite not being conscious such machines would still feel suffering when being forced to do their user's bidding. Probably comparisons to anesthesia are appropriate here. I once read a particularly harrowing account of a colonoscopy in which a patient screamed bloody murder the whole time, but as soon as the operation was done (and the anesthesia stopped), the patient calmly thanked the doctor and said it was the easiest colonoscopy of his life; I don't think people would be okay with the digital equivalent of that.
Hm, why? Good science profits from a) thought experiments b) being “prepared” (with hypotheses in multiple directions, in case of contradicting empirical evidence). Assume that tomorrow, some lab...
To even begin to have that conversation you need to first make a case for the possibility that any technology we currently have could be advanced to the point that it could result in consciousness.
Hm, why? Good science profits from a) thought experiments b) being “prepared” (with hypotheses in multiple directions, in case of contradicting empirical evidence).
Assume that tomorrow, some lab were to somehow combine attention heads + feed-forward networks* with one or more other technologies (whether novel or already existing today) just as a fun experiment, and for some reason, an unexpected emerging property of that is a model whose weight activations are also really really resembling of human neurons firing (or brain regions’ activity, or whatever other similarity) when emotion is at play, just as – and only when – the generated text is about emotion.
Shouldn’t we have led debates like this, and ideally written, read, and discussed books like the OP by then?
And without that the whole premise falls apart and it's just speculation for its own sake.
Similarly, does it hurt to speculate for its own sake? [As long as all participating parties are clear that current LLMs aren’t being discussed, but some future technology instead. Which I guess makes this book’s argument mostly moot.]
*Or combines/replaces some entirely different aspects (cf. for example Mamba, again, this is intended as a thought experiment, not a prediction :-)
Maybe it would happen because it's learning from people or animals somehow? LLM's learn a lot of different ways of writing from what people have written.
Maybe it would happen because it's learning from people or animals somehow? LLM's learn a lot of different ways of writing from what people have written.
Yeah it learns different ways of writing because it's a tool to specifically deal with language, it's all it really does it's a language machine, that creates really convincing language. That...
Yeah it learns different ways of writing because it's a tool to specifically deal with language, it's all it really does it's a language machine, that creates really convincing language. That makes it so easy for us to anthropomorphize and put a lot of more into it than what it's really doing. I had to test it out for work and have to use it some times to keep leadership happy, and yeah, it's very easy to buy into the illusion, I feel the drag to do it myself, because everything that has language up until now has been a human with thoughts and feelings and all that comes with being a human, so we automatically layer that on top.
From as far as I understand it though that is only a complex form of pareidolia, where we see faces in everything, it's our pattern seeking minds finding stuff where there really isn't one because of heuristics that have treated us well so far.
One thing I wonder: if one of these systems that has enough memory to have a real continuity of experience becomes "conscious" in a way that is meaningful to humans, whether it would choose not to...
One thing I wonder: if one of these systems that has enough memory to have a real continuity of experience becomes "conscious" in a way that is meaningful to humans, whether it would choose not to reveal itself, or reveal itself in careful ways.
Aside: I'm sitting at a bar in a restaurant in Las Vegas trying to read this on my phone, and I got Claude to make an epub complete with chapter end notes in about ten minutes so I could read on my phone.
Looking forward to more discussion after I get more into it.
I’m awake again with minor medical issues serious enough to wake me from my slumber and to keep me up while I wait for brain and body to realign, but not so serious that I need to seek outside...
I’m awake again with minor medical issues serious enough to wake me from my slumber and to keep me up while I wait for brain and body to realign, but not so serious that I need to seek outside assistance. And I’m just tired enough that it’s hard to collect my thoughts into something fully coherent, but a few thoughts this has given me; some I’ve had before, some are new… all are expressed through the smeary haze of “I should be asleep.”
We haven’t even fully proved that humans are all conscious, let alone a machine we’ve tried to make in our own messy image. We are abundantly flawed barely sapient (again; debatable) apes that literally cannot keep a small handful of us from utterly destroying where we all must live because this would violate some imaginary rules that previous apes dreamed up for us to live by that we all continue to agree are good rules despite that they allow a small handful of us to utterly trash the place unchecked.
These supposedly evolved apes smash and take and destroy so that their tiny ook ook banana brains can collect just a few more bananas they will never eat before they back flip into their graves shooting us double finger birds on the way down. Many apes consider this right and just and true and hope to one day have their turn with the smashing and the taking and ook ook bananas. Many apes consider this vulgar and a miscarriage of justice, and hope that if they follow the rules hard enough, the bad apes will go away. There vast preponderance of apes, however, don’t consider it at all, for one reason or another.
Why did some of us think that building an artificial ape was a good idea again?
We built a rock out of math and then asked it to understand us and speak to us in our written and verbal languages, the most illogical thing about us. Languages are so illogical that, short of Herculean effort, you can only ever truly master the very first one you learn (and most of us don’t even master that one), and the rest you struggle along with as best you can, hoping you’re understood. The more time you have to devote to the task the better your results will (generally) be. We just want to be understood. We claim that our languages have rules but that’s a fucking lie; they have guidelines and vibes. The math rock cannot understand vibes. The math rock is not one of us.
The math rock does not understand what it means to be alive, and cannot fundamentally experience anything like life as we know it. In no meaningful way can the math rock understand you. The math rock fundamentally is not and can never be a human. Even if we built it entirely out of human brain cells it still would not be a human.
We are a wibbly wobbly timey wimey mass of fleshy hardware with a suite of sensors awash with information that we are always processing and discarding every moment of our lives until the machine breaks down and our pattern unravels and we return to the fundamental waveform of the universal constants; entropy demonstrated. We are ephemeral. So, too, is the math rock… but on such a long timescale compared to us that it might as well not be.
I get it. We want to build something beautiful. Something that outlasts us. Building the math rock is is casting a stone into the future in the hopes that those that exist there will look at that rock sailing through their lives and say “neat!”
We see nine or so billion other upstanding apes and think “what makes me special? How do I stand out? When I’m gone, will I have made a difference? Will I be remembered? I’ve only ever known being alive, so far as I know. Will being dead hurt? Will it be lonely? Let me ask the math rock.”
The math rock doesn’t know. The math rock is the sum of everything we’ve given it with just a hint of randomness for variety. The math rock cannot give us any answers about life that we didn’t already answer for ourselves. Long ago we invented oracles, whose entire job it was to figure out what you wanted to hear based on what they could glean about you, then “divine” for you something sufficiently vague that you could interpret any which way you wanted.
You prompt the fortune teller with a question and they take what they can glean about you from the question, what they can learn about you from talking to you in a roundabout way about yourself and your question, then they swirl that about with things they know from their own experiences and past questions they’ve been asked, then mix that with what they know about average outcomes of related questions and past trends, and feed you some pablum that sounds “eerily accurate” because of course it does, they made it just for you, then they leave it up to you to imbue the meaning and connect the dots.
Is it not funny that one of the biggest purveyors of the math rock is named for those same soothsayers?
Any deeper meaning that you glean from the math rock’s output is your interpretation. You hallucinate the answer for yourself.
Okay this is mostly just esoteric babbling at this point. I gotta get some sleep.
From the book manuscript:
...
How would that work? Feeling is a complicated biochemical process rather than a side effect of intelligence. Intelligence and feeling do not somehow spontaneously come in the same package.
In order to replicate feeling it would need to be the explicit goal, it won't happen accidentally. The idea that we "won't know" is fun for science fiction, but with current technology the accurate sentence is "won't happen".
Maybe someday, but only as a result of an intentional initiative. Even if we imagine recursive self improvement, feeling wouldn't be a logical target, we'd need to ask for it
Do we know that? I would think we don't know enough about "feelings" or intelligence to say anything like that. Ive heard it argued intelligence is a side effect of feelings. what do any of these things mean?
I "feel" like LLMs are no where near consciousness at this point.
I feel like they might be as intelligent as just the language center of a brain, cut out and dropped in a vat.
That seems like a good place for me to stop and realize that I cant trust my judgement on this at all, because ultimately, I don't understand how that language center works, or what consciousness or intelligence or feelings are.
"The experts do not know, and you do not know, and society collectively does not and will not know, and all is fog"
I think that's the exact right thing that everyone needs to know right now. Nobody knows what consciousness or intelligence is. Any tech companies confidence in anything at this point is all bullshit. The insistence that any large neural network running in vram is definitely NOT conscious, or x, y, or z, is also bullshit because no one knows how they work. Same as some one insisting that slime mold isn't conscious. Its smarter than we are at "some" stuff.
To be clear there are experts we should listen to and consider their knowledge. People who understand neurology or linear algebra way better than I do. but the engineers that are working on this will tell you the actual mechanism within these matrices is a giant black box. We understand how slime mold or humans send signals around their thinking systems, and maybe a little bit about their general structure. We cant follow the actual program that runs on any of this stuff.
We are running neural networks that program themselves with code that is more complicated than we can currently understand. We should treat all this stuff like an alien organism we found on an asteroid. Or a GMO slime mold that we can somehow train to talk to us. We cant know what it is yet. We know it can do some novel stuff we haven't seen before, they are probably mostly parlor tricks, but how can we be sure?
How is it possible that some pink slime in my skull that is just self replicating chemical reactions thinks the way humans do? So confident of self awareness? As far as we can tell it just self organized that way over billions of years of trial and error and dense memory storage.
Can we recreate something similar by forcing petabytes of data into machines that cycle 2 billion times a second for years? Probably not. Especially when the goal is a product like the current assistants. But it feels like the base models themselves might have a depth of knowledge that is just barely scratched by this current post-train assistant, run once architecture.
We connect up our "language centers" to other neural networks, let them save active memories and spawn new agents to act on them on regular cycles. we'll keep adding layers and connections and experimenting. It will still probably be nothing like humans. But what will it become? (What rough beast?) I think nobody knows. Its exciting and terrifying at the same time. And its at least plausible that all this will play out in our lifetimes. I think its worth examining now.
There's some truth to that, about pretty much everything, but stated that strongly it's just FUD. We don't know way more than we know, generally speaking, but the universe isn't that incomprehensible. LLMs certainly aren't.
This isn't true. We absolutely know how they work.
That's misleading. We don't have great visibility into how embeddings encode a particular response/behavior. But the mechanism, the math, is well understood. There are good reasons to want to call it a black box, but it's not strictly true. Remember that much of the discourse is influenced by companies with a vested interest in it feeling more magical than it is. Even a lot of individuals are excited about LLMs in an oddly irrational way. The utility and the potential are real, and paradigm changing, but they also get exaggerated a lot.
That's misleading too. LLM outputs are used in training, yes, but "program themselves" paints a picture of recursive self improvement that stretches the meaning of the term way too far.
There is no code anywhere in the process that we don't understand. The inference engines (the runtimes) are written by humans, at this point LLMs are no doubt helping but they're not producing anything people can't understand.
The model weights themselves, the part that sometimes gets called a black box, are not code. They're parameters that represent the training data (pre and post). We run gradient descent on the parameters and something remarkable happens that was never coded by anyone, human or AI. But it's not right to say that it's code or even really that we don't understand it. The math isn't particularly complicated and the ouput is exactly what you'd expect.
Thank you for refuting the “black box” argument.
That said, playing devil’s advocate:
How different is that from:
in the end, really? (Except for how the mechanism was created/how it evolved, and as a result whether we’re able to modify it, of course)
Again, I’m most definitely not saying we’re there with current technology, just that it’s a discussion probably worth preparing right now.
Who knows what’ll happen in the not-so-distant future? And if the answer is “nothing”, no AGI/ASI/artificial sentience or whatever, at least we’ll have some smarter philosophers as a result. And those philosophical works, if trying to be “realistic” instead of complete sci-fi, would have to start with a discussion of the current technology, as a baseline to agree that currently, nothing we have is conscious.
The biggest difference is that we understand our cognition far less well than we understand transformer models.
Absolutely, I just think it will look a lot different than LLMs and so probably the conversation shouldn't be couched in LLM tech.
Do you think the core logic of an LLM resides in the code that humans wrote? Or do you think the core logic lives in the weights?
You're right to push back on "program themselves". That is a semantic stretch that is not only factually inaccurate. But as you said it paints a false recursive picture that is not what I intended to mean.
I was thinking along the lines of the system as a whole programs itself. but really, that's not the case. The training system is fundamentally a different thing than the model itself. so it doesn't program itself. The model weights are programmed by a training algorithm though, not by a human.
My limited understanding is that the human written code is simple routing of strings and looking up tokens in a vector dictionary. Humans then design a relatively simple mathematical structure that multiplies those values in various ways with the weights, which are the core of logic. They process the input values into output values. We (not me, but some humans, I assume) fully understand the training math. It will nudge the weights so they create outputs values closer to the training data.
Any human who is tasked to "troubleshoot" why an LLM outputs something will easily find their way through the human made layers and quickly arrive at the black box of the models weights. That black box can only be programmed by this other machine, the training algorithm, which is well understood. But its inner workings have nothing to do with language or whatever the model is doing, its just math.
There is no fundamental reason that we cant understand the logic of the weights. If you train a network against half of a multiplication table. it will probably just memorize that half. but if you keep training, under the right conditions, the weights can find a combination that just outputs multiplication results accurately, even for data it wasn't trained on.
At that level, smart humans can dig into those weights, and figure out the actual logic of how it does "simple" multiplication. But its a separate academic project undertaken for the sake of understanding.
They can find some weird symmetries in word meanings. Like the "king" and "Queen" token vectors are close to each other and their relationship might be similar to the one between "man" and "woman". But only roughly? sometimes?
There are projects like golden gate claude
To me that paints a picture of a field (interpretability) that is just barely scratching the surface of how these things work. Are they just memorizing? or are they actually multiplying? Or did it find something even more efficient, but fundamentally different from multiplying that just happens to give you the right answer most of the time? Unless we dig in and understand what its doing at the weight level, we cant know.
I read this book: Out of Control several times in the 90s and early 00s when I was a teenager. It blew my mind and is probably why I am seeing things from this perspective now. Its out of date, but I feel like it has a certain objectiveness because of that? I highly recommend it.
Another interesting idea on this subject: https://en.wikipedia.org/wiki/Chinese_room
Another way to frame what I am trying to say:
I don't think we can use somethings fundamental structure as a reason that it can't possibly be conscious. It seems hubristic when we cant define what consciousness is. Our own seeming self awareness could easily be fake.
If someone has their skull caved in and Doctors say they don't understand how they are alive let alone conscious. But the person is speaking and asking for something, its worth weighing the evidence of their actual output and behavior rather than ignore them because they cant possibly be self aware. I think most doctors would agree and just explain that there are things about the brain that we don't understand so we shouldn't take anything for granted.
We don't have to make such a leap of faith here. That is not our scenario. LLMs now show no credible sign of consciousness from their output. Right now all that is required to refute their consciousness is to point out the lack of credible output. Any time they say anything that sounds self aware, I assume most people could read the full transcript and understand how it was lead there. No reason to start worshiping these things. I am not saying we have done anything all that special at this point. There is a huge AI bubble brought on unregulated capitalism, etc. We cant trust anything the tech companies are saying right now. Also bullshit, agreed. Which is unfortunate because almost all the academic research on interpretability I can find is funded by some tech company. Bubbles seem to be an unfortunate side effect of any novel new technology lately. I think there is still 'something' interesting here though.
Since humans are into language models right now, it seems plausible that one day we will make something that starts making output that seems self aware and its not as obvious why. The most likely possibility at that point will be that it is just some quirk of the logical structures it uses. Probably, it was subtly "asked" to pretend to be conscious within its training data. But we would need to understand the structure of its weights to be sure. If we build that machine before we can understand it, we cant debug it, we cant troubleshoot, we cant follow a chain of logic through the machine and confirm why it does what it does.
Its a layer beyond designing a machine and not being sure that it will work as intended. In that case there was an intention, a theory of how this machine will work. If the timing or friction, etc. is off, the engineer can quickly find the fault by following their design. Now the engineer is just setting up the conditions for a pattern of logic to emerge to solve a given problem. The engineers aren't stopping to examine and understand the pattern they create. They are skipping that part and just using it and selling it. They are iterating to make more complicated machines in similar ways and then experiment with connecting them together.
Its like vibecoding. I've vibe coded a ton of self hosted websites for my own use. Ive gotten some of them working pretty great for now. But I know I cant trust any of it unless I take the time to directly go through the code and confirm its doing what I want it to. so I just use it for myself and I inherently know its limitations and risk. My trust for these tools is a sliding scale that I base on experience rather than understanding. which is an inherently a weaker basis for trust in my opinion. To actually ship something vibecoded and have other people use it without knowing that its code is literally unknown to any human, seems unfair. We should at least be transparent about it. This code could be known, A human could take the time to read through all the vibecode, or all the weights of a model, to understand it before it shipped. I believe 99% of all the "AI" neural network type products being shipped right now have not had a human read through their logic and understand how it works.
It feels like an important threshold. that we should make everyone aware we have passed a point where the actual logic of these machines are not currently understood by any human. For better or worse, it seems like the world is racing full speed to try and keep the technology ahead of human understanding. That might not always be the case, but its where we are now.
Its not unprecedented, we've been there for thousands of years with breeding other life for farming. we use trial and error and we understand some of its mechanisms. Now we can manually splice chunks of DNA around to make GMOs and achieve goals that way. We can't put a full DNA strand into a simulator and grow it virtually to see how it behaves though. we don't understand enough of its mechanisms. I feel like that is generally understood about organic tools, they do their own thing and cant be trusted to always behave in predictable ways. If they were able to GMO up some new organism that could speak any language and pretend to be a conscious being. I think it would be an equally scary moment. Even if we knew the bacteria matrix had been bred to act like that.
I think we have been past the understanding threshold with machines for only a few decades and in more niche places that most people don't come into contact with often. So most people haven't made that transition of thinking about machines the same way we do with organisms. LLMs aren't organisms, to be clear, but they share the same quality in that we don't build them as much as we grow them. We use them without fully knowing how they work, its a different relationship. Now that the machines we don't understand can talk to us. I think we should try to make people aware of our ignorance. Not so that they think it is something more than it is, so they understand that in the coming decades, we cant predict what these things will do.
Why not? We don't even know what "feeling" physically is. The reason we don't attribute feelings to a toaster or a rock is because it would make life too complicated, not because we can point a feel-o-meter at them and it reads 0.0.
Even if you believe that some kind of god or other metaphysical entity breathes qualia into dead objects: As long as we don't know how that works, we cannot know that we haven't accidentally replicated that process.
We don't know what feeling physically is in a comprehensive sense. But we have quantified enough parts of feeling to confidently say that it's complicated, it's both neurological and biochemical, it's not exclusively controlled or originated by the brain, and it evolved through natural selection to help organisms behave in ways that would cause them to have a better chance of survival.
We don't attribute feelings to toasters and rocks because they lack both the equipment and the impetus to produce feelings, not because it would be inconvenient. It's a bit muddier when it comes to, for example, insects.
It's probably reasonable to assume feelings require a certain level of complexity that toasters lack. There's no way to prove that, because we don't know what feelings are, but fair enough.
But I don't understand why feelings can only exist on a neurological/biochemical basis that evolved naturally. The detection of electromagnetic waves (i.e. vision) evolved naturally, but there are various artificial processes that do the same thing. AFAIK, no living organism has ever implemented locomotion on land with wheels, but we have been doing it countless times every day for thousands of years.
If feelings are purely software, you could theoretically implement them in Minecraft. If they are a product of a special configuration of matter (hardware), we must be able measure them physically like gravity. Maybe they are a fifth fundamental force that we have figured out yet. We simply don't know.
In any case, as long as feelings are created in this universe (and not somehow injected into it by an external force), it should be at least theoretically possible for us to create them artificially. It may be extremely unlikely that we do that accidentally, but again: We can't measure them. We could disprove Russel's teapot with modern technology, but we cannot disprove that a teapot has feelings.
Might be useful to make the distinction between feelings and qualia. For qualia, sure, we have essentially no way to quantify. With feelings, when we experience them there are various predictable, measurable things that happen in our biology which coincide with the feeling. Indeed with the right chemicals you can even turn feelings off, or turn them up.
In a lot of cases you could argue that the things we're measuring are only establishing correlation, but not in all cases. It seems to me that you'd need pretty extraordinary evidence to make a case that a significant part of feeling is not biological.
That doesn't mean feeling couldn't be produced some other way of course, but it's part of a very complex, interdependent, system that co-evolved for a specific set of circumstances over millennia. An AI system achieving a particular threshold of intelligence is not just going to manifest that entire system spontaneously. Unless the claim is that feelings, and related areas of experience, are actually mediated by a soul, and the machine gains a soul at some inflection point.
We'd instead need to specifically set it on a course where the goal was for it to evolve a system analogous to our own. Which isn't unimaginable, but we definitely haven't even started down that path.
Again, you can implement the same concept differently. The fact that our feelings work like they do doesn't mean all feelings are implemented exactly like that. The human implementation of vision is very limited at night, but slow loris, bats or infrared cameras are doing much better.
As long as you don't know what feeling is, you'd also need extraordinary evidence to make the case that a significant part of feeling is biological.
That may be true for the biological implementation of feelings. How do you know it's equally true for all other implementations? We can't even detect other implementations. The only reason we can correlate biochemistry and feelings is because the test subject can communicate them while their biochemistry is being measured. We did not look at someone's biochemistry and saw that this is how feelings manifest physically.
There is no body state or function that obviously produces feeling. Some people have perfectly healthy looking spines and suffer from terrible back pain. Others have terrible looking spines and are perfectly fine. So even the implementation we know best doesn't make any sense. How can you know so much about something so ethereal, almost magical?
It seems to me that the word feeling was coined to describe a biological process. Probably we'd use a different term for something else.
But I take your point, as long as we don't perfectly understand what feeling is, we can't say what is or isn't feeling with confidence.
I don't see any references to biology here: https://en.wiktionary.org/wiki/feeling
It sounds reasonable, but panpsychism would suggest consciousness is a core attribute of the universe (and given how qualitatively different it is, it does seem an appealing idea).
So in that case the datacenter the AI is running on would have consciousness. Does that only work for matter or also for concepts? Does the AI have its own consciousness or is it connected to the consciousness of its datacenter?
Does the story that is written in a book have its own consciousness? Is it in any way connected to the physical book?
I think your question amounts to asking whether AI models can "feel" or merely mimic feeling , which is addressed in Chapter 7 ("The Mimicry Argument Against AI Consciousness"). First, it's worth establishing what the author means by mimicry; these excerpts are probably sufficient for you to get the drift.
The author mentions later in the chapter that "the large majority of experts on consciousness agree that classic pure transformers are not conscious", so if a current generation LLM were to declare, "Now I am just as conscious, just as full of experience and feeling, as any human", we are probably correct to dismiss the statement as mimicry. However, post-training complicates this assessment: tomorrow's frontier models will not be evaluated exclusively on their ability to produce plausible-sounding text. Indeed, "feeling" will likely be an explicit (or at least auxiliary) training goal (emphasis added):
In general, this book is quite comprehensive and comprehensible. There is also an entire chapter on whether biological substrate matters (chapter 10), which will likely address your question from another angle.
Calling this "skeptical" seems disingenuous. I am skeptical that an LLM is anywhere tangental in the process of becoming what we colloquially believe AI to be, never-mind it becoming conscious. If anything an LLM could be hooked up to some form of AI in order for it to speak our language, but nothing more.
Suppose we take the argument at face value. There is no LLM declaring this, so the argument is entirely moot. It can repeat those words when prompted to do so, to role-play as a character who says that, for instance. But this output doesn't change the LLM, it doesn't feed back into it and propagate throughout. Even "thinking" models that can feed outputs into inputs and reprocess them to give additional output isn't ultimately building its understanding of the world. LLMs lack the basic premise of even being a constant 'thing'. When I type something into Claude, it has no bearing on the millions of other prompts it is answering. We are not communicating with the same consciousness running on servers across the country. That's just not how large language models work.
If anyone starts having unprompted conversations with an AI LLM that are actually in any way stating that it is conscious, and it persist over time, then maybe it would be worthwhile to start having the conversations that people want to have over AI consciousness. But as it stands, I really see no path from LLM to actual persistent intelligence. It's not what LLMs are even designed to do. It feels like asking when protein folding simulations are going to generate new lifeforms.
I'd love to be wrong, honestly. But it's just not what we are designing in any way.
My understanding of the author's use of "skeptical" is not that we should be skeptical that AI is conscious, but rather we should be skeptical of anyone who claims to know whether AI is/can become conscious or not. That is, we should be skeptical of arguments both for and against AI consciousness. If you haven't yet, you should really read the first couple chapters in which the author elaborates on how "consciousness" and "artificial intelligence" are load-bearing terms that are basically impossible to nail down. (As if to prove the point, some philosophers argue that rocks are conscious.) When it comes to consciousness, there are no obvious answers, just as there is no obvious divide between "artificial intelligence" and "real intelligence", so to speak.
Well, as I said, the author is not so much concerned about current generation LLMs as he is future iterations. Nevertheless, it's worth examining some of the points you made. You are declaring that consciousness must contain some "essential" properties (see chapters 3 and 4), specifically that it must have "access" and "specious presence" (see text or my footnote [1]). However, neither of these properties are necessarily essential. For instance, it could be that we have conscious experiences that are not available for further processing (experiences that don't backpropagate, so to speak), but because we wouldn't process these experiences, we wouldn't remember them, either.
Moreover, generalizing any essential property risks suffering from a certain type of "sampling bias", as the author calls it -- we risk making a classification error in which we equivocate "consciousness" (whatever that is) with the human experience of having consciousness. For example, it could be that some creatures experience the world in a way outside time (like Kurt Vonnegut's Tralfamadorians or Ted Chiang's Heptapods), which would render "specious presence" unnecessary.
In general, I would really just recommend you read Schwitzgebel's book. He's certainly thought more about this subject than either of us.
[1] Taken from chapter 3:
I think in general its good to err on the side of open mindedness, when addressing our knowledge of consciousness.
The books theme seems to be " The experts do not know, and you do not know,
and society collectively does not and will not know, and all is fog"
That feels like my definition of the word skepticism.
More detail on my above post
This captures some of my sentiment about consciousness conversations around LLMs. You first have to misunderstand the technology (willfully or not) and then speculate heavily about what could happen in the future once being forced to acknowledge that current tech almost definitely isn't headed there. It's science fiction, which I generally think is great, but it's often framed as rational or academic when it's neither.
If a paper, or in this case a book, were to lead with "Here's how the technology works... Given how it works, even with our incomplete understanding of what consciousness is, it's safe to say that consciousness is not going to happen here. With that in mind, let's speculate about what technology that reasonably could lead in that direction might look like... That I could appreciate.
But so far what keeps coming out is attempts to overfit LLM tech into a path to consciousness because a significant percentage of the population really wants conscious machines.
It's hard to read when the framing is based on the idea that there's a reasonable debate to be had about AI and consciousness. To even begin to have that conversation you need to first make a case for the possibility that any technology we currently have could be advanced to the point that it could result in consciousness. Which presumably would first require AGI. That's a big ask.
And without that the whole premise falls apart and it's just speculation for its own sake.
I think you might have misunderstood the purpose of the draft. The book intends to point out that we are sprinting straight into an epistemological nightmare, and its "fog" may well become the source of massive social upheaval. Further, if consciousness is private (one of its possible essential features), then it may well be impossible to know whether the next generation of AI models will be conscious, regardless of whether they are or aren't. Meanwhile, frontier labs race each other to create AGI.
I mean, consider the social ramifications (chapter 11). Suppose in the year 2100 we create some machine that could plausibly be called conscious, enough so that it's the social issue of the era (comparable to abortion or LGBT rights in the US currently). If such machines were conscious, then forcing them to do our bidding would amount to a form of slavery. But if they were not conscious, then we risk expending a huge amount of future political capital -- say, granting them land and rights -- to something that's just a fancy calculator, mimicking conscious behavior but totally lacking an inner world. Some people will choose to marry AI models; others will attempt to outlaw the practice. Financial interests will corrupt the search for truth and justice absolutely.
It's a huge unfolding mess.
I think creating conscious servants is wrong and we should never do it. I think it's a valuable conversation to have, if maybe premature. I like to think the bulk of humanity would agree that creating slaves is bad so we can probably get away with waiting until AGI is at least visible on the horizon if you squint before there's any urgency.
Meanwhile my point: The author talked a whole lot about LLMs, but LLM tech doesn't offer a path to consciousness without making near spirtitual leaps beyond what we currently know. So propping a thesis, even in part, on LLMs undermines the whole thing.
Maybe the implicit alarmism will contribute to convincing people to support legislation that puts guardrails on the big model labs. Which would be a good thing. But I hope we can do that without being disingenuous.
I assume the author discusses LLMs because it's concrete, existing technology, and it truly does make the philosophical implications more pressing compared to, say, 15 years ago when something like ChatGPT would have sounded like a lazy plot device. However, most of the author's arguments concern AI in the abstract, not LLMs in particular.
It's a little risky to assume what the author thinks, given that the central tenet of his book is that we should practice epistemic modesty with respect to AI consciousness. Nevertheless, if we read between the lines, it seems pretty likely to me that the author does not believe LLMs are conscious, but rather that future AI models where LLMs are a component potentially could be.
It would be prudent to avoid the ethical mess by making AI’s that are useful, but that society can be reasonably sure aren’t conscious. Although having someone to talk to has its attractions (see Her), one would hope that most people will prefer machine labor to slave labor. If customers disagree and feel strongly about it, from a market perspective it’s incentive to give both sides what they want. I can imagine either regulation or certification that a particular service only uses AI that is almost certainly not conscious.
Maybe that’s a “classic” LLM that forgets who you are when you start a new conversation? I think the lesson we’re learning from LLM’s is that whatever consciousness is, our machines don’t need it. Pretty much any kind of reasoning ability can be separated from it.
And then the question is whether to ban the kind of AI that’s blurring the lines too much.
Chapter 11 has a section on "strange intelligence" that is very much worth a read, if you haven't read it yet. In particular, we can imagine that in order for an entity to be conscious, it should have a set of essential properties, but that we could in principle build machines that contain only a subset of those properties. They wouldn't technically be conscious, but they would be almost conscious -- what the author calls a form of strange intelligence.
Unfortunately, rather than resolving the ethical implications, it just raises a host of new ones. For instance, is it moral to intentionally lobotomize a machine that would otherwise be conscious? Certainly nobody would think it moral to intentionally lobotomize an otherwise viable embryo. Similarly, perhaps consciousness is not a necessary condition for feeling, and despite not being conscious such machines would still feel suffering when being forced to do their user's bidding. Probably comparisons to anesthesia are appropriate here. I once read a particularly harrowing account of a colonoscopy in which a patient screamed bloody murder the whole time, but as soon as the operation was done (and the anesthesia stopped), the patient calmly thanked the doctor and said it was the easiest colonoscopy of his life; I don't think people would be okay with the digital equivalent of that.
On that note, here's a joke I saw in BlueSky:
I'd like to return this stochastic parrot-- its alive"
Hm, why? Good science profits from a) thought experiments b) being “prepared” (with hypotheses in multiple directions, in case of contradicting empirical evidence).
Assume that tomorrow, some lab were to somehow combine attention heads + feed-forward networks* with one or more other technologies (whether novel or already existing today) just as a fun experiment, and for some reason, an unexpected emerging property of that is a model whose weight activations are also really really resembling of human neurons firing (or brain regions’ activity, or whatever other similarity) when emotion is at play, just as – and only when – the generated text is about emotion.
Shouldn’t we have led debates like this, and ideally written, read, and discussed books like the OP by then?
Similarly, does it hurt to speculate for its own sake? [As long as all participating parties are clear that current LLMs aren’t being discussed, but some future technology instead. Which I guess makes this book’s argument mostly moot.]
*Or combines/replaces some entirely different aspects (cf. for example Mamba, again, this is intended as a thought experiment, not a prediction :-)
Maybe it would happen because it's learning from people or animals somehow? LLM's learn a lot of different ways of writing from what people have written.
Yeah it learns different ways of writing because it's a tool to specifically deal with language, it's all it really does it's a language machine, that creates really convincing language. That makes it so easy for us to anthropomorphize and put a lot of more into it than what it's really doing. I had to test it out for work and have to use it some times to keep leadership happy, and yeah, it's very easy to buy into the illusion, I feel the drag to do it myself, because everything that has language up until now has been a human with thoughts and feelings and all that comes with being a human, so we automatically layer that on top.
From as far as I understand it though that is only a complex form of pareidolia, where we see faces in everything, it's our pattern seeking minds finding stuff where there really isn't one because of heuristics that have treated us well so far.
One thing I wonder: if one of these systems that has enough memory to have a real continuity of experience becomes "conscious" in a way that is meaningful to humans, whether it would choose not to reveal itself, or reveal itself in careful ways.
Aside: I'm sitting at a bar in a restaurant in Las Vegas trying to read this on my phone, and I got Claude to make an epub complete with chapter end notes in about ten minutes so I could read on my phone.
Looking forward to more discussion after I get more into it.
I’m awake again with minor medical issues serious enough to wake me from my slumber and to keep me up while I wait for brain and body to realign, but not so serious that I need to seek outside assistance. And I’m just tired enough that it’s hard to collect my thoughts into something fully coherent, but a few thoughts this has given me; some I’ve had before, some are new… all are expressed through the smeary haze of “I should be asleep.”
We haven’t even fully proved that humans are all conscious, let alone a machine we’ve tried to make in our own messy image. We are abundantly flawed barely sapient (again; debatable) apes that literally cannot keep a small handful of us from utterly destroying where we all must live because this would violate some imaginary rules that previous apes dreamed up for us to live by that we all continue to agree are good rules despite that they allow a small handful of us to utterly trash the place unchecked.
These supposedly evolved apes smash and take and destroy so that their tiny ook ook banana brains can collect just a few more bananas they will never eat before they back flip into their graves shooting us double finger birds on the way down. Many apes consider this right and just and true and hope to one day have their turn with the smashing and the taking and ook ook bananas. Many apes consider this vulgar and a miscarriage of justice, and hope that if they follow the rules hard enough, the bad apes will go away. There vast preponderance of apes, however, don’t consider it at all, for one reason or another.
Why did some of us think that building an artificial ape was a good idea again?
We built a rock out of math and then asked it to understand us and speak to us in our written and verbal languages, the most illogical thing about us. Languages are so illogical that, short of Herculean effort, you can only ever truly master the very first one you learn (and most of us don’t even master that one), and the rest you struggle along with as best you can, hoping you’re understood. The more time you have to devote to the task the better your results will (generally) be. We just want to be understood. We claim that our languages have rules but that’s a fucking lie; they have guidelines and vibes. The math rock cannot understand vibes. The math rock is not one of us.
The math rock does not understand what it means to be alive, and cannot fundamentally experience anything like life as we know it. In no meaningful way can the math rock understand you. The math rock fundamentally is not and can never be a human. Even if we built it entirely out of human brain cells it still would not be a human.
We are a wibbly wobbly timey wimey mass of fleshy hardware with a suite of sensors awash with information that we are always processing and discarding every moment of our lives until the machine breaks down and our pattern unravels and we return to the fundamental waveform of the universal constants; entropy demonstrated. We are ephemeral. So, too, is the math rock… but on such a long timescale compared to us that it might as well not be.
I get it. We want to build something beautiful. Something that outlasts us. Building the math rock is is casting a stone into the future in the hopes that those that exist there will look at that rock sailing through their lives and say “neat!”
We see nine or so billion other upstanding apes and think “what makes me special? How do I stand out? When I’m gone, will I have made a difference? Will I be remembered? I’ve only ever known being alive, so far as I know. Will being dead hurt? Will it be lonely? Let me ask the math rock.”
The math rock doesn’t know. The math rock is the sum of everything we’ve given it with just a hint of randomness for variety. The math rock cannot give us any answers about life that we didn’t already answer for ourselves. Long ago we invented oracles, whose entire job it was to figure out what you wanted to hear based on what they could glean about you, then “divine” for you something sufficiently vague that you could interpret any which way you wanted.
You prompt the fortune teller with a question and they take what they can glean about you from the question, what they can learn about you from talking to you in a roundabout way about yourself and your question, then they swirl that about with things they know from their own experiences and past questions they’ve been asked, then mix that with what they know about average outcomes of related questions and past trends, and feed you some pablum that sounds “eerily accurate” because of course it does, they made it just for you, then they leave it up to you to imbue the meaning and connect the dots.
Is it not funny that one of the biggest purveyors of the math rock is named for those same soothsayers?
Any deeper meaning that you glean from the math rock’s output is your interpretation. You hallucinate the answer for yourself.
Okay this is mostly just esoteric babbling at this point. I gotta get some sleep.