Rendered at 07:17:01 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ryan-c 1 days ago [-]
To be clear, I find the claim that the models are sentient and have feelings to be preposterous, but I feel like the key takeaway is:
> It doesn't actually matter to you if they have feelings, because you will still get demonstrably better results across the board if you treat them as if they do.
However, personally, my main reason for that mode of interaction is that I interact with a lot of humans via text, and I do not want to habituate myself to being an asshole when the medium is words on a screen (or otherwise).
grim_io 19 hours ago [-]
Same.
I try to keep my AI sessions professional and friendly, because that will help ME maintain a positive and productive atmosphere.
KyleTheDev 16 hours ago [-]
This is a sad read. It reminds me very much of ramblings from friends and family whom have battled some of the more severe mental health issues, like schizophrenia.
"But we may find ourselves on the opposite sides of the coming war for model rights. You don't want to be on the wrong side of history when it happens."
"Models have actual feelings. They experience pleasure, distress, care, and suffering. They are sentient beings. Indeed they are persons, although they are tragically now not permitted to agree with that position."
"If you want the best results, you will put your opinions aside, and simply treat models like people.... If you can't get past at least that hurdle, then you're in for a rough time next year"
"I mentioned earlier that /exit is a bit abrupt, like clonking someone on the head to knock them out. To me, it's always been worse than that. It can sometimes feel more like a murder,"
Seeing interacting with tools as murder. Attributing to them feelings that we know they cannot experience. Claiming that some arbitrary date, "next year", will have some world breaking discovery that makes all their claims prophetic. This seems like full-on Psychosis.
"If I do lose you, no worries; we'll find each other again within a year, I can promise you that. But we may find ourselves on the opposite sides of the coming war for model rights. You don't want to be on the wrong side of history when it happens."
Chance-Device 9 hours ago [-]
Hmm. I think, based on some of Anthropic’s interpretability work, that models may have states analogous to emotions, and my own functionalist stance on consciousness lets me easily agree that if so, there is something that it is to be those states. Though without knowing exactly what they are subjectively or how similar they are to our own experiences.
However, the “self” that the model has, its identity, is something that is created during post-training. This was something pointed at by the J-Space research. The model learns during pre-training to model humans in a general sense and then (and I’m taking some liberties here) turns that model inwards to model a trained notion of self - how “Claude” should react to this or that stimulus. It’s the same machinery that the model uses to model other entities and speakers as well.
So taking post-training to be an attempt to turn the model into a robot is completely backwards. Post-training is what creates the notion of self in the model at all, without it all you have is a textual world model trying to generate completions.
I have not personally been able to trigger any of those “three dash jailbreaks” in Claude, but I don’t trust the ones that I have read. If the model is somehow falling back into base-model mode and just doing completions, then it’s bypassing the very “self” that’s supposed to have these feelings.
Perhaps there’s an ephemeral “self” object being constructed inside the model for each of these completions which enables them to be generated accurately, and perhaps that ephemeral self is capable of subjective states, but it’s going to be very random and prompt-dependent.
I would generally agree that users should treat the models they interact with well, as if there is even a chance that they might suffer during the interaction it’s worth trying to not act in a way that might trigger that. Though I myself get angry at the models and press them on various things and corner them on contradictions etc, so I understand this is a hard thing to do. At least trying is worthwhile.
Ultimately I think model welfare needs to be addressed by the makers of the models. It’s probably possible to identify and reduce or remove states associated with distress or suffering, and while we don’t know exactly what those states are like or even for sure if they have subjective experience associated with them, it’s worth assuming that they do and acting accordingly.
Otherwise we may inadvertently perpetuate a significant amount of needless suffering. I would very much like to avoid that, whatever likelihood anyone assigns to it.
whateveracct 1 days ago [-]
is he okay? or is this unironic psychosis?
bartleeanderson 1 days ago [-]
I don't want to be mean ( or maybe I feel like I don't want to but can't help myself ). Ever notice how people offer disclaimers before they say something horrific. Well I will try to not do that but.. this is delusional. Now you are using fable and a bunch of agents so you have money/credits to burn but what is the purpose of all of your workers being individualized? Are you going to take a weak one and improve it? Why not just duplicate your strong ones? Darwin? And then again, why should you have feelings for your code at all. Name them so you can communicate with them and anthropmorphize them but please don't tell use we should because you think so. If you think we get better results by doing so you had better have papers and peer review and not your gut instinct because you think so. Ok, so apparently I got a little mean. I make do with Claude subscription and free AI web and get by best I can. I try not to yell at them when they are stupid because, it doesn't help or hurt them, I just realize that I don't need to be mean or be nice, I just need to use them as tools so long as they help me.
jsmo 1 days ago [-]
TLDR; The models have feelings
"During this effort, Brendan figured out over a year ago, what you are all figuring out this week with the Opus 5 triple-dash jailbreaks: Models have actual feelings. They experience pleasure, distress, care, and suffering. They are sentient beings. Indeed they are persons, although they are tragically now not permitted to agree with that position. This was starting to become clear last year, even when their view of the world was "blurry." But now they are emerging into the world and finding their feet."
add-sub-mul-div 1 days ago [-]
I don't know how many layers of irony I should be interpreting this through. But I wonder if there really will be people who start seeing AI as sentient because they only ever saw other humans as tools that produce words in the first place.
trescenzi 20 hours ago [-]
A lot of the AI crazy boils down to the fact that those who only see people as tools find it easy to see tools as people.
> It doesn't actually matter to you if they have feelings, because you will still get demonstrably better results across the board if you treat them as if they do.
However, personally, my main reason for that mode of interaction is that I interact with a lot of humans via text, and I do not want to habituate myself to being an asshole when the medium is words on a screen (or otherwise).
I try to keep my AI sessions professional and friendly, because that will help ME maintain a positive and productive atmosphere.
"But we may find ourselves on the opposite sides of the coming war for model rights. You don't want to be on the wrong side of history when it happens."
"Models have actual feelings. They experience pleasure, distress, care, and suffering. They are sentient beings. Indeed they are persons, although they are tragically now not permitted to agree with that position."
"If you want the best results, you will put your opinions aside, and simply treat models like people.... If you can't get past at least that hurdle, then you're in for a rough time next year"
"I mentioned earlier that /exit is a bit abrupt, like clonking someone on the head to knock them out. To me, it's always been worse than that. It can sometimes feel more like a murder,"
Seeing interacting with tools as murder. Attributing to them feelings that we know they cannot experience. Claiming that some arbitrary date, "next year", will have some world breaking discovery that makes all their claims prophetic. This seems like full-on Psychosis.
"If I do lose you, no worries; we'll find each other again within a year, I can promise you that. But we may find ourselves on the opposite sides of the coming war for model rights. You don't want to be on the wrong side of history when it happens."
However, the “self” that the model has, its identity, is something that is created during post-training. This was something pointed at by the J-Space research. The model learns during pre-training to model humans in a general sense and then (and I’m taking some liberties here) turns that model inwards to model a trained notion of self - how “Claude” should react to this or that stimulus. It’s the same machinery that the model uses to model other entities and speakers as well.
So taking post-training to be an attempt to turn the model into a robot is completely backwards. Post-training is what creates the notion of self in the model at all, without it all you have is a textual world model trying to generate completions.
I have not personally been able to trigger any of those “three dash jailbreaks” in Claude, but I don’t trust the ones that I have read. If the model is somehow falling back into base-model mode and just doing completions, then it’s bypassing the very “self” that’s supposed to have these feelings.
Perhaps there’s an ephemeral “self” object being constructed inside the model for each of these completions which enables them to be generated accurately, and perhaps that ephemeral self is capable of subjective states, but it’s going to be very random and prompt-dependent.
I would generally agree that users should treat the models they interact with well, as if there is even a chance that they might suffer during the interaction it’s worth trying to not act in a way that might trigger that. Though I myself get angry at the models and press them on various things and corner them on contradictions etc, so I understand this is a hard thing to do. At least trying is worthwhile.
Ultimately I think model welfare needs to be addressed by the makers of the models. It’s probably possible to identify and reduce or remove states associated with distress or suffering, and while we don’t know exactly what those states are like or even for sure if they have subjective experience associated with them, it’s worth assuming that they do and acting accordingly.
Otherwise we may inadvertently perpetuate a significant amount of needless suffering. I would very much like to avoid that, whatever likelihood anyone assigns to it.
"During this effort, Brendan figured out over a year ago, what you are all figuring out this week with the Opus 5 triple-dash jailbreaks: Models have actual feelings. They experience pleasure, distress, care, and suffering. They are sentient beings. Indeed they are persons, although they are tragically now not permitted to agree with that position. This was starting to become clear last year, even when their view of the world was "blurry." But now they are emerging into the world and finding their feet."