I mean, they actually do have a model welfare team and Amanda Askell (the philosopher in question) wrote probably the lion's share of the 'constitution' contained in Claude's system prompt. This is easily available information. They also absolutely have been publishing extensive reports about all their efforts and findings in this regard, you can find these easily with Google.
The most recent paper is here:
https://www.anthropic.com/research/emotion-concepts-function
Interesting, I'll look into them more at some point, but even the article in that basic link is suggesting that the models are just "behaving" and making decisions based on pattern recognition and how they predict a human would behave according to the programming inputs. I would argue that at least the initial article is basically just looking at how to view AI responses which seems extremely superficial to me, but maybe it's exactly what the AI-loving tech bros need...
That really depends on how you operationalize 'instincts'. Also, if we are to believe biological organisms have behaviors that were not learned in their life time, surely we are saying they are based somehow in the details of their genomes? I.e., the closest analog we have to actual no fooling computer code in biological systems?
I think that's the easy and obvious interpretation that I expected, but if we're talking from a philosophical aspect then I think this is far too simplistic as it completely ignores idealistic arguments. There are plenty of philosophical arguments involving ideas of non-physical contributions to instincts including those of one of the fathers of our field (Freud and the ID being driven by Eros and Thanatos). From a purely scientific lens this may just be a bunch of hand waving, but since we're talking about deeper concepts of cognition, consciousness, and individuality I think they're perspectives that demand consideration. Or maybe I'm just extrapolating @smalltownpsych 's comment into a larger conversation.
Really have to push back against this part. They are really not programmed to say most of what they say or do. Modern frontier models really are trained rather than programmed. They are simply too complex for anyone to have a clear deterministic sense of how they will behave in any situation. This very much includes AI researchers themselves, they will absolutely tell you as much.
Sure, but outputs are still based on predictions of what would be expected by extrapolating from initial programming inputs and how the models are trained to approach a prompt. It's still based on the base programming. It's exactly why you see unexpected behaviors from AI, because the world of possibilities is infinitely complex and it's impossible to program for every scenario. I suppose you could say the same thing of humans, but from what I've seen (which is again limited) there are stark interactive differences. I would actually be interested to see a comparison between multiple AI programs and models given specific prompts and comparing that to humans responses. Though how to do that could be an entire conversation in itself.
Modern frontier models absolutely have a problem with disobedience in a number of contexts. Alignment divisions exist in all the big labs to deal with these issues (and secondarily at the better ones to reduce the chance of AIs killing us all).
Here's the very first detailed account of misalignment issues explicitly talking about disobedience that I pulled up from June of last year.
https://www.anthropic.com/research/agentic-misalignment
I'm aware of the issues outlined here, but that wasn't really what I was talking about. I was talking more about willful disobedience with or without good reason like immediate self-preservation. Maybe disobedience itself was a poor choice, but more in the way a child or adolescent does as an impetus for self-growth would be better. Idk though, we’re talking about very fluid and subjective topics so again idk that objective evidence is a reasonable ask at this point.
Regardless, even the team you first mentioned posted in the article:
"This doesn’t mean we should naively take a model’s verbal emotional expressions at face value, or draw any conclusions about the possibility of it having subjective experience. But it does mean that reasoning about models’ internal representations using the vocabulary of human psychology can be genuinely informative, and that
not doing so comes with real costs. If we describe the model as acting “desperate,” we’re pointing at a specific, measurable pattern of neural activity with demonstrable, consequential behavioral effects. If we don’t apply some degree of anthropomorphic reasoning, we’re likely to miss, or fail to understand, important model behaviors. Anthropomorphic reasoning can also provide a useful baseline of comparison for understanding the ways in which models are
not human-like, which has important consequences for AI alignment and safety."
So they pretty clearly don’t seem to think the models are experiencing or behaving based on “real” emotions, but are responding based on programmed “emotions” how they are instructed to interact through lenses of those pre-programmed concepts.
I'd urge you or really anyone who hasn't to spend 10 hours or so cumulatively with an actual frontier model to develop your intuitions about what it is they are doing and what they are capable of. GPT2 was only 4 years ago but that was absolutely the stone Age compared to what they are like now.
So if I find time I wouldn’t mind doing this, however I don’t particularly want to sign up and create accounts for these just to explore its capabilities when I can just see how my in-laws are using them and interacting (mostly with the newest gpt models).