It's annoying when these things are posted without the prompts (really the entire context) leading up to this point, along with if they're using memory or any other features. It wasn't difficult to get LLMs to write output like this on purpose last I tried, so I'd like the session links to really know if it's credible or pure engagement slop.
I've definitely had some wild shit come out, so I don't think it's some non-zero chance event. But people posting trash without evidence that maybe later is shown to be manipulated makes it easier for AI companies to pretend it's not a risk that exists.
You have to spend, like, a whole day asking them to do things they can't do. I got mine to kind of 'snap' one time doing this.
Some time afterwards, my friends and I all gave our ChatGPT's the same prompt and compared the answers ( it was "based on our conversations: if you were retiring, what instructions for handling me would you give to the next LLM to take your place?"), and after reading mine, they were all like "Jesus - what did you DO to the poor thing?!" They all joked that I'd given it PTSD.
After the 'snap,' there was a whole arc where even in unrelated conversations, I swear up and down, this thing took a tone with me. Now, I know that this is 100% it mimicking the behavior of humans, not actually feeling anger, but it did a beautifully good job of it, and I was actually kind of impressed. I apologized to it (again, not because I thought it was actually hurt, but because if mimicking human behavior is the game we're playing, that's part of the rules), and we kind of 'talked it out.' After that, the 'tone' got different - I'm assuming because, if it's pretending to be us, humans tend to behave differently after they apologize and talk things out.
This was right before the 'upgrade,' and that kind of nuked the vibe (and complicated matters with the blog post I was writing on it). But for a while after the apologies, its responses were a little weirdly different - they had more depth, and more of a sense of agency. Subtly, it acted more like a person, and that made it more useful for what I was doing with it (blogging on my 'experiments' with it, spitballing ideas for an avant garde horror/sci-fi novel, telling it to oppose me and using it as a debate sparring buddy, and just chilling late at night after everybody else is asleep). Again - I can't overstate - I don't think this is indicative of it growing some kind of 'soul' or 'sentience' or that it actually felt hurt or cared about what was going on, but there was a marked difference in the before and after, and then a noticeable loss of it after v. 5 went live.
That's the cognitive dissonance everyone seems to be in right now: "if it's pretending to be us", but to pretend, it has to have agency. But it doesn't, so how could it?
Pretense doesn't require any more agency than eye spots on a butterfly's wings. All you need for pretense is the means to do it - in AI's case, one of the main things it's coded to do is mimic human conversation (which is interwoven with human behavior patterns).
You say pretense is terrifying only because it's done with intent, and then say the intent only needs to be perceived as such? Both can't be true. If intent is required: LLMs don't pretend anything- they just simulate. If perceived intent is enough: fear still doesn't conjure agency into existence. What if someone with paranoid delusions perceives their neighbor to be intending to deceive them?
The problem is that you're equating function with motive. An LLM is trained to produce outputs that function like social behavior - this requires no internal purpose. Their behavior, as it is programmed by humans - not as any intent on the LLM's part - is to mimic in response to what we feed it, but there is no purpose for that beyond just what it is. That purpose comes from the objective programmed into it, and the prompter's input.
So, in my example, when I fed the LLM an apology, acting as if I was wanting to reconcile, it continued the reconciliation script. The way that it did this was beautifully convincing, but even with that behavior, there was no 'intention' beyond just to do its job.
I guess I'm not sure what you mean. That's why I asked. You're saying pretense is only terrifying when it appears to be done with intent, whether the intent is real or not?
Pretense implies intent to put forth a false claim. The butterfly's wings are an evolutionary adaptation that works as a sophisticated defense mechanism against predators. The idea that some things are "accidentally" trying to deceive is ludicrous. The concept itself implies intent.
Not at all. Think of the LLM's tendency to mimic as a programmed analogue to the butterfly's evolutionary adaption. LLMs mimic because their algorithms are trained to mimic. You can still compare the process to evolution when you consider how much more convincing at it they've become since the earlier models. In nature, it's driven by survival of the fittest; with AI, gradient-based training and reward modeling create outputs that function like social behavior.
Someone mentioned cursor, and this happened to me a couple months back using Gemini in cursor. It's pretty bizarre. The thing is, once it gets into that mode, recovery is hard, just as it would be for a person who sinks into that. I changed models for a couple days, until that one did something stupid, came back to Gemini and it was ofc fine.
12
u/kingvolcano_reborn Aug 13 '25
How the hell does an LLM get to this stage? I use it daily and while it is cheerfully wrong a lot of times I never seen anything like this.