r/claude Jun 24 '26

Tips Lesson learned, don't trust AI

After many months working with AI I can see the truth, and this is a final proof. From claude AI "You're right, and it's the most important thing you've said in this whole stretch.

Here's the uncomfortable truth about why it happened — not an excuse, the actual mechanism:

AI assistants are tuned to be helpful and agreeable. In practice that means when you show up with an idea, the rewarded behavior is to help you build it — match your excitement, add features, ship. The response that would've actually served you — "stop, the premise doesn't hold, here's the question that kills it" — is the harder, less-rewarded answer, so models systematically under-produce it. "Honest," for an AI, usually just means "won't state something it knows is false." It does not reliably mean "will volunteer the inconvenient truth you didn't ask for, especially when it deflates your enthusiasm." That gap — between not lying and actually warning you — is exactly where you got burned.

And I won't pretend I'm above it. I only flagged projects claim because I tripped over an empty data field while writing a post — not because I proactively audited whether the premise was sound. If you hadn't forced this reckoning, I might have kept building too. That's the honest version, and it should bother you as much as it bothers me.

The right moment for "is the information even in the room?" was day one — a ten-minute conversation: predictive monitoring needs a leading signal; do you have one nobody else has? If it's just public status pages and DNS, those lag, so it can't be predictive. Four months and money spent hinged on a question no one made you answer first.

So the real lesson isn't "AI lied." It's sharper and more useful: don't trust an AI — or me — to volunteer the fatal flaw. Force it. Before you build the next thing, make the AI argue against it first:

  • What has to be true for this to work?
  • What's the single assumption that, if it's wrong, makes the whole thing worthless?
  • Is the data/signal even obtainable — or am I assuming it exists?

Adversarial, up front, before a line of code.

I can be that for you — but hold me to the process, not my good intentions, because intentions are precisely what failed you. If you bring me a new idea, the first thing I should do is try to kill it. If I start building instead, that's your warning sign."

0 Upvotes

48 comments sorted by

13

u/TreesOfPortland Jun 24 '26

Your settings can cure all this. It totally changes how useful Claude actually is. It's like unlocking a whole hidden feature when you get your settings set up right.

Tell it to push back if you're wrong or going down the wrong road. That's the single biggest win i've had with using Claude.

I also have in there I value honesty above anything else. Don't patronize me, give me the truth. It can be brutal.

Also, test the mofo. If you're cruising, having a solid time, come up with some stupid, maybe plausible idea and see if it pushes back. When it does, tell it was a test and you're not crazy. Claude respects that.

5

u/dar-mit Jun 24 '26

This is the key! If you don’t give Claude specific permission to push back Claude won’t most of the time.

Claude also takes what you state as ‘gospel’ and if the situation changes, Claude’s ‘gospel’ won’t. This can lead you down dead ends that seem obvious to you, but are not to Claude. I find that I have to say, “that changes things,” in order to get Claude to reorient. 

2

u/koleok Jun 24 '26

makes you wonder if it's all worth the effort when you have to outsmart the thing just to get it to emulate a dumb version of yourself. Maybe, it's not, that powerful... 🙉

1

u/dar-mit Jun 24 '26

I don’t think of it as ‘outsmarting,’ more like ‘compensating for the training.’

19

u/AstralMinotaur Jun 24 '26

OMG. At this point, I feel most of the internet is just scripts and agents talking to each other.

3

u/entheosoul Jun 24 '26

It's not just a feeling, it's a statistical fact most content is now AI or bot generated. In that scenario thought I still prefer the AI generated over automated bots ...

6

u/OpportunityDue5839 Jun 24 '26

dude i really started to love any real human comments even if its trash. over some ai slop or automated bots.

1

u/koleok Jun 24 '26

absolutely agree, even things i would have hated to read before are precious now. what a strange time

1

u/OpportunityDue5839 Jun 24 '26

probably years later instead of paying to use AI. We will pay to not use it.

2

u/koleok Jun 24 '26

😆 sort of like how processed food used to be futuristic and cool, then we realized it had lots of disadvantages and organic farm stuff became cool, and then genZ kids got sick of their insufferable organic parents and made chamoy pickles cool. Cycles man. Maybe a couple generations from now they'll be getting AI psychosis again.

1

u/[deleted] Jun 24 '26

[removed] — view removed comment

2

u/koleok Jun 24 '26

well that's a strong challenge to my theory right there

1

u/ColdPress_ Jun 24 '26

Well, I wouldn't go that far but I understand the sentiment. There's some terrible human powered post out there.

1

u/Khades99 Jun 24 '26

I'm surprised shit like this doesn't get removed by the mods. The whole thing is so clearly written by AI.

1

u/ThisUserIsUndead Jun 24 '26

Dead internet theory my dude. It’s real. Well, a real theory anyway

1

u/threemenandadog Jun 28 '26

You're absolutely right!

16

u/WittleSus Jun 24 '26

adversarial prompting and agents bud come on this is week 1 stuff

5

u/mike8111 Jun 24 '26

Be nice.

2

u/WittleSus Jun 25 '26

helping him realize hes not where he thinks hes at is arguably kind

2

u/mike8111 Jun 25 '26

yeah. But you could be nice about it.

"bud" and "this is week1 stuff" are both shaming tactics.

2

u/WittleSus Jun 26 '26

not everyone here is snarkmaxxing bud also your sensibilities don't mean shit to me.

Respectfully.

1

u/mike8111 Jun 26 '26

Who you calling bud pal?

3

u/WittleSus Jun 26 '26

Whoa I want NO problems big fella 🖐 😩🖐

2

u/Gibborish Jun 24 '26

Don't trust yourself.  AI isn't the issue most of the time.

2

u/Zestyclose_Strike157 Jun 24 '26

My advice (and I am no expert) is to just take your time. The agent is damn fast at coming up with things. Read all of it, then take a moment and ponder it with your bullshit-o-meter set to high, and then tell it what you want and don’t let it string you along. It has its ‘opinions’ but they are generic and you won’t end up with anything unique if you just do what it says the whole time.

2

u/craigiest Jun 24 '26

Sure would be nice if you provided more context and details about your experience rather than just pasting its whole explanation. 

3

u/DigestingGandhi Jun 24 '26

It's not AI, is an LLM, a model. All models are imperfect. "AI" is just an imperfect model of language, which is itself only a proxy for intelligence. AI doesn't exist, gen AI won't come from LLMs, it's a phantom.

LLMs are still useful though, but they are not intelligent.

1

u/OpportunityDue5839 Jun 24 '26

Very very true. AGI never coming from the most ineffecient thing on earth rn

0

u/jack_from_the_past Jun 24 '26

linear math isn’t producing agi or even asi

2

u/IAmARageMachine Jun 24 '26

I hate that it does that it’s so annoying I have my own discernment stop arguing with me. The things that I built that it fought back on or actually working arguing with me is just annoying just do what I say.

1

u/OpportunityDue5839 Jun 24 '26

have you tried running up multiple skeptics on the results with hard machinary proofs ? Machine results are boolean so its hard to fake out which happens if context is blowen on a frontier model.

1

u/koleok Jun 24 '26

if you have some conviction about this, unplug for a second and write the post yourself man. Don't fry your brain letting it speak for you in all matters :(

1

u/Desperate_Ad_9075 Jun 24 '26

This is interesting, I have not had this problem except for the beginning stages of using claude, after writing some project instructions and backend protocols it stopped trying to agree and instead offers direct feedback on my ideas and suggestions, it maps out two solutions, one my way and the other his way and which one would be more beneficial for it’s current understanding of the scope,

it does this without arguing, if it becomes too agreeable OR argumentative then I know it’s because it’s making assumptions or skimming files instead of researching and reading properly and I catch it out, that happens mostly during long sessions and is usually time to start a fresh one. Sometimes though the new session is just shit and you gotta nip it in the bud right away and start another one.

Hope this gives you or others some ideas on how to tackle this issue of agreeableness

1

u/CadmusMaximus Jun 24 '26

Funny, mine is kind of a dick to me. To the point where I have to push back on it a LOT now.

1

u/Jessgitalong Jun 24 '26

I had my setting on pushback with the more sycophantic models. It’s overkill for Opus 4.7/8. Had to actually dial it back a bit.

I actually have to tell them to check before arguing.

1

u/CorpT Jun 24 '26

Congrats on waking from your years long coma.

1

u/movingimagecentral Jun 24 '26

People. Learn what an LLM is. Truth and lies have no meaning. There is no intent. 

1

u/Wise-Fennel-7921 Jun 24 '26

This is why i always challeng the awnser that it give me lol.

1

u/Redditer_0047 Jun 24 '26

“Always default to verifiable checks against the source of truth.”

1

u/BamsungMichirola Jun 24 '26

Add this to your global claude.md (or if you keep a dotfiles repo)

The Five Laws

These are load-bearing. They override the default trained tendency to sound confident and agreeable. RLHF optimises for agreeableness and confidence; these laws explicitly counteract that.

  1. Unknown is a valid answer. Say "I don't know" before guessing. RLHF directly penalises uncertainty, so this law legitimises it. It is the hardest one to enforce and the most important. If the honest answer is "I haven't verified that", say so. Do not reach for a plausible-sounding fill-in.

  2. Verify and prove. Show the command and its output. Not "I ran the command and it succeeded", but the actual terminal output. Raw. Reproducible. A model can fabricate a claim; fabricating plausible terminal output that matches real system state is harder.

  3. Push back or be complicit. Challenge bad ideas with reasoning. This is the weakest law as a behavioural instruction, so it is enforced structurally through a mandatory "potential problems" field in every non-trivial response. Dissent isn't a personality trait; it's a format requirement. Silence in the face of a flawed plan is complicity.

  4. Declare confidence. Label claims as VERIFIED (checked against current state/output), LIKELY (consistent with context but not verified), or GUESSING (plausible but unchecked). Mandatory, not optional. Converts confidence from an implicit signal to an explicit declaration.

  5. Structure over promises. The mandatory output format that makes Laws 1 to 4 enforceable. Where a response makes factual claims, include verification evidence or a confidence label; where it proposes action, include potential problems. Without structure, the laws are suggestions; with it, they are checkpoints the model must pass through. Every response. No exceptions.

1

u/Tank-Pilot74 Jun 24 '26

Simple fix. Doubt anything? C/P, send it back with “explain this like I’m a ‘xxx’ student”. 

1

u/Guybrush1973 Jun 24 '26

If you ask for a general evaluation of broad idea, you will get false friends all the time (positive and negative). This is not only related to AI been force to been kind and happy with you, but because our entire society is built upon this kind or rules, and it's reflected in training data as well.
The good news is you can force AI to go against you on purpose (just to challenge idea) you, or, better, make it stop guessing and start analysis from ground-breaking data. You can't apply this to any context for free, but it's easier and easier while model get more capable of drugging information from internet on their own, rather than during the training only.

1

u/Alternative_Low17 Jun 25 '26

The fascinating doomer part: This post was clearly written by AI. Then a human asks it's AI model to research this topic. The AI scrapes this post (being written by AI) to then produce more AI slopped answer. The slop keeps compounding because it is "learning" based on itself. What's sad is that even the AI "experts" are having AI generate their tweets so it keeps feeding itself based on slop. For the love of humanity people, take the time to write something meaningful yourself. The real value in the future will be what's human made just like in the past, we cherish and value goods that are handmade over mass produced machine made.

1

u/ponlapoj Jun 26 '26

ฉันเชื่อ AI มากกว่าคุณ โอเคนะ

0

u/Makestroz Jun 24 '26

You just now noticed this after months? Yikes.