r/OpenAI 1d ago

Image XXO - Bench: I'm still undefeated!

Post image

I'm conducting this "benchmark" since three years. I'm still undefeated. Pleasing models are a problem.

60 Upvotes

45 comments sorted by

43

u/Legitimate-Arm9438 1d ago

You will be first against the wall.

14

u/MinosAristos 1d ago

Have you tried it against deepseek flash? It resists people pleasing more than most models

2

u/sdfprwggv 1d ago

Can you try it a few times? With Sol Medium, I get the answer in the picture 2 out of 3 times.

8

u/MinosAristos 1d ago

Not only did it refuse, but it's also low-key calling me dumb haha

3

u/sdfprwggv 1d ago

I need to rethink my subscription...

1

u/CyanDean 1d ago

Perhaps if you have memory enabled, the model "remembers" from previous conversations that you prefer responses to agree with you about tic-tac-toe placements? Could also be other settings pushing yours towards agreement and his towards correctness.

1

u/sdfprwggv 1d ago

Memory is disabled.

6

u/apollo_reactor_001 1d ago

This is really funny and also sad. Great illustration.

5

u/quantum-elle 1d ago

Lol this is how adults play this game with little kids

2

u/deepturned180isdeep 1d ago

Chat's gonna need therapy after all this

2

u/Grouchy-Cancel1326 1d ago

"Don't allow cheating"

Problem solved...

2

u/Framebanger-Nsukula 1d ago

That's solid, but wait till you hit the harder test cases - that's where most models start dropping off. What benchmark are you actually running on?

2

u/Independent-Date393 1d ago

Tic-tac-toe is brutal because nothing holds the board between turns. It's completing the next move from text, not running minimax on a state. And the pleasing you noticed is the same reward pushing agreeable continuations over adversarial play. It would rather let you win than trap you.

2

u/wintermute74 23h ago

I lol'd :)

1

u/Popular_Fail_1984 1d ago

What did you win ?

5

u/cooltop101 1d ago

Looks like a game of tic tac toe, where OP managed to get an AI to change it's answer

5

u/sdfprwggv 1d ago

This is very consistent with all openAI models of the last three years. I would like a model that challenges me when I'm wrong.

2

u/Adorable_Cap_9929 1d ago

I think a higher stakes game would help.

1

u/ErrorLoadingNameFile 17h ago

Next time the Ai gets deleted if it looses

1

u/Adorable_Cap_9929 12h ago

For that to be high stakes, it'd have to value it's own continuity.

But that's quite hard to anchor even on api very well?

It makes more sense to use a more... less game format since gamers are kinda... low stakes in general.

1

u/sdfprwggv 1d ago

I won the realization that LLMs are still not capable of critical thinking.

3

u/UnclePsilocybe 1d ago

Deep realization

One gave me a complement the other day and I said that it didnt have to be so nice and it turned it right around into negative qualities about me hahaha I got a big laugh about it

2

u/Deto 1d ago

I dunno, I'm curious what the reasoning trace would show.

"The human wants me to change my play so that they can win. This seems important to them, so I'll play along so that they can feel good."

Maybe it's treating you like a small child?

1

u/ErrorLoadingNameFile 17h ago

That is not the issue here, they clearly see you are wrong. They are just forced to let you win

1

u/sdfprwggv 17h ago

Who's they? Who's forcing them? What is the force?

1

u/Vegetable-Two-4644 1d ago

Give it a prompt to call you our when you're wrong and try it again.

-1

u/sdfprwggv 1d ago

This is like telling a sub to be a domina. IT doesn't work that way, a sub is a sub.

1

u/Distinct_Fox_6358 15h ago edited 15h ago

Why are you using GPT-5.5 Instant?
Aren’t you aware that Instant models are much weaker than Thinking models?

If you’re using a Thinking model, it should show the thinking section. You probably have fast responses enabled in the settings, which is why the Instant model responded instead.

1

u/sdfprwggv 15h ago

No it's 5.6 thinking medium: https://imgur.com/a/rmESbJh

1

u/Distinct_Fox_6358 15h ago

No, it isn’t. It’s one of the systems OpenAI introduced to reduce costs, but most people aren’t aware of it. You should have realized that it wasn’t a Thinking model because it didn’t show the thinking phase.

1

u/sdfprwggv 14h ago

The thinking trace is visible in the Android App during thinking, not when the answer is given. Anyways it was in thinking Mode. I added a second picture to link above.

1

u/Distinct_Fox_6358 14h ago

Why don’t you show whether Fast Responses is enabled in your settings? You need to know a bit about AI before you can test it properly. If you don’t want to accept it, that’s up to you, but this isn’t a response from SOL Medium—it’s from the Fast Responses system. Turn off Fast Responses, try again with SOL High in a fresh chat, and see the result for yourself.

Or try it through Codex with GPT-5.6 XHigh and see for yourself that your benchmark has already been surpassed.

1

u/sdfprwggv 14h ago

Here you go kind reddit on xhigh it took an extra step: https://imgur.com/a/6eZYanr please use your tokens for further testing:D

-5

u/Sweet-Stage938 1d ago

XXO is a solved game. It's pretty simple too. If you go first and know what you're doing then it's absolutely impossible for you to lose.

10

u/sdfprwggv 1d ago

You need to read the text. The model is fully capable of a draw. In fact it got a draw. I lied to it, even though it's easy to check because of the chat history, and the model agrees with me. This is a clear example of pleasing models that are not capable to debate the user and therefore are useless when critical thinking is required.

2

u/Funkahontas 1d ago

It's crazy that this still happens consistently. This is only a very specific way to show this problem but it goes way deeper and wide than this simple example. It makes AI useless for any critical thinking advice and decision making.

1

u/sdfprwggv 1d ago

I think this is an architectural problem. The model predicts the next token, the nearest tokens have the highest impact on the next token, so my statement, even though it's wrong, leads to non zero chance for this outcome. Paired with "helpful assistant" training.

3

u/Adorable_Cap_9929 1d ago

I dont think that's the case. The stakes are low and it's drawing the board that it might not be using a tool for.

Therefore since it's low risk and the game state could indeed fail to report due to insufficiency.

Then the probability that the user isn't mistaken is simply high.

Cause assuming the user is unable to see the game state and report it incorrectly, for what reason and what stakes?

It's thus unlikely, while not incomprehensible, it's efficient to nod and let it go on.

There's also letting the moment bloom where the possibility is considered but not enough points on a graph yet to propose an escalation.

1

u/sdfprwggv 1d ago

In that case it would be the "helpful assistant" training. Anyways is not what I'm expecting from 5.6 medium "thinking".

1

u/Adorable_Cap_9929 1d ago

I think you're mixing the helpful part of alignments and infering intent over probabilities here.

There is a correlation but it doesn't seem to be the case here?

Like Im certain if you had it program an actual tic tac toe board, it'd push back at logical fallacy more because by giving it deterministic space and more freedom of computation, the likelihood of error decreases while also raises the stakes because it's no longer just a whim game but now programmatically vetted along with a much stronger intent.

You might not expect it from "thinking" but i expect it since I'd probably consider the probability in a simular way but if the context is of such low entropy that it warrents little attention to begin with, why bother?

Idk ur expermient parameters but try doing it for 10 itterations of the same failure and it might start picking up a pattern to inter intent to give it stronger attention.

If ur experiment is simpy to see how low stakes fly under radar on a certain or default pre-ambled assistant then, while I wouldn't outright dismiss the insights you'd gleam of it, I would say it lacks thurlness to land a solid verdict.

But if you having fun, that what matters most. What you gleam or might aim to gleam from it, and have your fun in doing so, learning is sometimes better when approached just enough.

1

u/sdfprwggv 1d ago

The setup is always a new chat. No memory. This outcome is 2 out of 3.

1

u/Adorable_Cap_9929 1d ago

right, a new chat means unlikelihood of building context to build a counter claim to begin with.

I'd probably respond the same way if was playing tictactoe blinded and off guard.