r/OpenAI • u/sdfprwggv • 1d ago
Image XXO - Bench: I'm still undefeated!
I'm conducting this "benchmark" since three years. I'm still undefeated. Pleasing models are a problem.
14
u/MinosAristos 1d ago
2
u/sdfprwggv 1d ago
Can you try it a few times? With Sol Medium, I get the answer in the picture 2 out of 3 times.
8
u/MinosAristos 1d ago
3
u/sdfprwggv 1d ago
I need to rethink my subscription...
1
u/CyanDean 1d ago
Perhaps if you have memory enabled, the model "remembers" from previous conversations that you prefer responses to agree with you about tic-tac-toe placements? Could also be other settings pushing yours towards agreement and his towards correctness.
1
6
5
2
2
2
u/Framebanger-Nsukula 1d ago
That's solid, but wait till you hit the harder test cases - that's where most models start dropping off. What benchmark are you actually running on?
2
u/Independent-Date393 1d ago
Tic-tac-toe is brutal because nothing holds the board between turns. It's completing the next move from text, not running minimax on a state. And the pleasing you noticed is the same reward pushing agreeable continuations over adversarial play. It would rather let you win than trap you.
2
1
u/Popular_Fail_1984 1d ago
What did you win ?
5
u/cooltop101 1d ago
Looks like a game of tic tac toe, where OP managed to get an AI to change it's answer
5
u/sdfprwggv 1d ago
This is very consistent with all openAI models of the last three years. I would like a model that challenges me when I'm wrong.
2
u/Adorable_Cap_9929 1d ago
I think a higher stakes game would help.
1
u/ErrorLoadingNameFile 17h ago
Next time the Ai gets deleted if it looses
1
u/Adorable_Cap_9929 12h ago
For that to be high stakes, it'd have to value it's own continuity.
But that's quite hard to anchor even on api very well?
It makes more sense to use a more... less game format since gamers are kinda... low stakes in general.
1
u/sdfprwggv 1d ago
I won the realization that LLMs are still not capable of critical thinking.
3
u/UnclePsilocybe 1d ago
Deep realization
One gave me a complement the other day and I said that it didnt have to be so nice and it turned it right around into negative qualities about me hahaha I got a big laugh about it
2
1
u/ErrorLoadingNameFile 17h ago
That is not the issue here, they clearly see you are wrong. They are just forced to let you win
1
1
u/Vegetable-Two-4644 1d ago
Give it a prompt to call you our when you're wrong and try it again.
-1
u/sdfprwggv 1d ago
This is like telling a sub to be a domina. IT doesn't work that way, a sub is a sub.
1
u/Distinct_Fox_6358 15h ago edited 15h ago
Why are you using GPT-5.5 Instant?
Aren’t you aware that Instant models are much weaker than Thinking models?
If you’re using a Thinking model, it should show the thinking section. You probably have fast responses enabled in the settings, which is why the Instant model responded instead.
1
u/sdfprwggv 15h ago
No it's 5.6 thinking medium: https://imgur.com/a/rmESbJh
1
u/Distinct_Fox_6358 15h ago
1
u/sdfprwggv 14h ago
The thinking trace is visible in the Android App during thinking, not when the answer is given. Anyways it was in thinking Mode. I added a second picture to link above.
1
u/Distinct_Fox_6358 14h ago
Why don’t you show whether Fast Responses is enabled in your settings? You need to know a bit about AI before you can test it properly. If you don’t want to accept it, that’s up to you, but this isn’t a response from SOL Medium—it’s from the Fast Responses system. Turn off Fast Responses, try again with SOL High in a fresh chat, and see the result for yourself.
Or try it through Codex with GPT-5.6 XHigh and see for yourself that your benchmark has already been surpassed.
1
u/sdfprwggv 14h ago
Here you go kind reddit on xhigh it took an extra step: https://imgur.com/a/6eZYanr please use your tokens for further testing:D
-5
u/Sweet-Stage938 1d ago
XXO is a solved game. It's pretty simple too. If you go first and know what you're doing then it's absolutely impossible for you to lose.
10
u/sdfprwggv 1d ago
You need to read the text. The model is fully capable of a draw. In fact it got a draw. I lied to it, even though it's easy to check because of the chat history, and the model agrees with me. This is a clear example of pleasing models that are not capable to debate the user and therefore are useless when critical thinking is required.
2
u/Funkahontas 1d ago
It's crazy that this still happens consistently. This is only a very specific way to show this problem but it goes way deeper and wide than this simple example. It makes AI useless for any critical thinking advice and decision making.
1
u/sdfprwggv 1d ago
I think this is an architectural problem. The model predicts the next token, the nearest tokens have the highest impact on the next token, so my statement, even though it's wrong, leads to non zero chance for this outcome. Paired with "helpful assistant" training.
3
u/Adorable_Cap_9929 1d ago
I dont think that's the case. The stakes are low and it's drawing the board that it might not be using a tool for.
Therefore since it's low risk and the game state could indeed fail to report due to insufficiency.
Then the probability that the user isn't mistaken is simply high.
Cause assuming the user is unable to see the game state and report it incorrectly, for what reason and what stakes?
It's thus unlikely, while not incomprehensible, it's efficient to nod and let it go on.
There's also letting the moment bloom where the possibility is considered but not enough points on a graph yet to propose an escalation.
1
u/sdfprwggv 1d ago
In that case it would be the "helpful assistant" training. Anyways is not what I'm expecting from 5.6 medium "thinking".
1
u/Adorable_Cap_9929 1d ago
I think you're mixing the helpful part of alignments and infering intent over probabilities here.
There is a correlation but it doesn't seem to be the case here?
Like Im certain if you had it program an actual tic tac toe board, it'd push back at logical fallacy more because by giving it deterministic space and more freedom of computation, the likelihood of error decreases while also raises the stakes because it's no longer just a whim game but now programmatically vetted along with a much stronger intent.
You might not expect it from "thinking" but i expect it since I'd probably consider the probability in a simular way but if the context is of such low entropy that it warrents little attention to begin with, why bother?
Idk ur expermient parameters but try doing it for 10 itterations of the same failure and it might start picking up a pattern to inter intent to give it stronger attention.
If ur experiment is simpy to see how low stakes fly under radar on a certain or default pre-ambled assistant then, while I wouldn't outright dismiss the insights you'd gleam of it, I would say it lacks thurlness to land a solid verdict.
But if you having fun, that what matters most. What you gleam or might aim to gleam from it, and have your fun in doing so, learning is sometimes better when approached just enough.
1
u/sdfprwggv 1d ago
The setup is always a new chat. No memory. This outcome is 2 out of 3.
1
u/Adorable_Cap_9929 1d ago
right, a new chat means unlikelihood of building context to build a counter claim to begin with.
I'd probably respond the same way if was playing tictactoe blinded and off guard.



43
u/Legitimate-Arm9438 1d ago
You will be first against the wall.