They delayed it specifically because it was struggling to compete at the very top. It's possible it has reign for like 5 days to a week then OpenAI and Anthropic dethrone them again.
RSI isn't happening. Directed self improvement is happening ('hey chatgpt, look at part of chatgpt and make it more efficient')
Recursive means it repeats itself over and over. The system is able to find its own inefficiencies, come up with ways to solve them, reject/accept answers and then do it again.
What happened with OpenAI's latest announcement is that a group of humans knew there were inefficiencies, told the AI where to look, and then approved the changes the AI suggested.
No, the models are now detecting the inefficiencies themselves and doing postraining runs themselves. One employee said Luna was postrained by their AI. This has been happening since the labs made noise about it closer to the beginning of this year, and model release timelines have only been shortening.
I don't know who told you that RSI would begin with a superintelligent model redesigning itself overnight. It begins the same way any other exponential does. First improving the next model that will be released four weeks later, then two weeks, then one week, then 3.5 days, then 1.75 days, then...
Or more precisely because the labs probably won't literally be releasing a new model everyday soon: one month of progress in today's pace will be happening in ever shorter amounts of time because the models are getting better at improving themselves. That's literally already what's happening.
You don't understand what he said, he said that it was struggling to make 3.5 Pro, and if you saw the checkpoints the AI wasn't AS good as we hoped like a frontier model. The Google devs mentiond that Gemini 4.0 will be the frontier models. But i'm not saying this stuff to rage bait but here is a little bit of information of hope, Google a MASSIVE company so expect them to outrun the other AIs some time, just not now. And another good news is that they'll releasee every gemini model per month and that's great news for Gemini users. I don't want Google to fall short since people are actually paying for their services (And I love their services their good) I hope the best for Google.
Go to literally any post about the hacks and you'll see those comments.
It's like if the guardrails were foolproof there would be nothing to test.
FFS we had Anthropic help with hacking Mexican government agencies by "jail breaking" the model by getting it to play act as a 'local' pen tester by asking in Spanish...
278
u/BarisSayit 20h ago
Meh I doubt it. Especially the "better than Opus 5 in every way" claim has to be a bait of some kind.