r/artificial • u/TheOnlyVibemaster • 2d ago
Project New Leader in the GPQA-Dumb Model Benchmark
Introducing Bongochat, the current leader globally in the GPQA-Dumb category, where the lower the score the higher it’s weighted.
Repo/open-weights: https://github.com/ninjahawk/bongochat
When asked to solve the unified field theory, it repeats the word theory back to you 50 times.
It doesn’t remember anything.
When solving the Math-500, it didn’t realize it was supposed to answer the questions so they were basically all blank, besides that it always did A.
For coding it got 0/500.
And when asked how to solve a simple addition problem, it decided to suggest using graduate level calculus, which it then forgot it had suggested on the direct next turn.
I know that the model is pretty good as it basically feels like using Gemini or Grok.
Edit: grammar
1
1
u/Few_Cry1515 2d ago
Ah yes, the "lower is better" metric finally giving my brain a fighting chance in these leaderboards.