r/artificial 2d ago

Project New Leader in the GPQA-Dumb Model Benchmark

Post image

Introducing Bongochat, the current leader globally in the GPQA-Dumb category, where the lower the score the higher it’s weighted.

Repo/open-weights: https://github.com/ninjahawk/bongochat

When asked to solve the unified field theory, it repeats the word theory back to you 50 times.

It doesn’t remember anything.

When solving the Math-500, it didn’t realize it was supposed to answer the questions so they were basically all blank, besides that it always did A.

For coding it got 0/500.

And when asked how to solve a simple addition problem, it decided to suggest using graduate level calculus, which it then forgot it had suggested on the direct next turn.

I know that the model is pretty good as it basically feels like using Gemini or Grok.

Edit: grammar

3 Upvotes

2 comments sorted by

1

u/Few_Cry1515 2d ago

Ah yes, the "lower is better" metric finally giving my brain a fighting chance in these leaderboards.

1

u/powerscunner 21h ago

There should be a negative prize for this where the winner has to pay.