Man I'm so sick of AI benchmarks lol. You're telling me GPT-6 got 98% on the PissX Ultramark? Well that plus the freaking 812 they pulled on the FPR EufRA EEE 4.2.3 means this thing is really gonna rip. I gotta say though I really was expecting a bigger improvement on the BSMT:3 (Big Scawy Math Test :3). Then you look up what they measure and it's always like "The PissX benchmark suite measures how efficiently a model can manipulate an image." Oh ok thank u! Something something Goodhart's law
@theodoraward remember when the first thing everyone learned in every "introduction to ai" class was that you should always separate your training data and test data
@operand @theodoraward mixing them saves a lot of time though so who's to say if it's good or bad
I am 100% certain this would be at least slightly less annoying if I was less lazy but somewhere around the 10th model announcement where I scrolled down and the website did that thing where it all fades from section to section to reveal a row of vertical bar graphs showing a final bar slightly higher than the other bars with a headline like "Unprecedented performance. Superb efficiency." i was like Ok dog but like can it actually do stuff better. Like actual stuff