2 pointsby delis-thumbs-7ean hour ago3 comments
  • spottedmarley24 minutes ago
    I built my own benchmarking arena that tests local models on all of things the I need a model to do well. I don't look at any of the existing benchmark data that is out there. When a new model drops, I run it through my arena and see how it compares to previous models. If I talk about a model being good I am referencing my own accumulated knowledge on how a model performs for me on tasks that I care about. I generally will never be heard talking negatively about a model (except maybe a frontier/hosted model, they all suck in their own ways) because if a model sucks it just gets deleted and I move on to other things. I suppose I'd consider myself somewhat of an 'expert' when it comes to analyzing local model performance, but I don't really listen too much to what anyone else says about them, or which benchmarks tell them which things about a model. Just test them on the things that are important to you.
  • tolugeniusan hour ago
    There is probably far, far more people trusting benchmarks and "I remade x thing in 1 prompt with y model" claims than you'd imagine, just ignore all of it. You know your workflow and what better should be and could be, measure on what works for you. You should note (and I may be wrong, not active in these part) a lot of those demos are very toy, recreating a known game, known app, known workflow, etc. very interesting but again a very toy example that should be taken with that in mind.
  • bigyabaian hour ago
    Benchmarks can still be useful, even for benchmaxxing labs. For example, the recent Beam model benchmarked much worse than DS4.1 Flash and GLM 5.3, both of which perform extremely well outside of benchmarking. Regardless of whether or not Beam was benchmaxxed, it's performance suggested that it wasn't capable of solving problems that other models in it's weight class could do easily.

    The best-case scenario is that Reflection was being honest about their model's disappointing performance. The worst-case scenario is that they benchmaxxed, and it still managed to underperform compared to it's peers.