1 pointby alexellisuk3 hours ago1 comment
  • alexellisuk3 hours ago
    Having been frustrated with benchmaxxing figures shared on X (as described here: https://blog.alexellis.io/how-and-why-we-bought-4-dgx-sparks...), I developed an independent benchmark to get stable and comparable numbers for Prose, Code and Structured output, plus Prefill.

    I didn't expect it - but the majority of folks building recipes for DGX Sparks on X/Twitter are now adding/using RigMark as one of their primary benchmarks when sharing figures.

    We've gone from a random claim of "I get 65-80 tok/s with this model" to actual receipts that can be traced back to a particular SHA/version of RigMark.

    We're on the third version now, and PRs / feedback is welcome.

    Where I think we could all use some help going forward is on quality - running local AI sometimes means using compression of weights, which is lossy.