20 pointsby flifenstein2 hours ago3 comments
  • michalpleban11 minutes ago
    I would love to see such comparisons but with quantized versions, because quantization allows running models on smaller hardware with some quality loss. I am running Qwen3.6-35B-A3B quantized to int4 on an A6000 card that was otherwise just sitting around idle. It works up to a degree, but I would love to see benchmarks comparing different quantizations of several models, especially in quality. This is an important dimension in deciding whether to buy a GPU is worth it, and it is missing from this (otherwise comprehensive) article.
  • Lord-Jobo11 minutes ago
    >which box to buy

    > spark, costs less than a conference trip.

    I know putting actual prices regionally localizes your article and temporally, with how prices are so unstable. But analysis of “what to buy” without actual prices is borderline meaningless.

    Overall, good article, very interesting to see a real deployment that’s actually attainable and not just a subscription to a big 3 token plan.

  • flifenstein2 hours ago
    [dead]