4 pointsby zodwick4 hours ago1 comment
  • zodwick4 hours ago
    Author here, Life lately: build a hard eval → new model mogs it → build a harder eval → mogged again → repeat. Cafe Bench is the latest round this, it is a small eval ( world sim ) we built at Dot to test and track progress of the product.

    But the data points transfer well into generalised model benchmakrs.

    surprises: Opus 5.5 made +$227k for $8.57 in 22 minutes; GPT-6 Astra made +$157k but cost $56; and GPT-6 Luna lost $60k on average.

    • zurfer3 hours ago
      really interesting results, it seems we really crossed another threshold with yesterdays releases where sol 5.6 still lost money, but sol 6 is in the green.