2 pointsby bassamtabbara7 hours ago1 comment
  • sumbry2 hours ago
    Having spent a chunk of my career in Platform Engineering, I feel what platform teams are about to go through. They're going to get stuck running self-hosted Inference workloads and it's going to come out of the blue. The training/inference worlds are completely different than what we're used to with very deep complex stacks and technologies that you'll have to come up to speed with very quickly.

    The catalyst will probably come from outside of Engineering; finance asking for ways to control token spend, Security asking for more control over our data, even engineers asking for more control over model output (and stability of tools/harnesses,etc).

    Either way the day of reckoning is coming. We're doing this ourselves at Upbound and decided to share our journey along the way. I do think SRE for Models is going to pop-up at some point as we learn how to reliably operate these things at scale but as we've already learned, getting the model up and running is the easy part.