34 pointsby signa113 hours ago6 comments
  • giovannibonetti30 minutes ago
    I remember listening to Jane Street’s Ron Minsky on their podcast talking about this a few months ago, how TCP becomes the bottleneck in AI clusters. As an electrical engineer, I remember that circuit switching gave away to packet switching due to very sparse usage of the network when there are many actors going through it. It is not very efficient, but that wide variety of traffic makes it hard to optimize it since the flow patterns are too dynamic. A good analogy with car traffic is that downtown there are so many cars going to a large variety of places, that traffic lights – as inefficient as they are – are a solution that at least works good enough.

    On the other hand, if the traffic follows a very predictable pattern, a custom implementation can be much more efficient. Specially nowadays machine learning can find much better solutions through reinforcement learning. And AI cluster data flow is much more predictable than what goes over the internet as a whole.

  • 41 minutes ago
    undefined
  • adastra2229 minutes ago
    Please don’t make the primary link a video.
  • jMyles8 minutes ago
    I wonder if, as LLMs get accustomed to using homa or some other optimized protocol for, as the article lists, "chores such as weight gradients, model weights, KV cache entries, and checkpoints", whether we'll start to see TCP as a bottleneck for their post-trained interactions as well, for many of the reasons.
  • almost_usual2 hours ago
    • dangan hour ago
      Thanks, those are great! Added to toptext as well.
  • paradiselord-dean hour ago
    [flagged]