Homa: The end of TCP for AI clusters [video]

(youtube.com)

39 points | by signa11 3 hours ago

8 comments

  • Animats 7 minutes ago
    Homa has been around for a while. Here's the 2018 paper.[1]

    The core idea: When a message arrives at the sender’s transport module, Homa divides the message into two parts: an initial unscheduled portion (the first RTTbytes bytes), followed by a scheduled portion. The sender transmits the unscheduled bytes immediately, using one or more DATA packets. The scheduled bytes are not transmitted until requested explicitly by the receiver using GRANT packets.

    So it sends blind for short requests, then needs a go-ahead from the receiver. That's reasonable when the main application is a remote procedure call. It's reminiscent of QNX's networking protocol, which is also single packet message request/response but can also handle arbitrarily long messages.

    What makes this work today is that per-packet processing overhead in hardware switches is low vs. per-byte overhead. In early software driven switches, per-packet overhead tended to dominate, and sending small packets was very inefficient. In modern hardware switches, where FPGAs are doing the processing, the per-packet overhead is low enough that small packets are not inefficient.

    It's amusing that web stuff is so bloated today that any transaction under 1MB is considered "small". So this is not a suitable protocol for open web use.

    [1] https://people.csail.mit.edu/alizadeh/papers/homa-sigcomm18....

  • giovannibonetti 1 hour ago
    I remember listening to Jane Street’s Ron Minsky on their podcast talking about this a few months ago, how TCP becomes the bottleneck in AI clusters. As an electrical engineer, I remember that circuit switching gave away to packet switching due to very sparse usage of the network when there are many actors going through it. It is not very efficient, but that wide variety of traffic makes it hard to optimize it since the flow patterns are too dynamic. A good analogy with car traffic is that downtown there are so many cars going to a large variety of places, that traffic lights – as inefficient as they are – are a solution that at least works good enough.

    On the other hand, if the traffic follows a very predictable pattern, a custom implementation can be much more efficient. Specially nowadays machine learning can find much better solutions through reinforcement learning. And AI cluster data flow is much more predictable than what goes over the internet as a whole.

  • wmf 31 minutes ago
    I've been hearing about Homa for years and wondered what's new. I found a changelog in the readme of the git repo: https://github.com/PlatformLab/HomaModule
  • adastra22 1 hour ago
    Please don’t make the primary link a video.
  • jMyles 48 minutes ago
    I wonder if, as LLMs get accustomed to using homa or some other optimized protocol for, as the article lists, "chores such as weight gradients, model weights, KV cache entries, and checkpoints", whether we'll start to see TCP as a bottleneck for their post-trained interactions as well, for many of the reasons.
  • almost_usual 2 hours ago
    • dang 1 hour ago
      Thanks, those are great! Added to toptext as well.
  • paradiselord-de 2 hours ago
    [flagged]