2 pointsby florianherrengt4 hours ago1 comment
  • florianherrengt4 hours ago
    Works with GeForce RTX 20 series and newer, RTX PRO (Turing and newer), DGX Spark and Apple M4 or newer.

    It does not pool memory or split one inference request across machines. Adding machines increases parallel throughput but won't let you run bigger models.