1 pointby sourav_biswas4 hours ago1 comment
  • minimaxir4 hours ago
    No, because the cache is local to the GPUs and there would be no cost/compute reason to have long-lived caches outside of the 1 hour cache already done by LLM APIs.
    • sourav_biswas4 hours ago
      What if the real business isn't caching generic prompts/but caching outputs for a specific vertical where one can know two requests are truly equivalent n not just embedding similar?