1 pointby anitroves5 hours ago4 comments
  • bydhruvil5 hours ago
    First, check what LLMs your system can actually handle based on your RAM/VRAM, then choose the model variant accordingly.

    If you want to use the model through an API for your website, a good setup would be a dedicated machine/server to host it, since running LLMs locally can consume a lot of memory.

    Depending on your hardware, you can look at smaller variants of Gemma or Qwen Coder.

    Another option is to use a hosted inference provider. It’ll cost a little, but you’ll usually get much faster inference compared to running locally especially if your system isn’t high-end.

    • anitroves4 hours ago
      What should be the minimum hardware requirements I got Nvidia processor with a high quality graphic card and 16 gb ram
  • mrsalty4 hours ago
    go to https://huggingface.co/settings/hardware?next=%2Fmodels and add your hardware, it will show models you can run on it
  • jaggs4 hours ago
    Try TurboLLM (https://turbollm.dev/).
  • andy_parhelia38 minutes ago
    [flagged]