80 pointsby anerli2 hours ago17 comments
  • lxe2 hours ago
    On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.

    Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.

    Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.

    • anerli6 minutes ago
      Yeah we heavily leverage coding agents for optimizing our kernels. Since it's highly verifiable and takes time to measure we often leave multiple running and improving performance on different model architectures.

      Definitely still helps to reference relevant academic work as well, or even just encouraging the agent to make bigger structural leaps, otherwise it will often get stuck working on low impact micro-optimizations.

  • MaxikCZ6 minutes ago
    if fully custom compiler would find best settings for given setup, upload the setup to mothership and allow new peers to download it as good starting point.

    Can it do all the shenanigans that allows to run qwen flash on 12GB vram over 40 toks like people seems to be getting in this thread?: https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38f...

  • c7b27 minutes ago
    Cool idea! Do you happen to have benchmarks for Strix Halo (AMD Ryzen AI Max+ 395)? I take it that Qwen3.8-Flash-Next is not supported?

    And a more general question: does your engine detect and optimize for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.

    • anerli10 minutes ago
      No specific benchmarks for Strix Halo yet but planning to release more results for different hardware and models soon!

      Qwen3.8-Flash-Next support will also be added very soon.

      Taking full advantage of all the hardware on your machine in the most performant way possible is the overall goal of the inference engine. This includes a lot of what you're describing. We want to map out the full hardware topology of your system (one or more GPUs, CPU, memory), and compile a combination of kernels to serve a given model optimally across that stack, allocating different parts of the workload wherever it fits best.

      Currently we're writing tunable kernels that optimize themselves for one device, but we're working on a kernel compiler that will be able to compile and distribute kernels across any number of devices in a system.

  • kmike842 hours ago
    This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

    I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.

    3 main failure modes I observed in the engines:

    * Not using best available spec decoding

    * Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)

    * Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

    • anerli2 hours ago
      Yeah these are all things that we directly tackle!

      Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).

      Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.

      Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.

      • skohanan hour ago
        Do you have anything published on the quality benchmarking using your caching strategy?
    • an hour ago
      undefined
  • mncharityan hour ago
    Fwiw, top of my own pain-point list (I suppose given the first item, that's a pun) includes:

    External/policy-based throttling for temperature control. Unthrottled, my laptop bottom goes skin-burn hot. But fixed compute caps can have non-linearly dreadful performance impacts in particular cases. Plan is a runtime knob, to replace manual limits-kludgery.

    I'll use models which barely fit in VRAM+RAM, and are order-1 tok/s slow. So tool call step overhead can be painful - a world where `ls` costs tens of seconds. Plan is blending harness plugins with inference loop, for "no, don't stop - I already have the call result for you - just keep going" (and also some logit games).

  • herf2 hours ago
    I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti).

    Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:

    set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0

    • anerlian hour ago
      Thanks for reporting the issue.

      Currently we don't support multi-GPU setups, that is on our near-term roadmap. It saying the model is too big for that GPU might be a bug - would you be willing to open a github issue with more detail on your setup? https://github.com/magnitudedev/magnitude/issues

      As for performance, there may be some variability still depending on the model and backend. We have room for improvement for various setups that we are closing as we work out some details with our kernels and tuning system, so appreciate the data point and will look into that combination.

  • cedricd35 minutes ago
    Looks interesting! Is there any way to skip or speed up the 'Assessing Models' step? I'm unable to download anything because it's been taking forever. I'm sure you could apply some quick heuristics or do a lookup or something to filter models. Or trust the user a bit more -- I already know which models fit on my machine. As it stands I'm stuck at that step and can't use the app.

    Maybe have it run silently in the background and assess on demand when a user selects / attempts to download a model. It's not quite clear why all need to be assessed before I can download the first model to try.

    • anerli21 minutes ago
      Thanks for reporting this issue - assessing is not supposed to take more than a minute or so. This is not strictly necessary but filters out models that don't fit in memory and gives speed estimates. This should ideally be a very short step so skipping hopefully wouldn't feel necessary if we patch this.

      Could you share your hardware and OS details to help us identify what might be the issue here?

      There's also a github issue open on this topic if you want to leave a comment there: https://github.com/magnitudedev/magnitude/issues/142

  • nateb20222 hours ago
    Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
    • anerli2 hours ago
      The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.

      For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.

      We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.

      The source for the benchmark is available here also: https://github.com/magnitudedev/magnitude/tree/main/inferenc...

    • anerli2 hours ago
      Compared to MLX - we've done some rough benchmarking and we are outperforming any of the MLX-based engines we've compared to so far. Going to do more in depth benchmarking and release it soon.
  • msdz2 hours ago
    Congratulations on the launch, it looks like an impressive product and tool!

    Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is model- or at least architecture-specific code for most, if not each new open-weight model coming out. Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?

    • anerli2 hours ago
      Yeah, generally being able to focus on specific architectures lets you optimize better for those. However models of the same family (for example Qwen 3.5/3.6/ some 3.8 models) share the same architecture, so you only need to optimize once and new models can use the same kernels. There's also shared algorithms and kernels that can be optimized once and used across different families, so it's a bit nuanced.

      We plan to support any model architecture that we believe is somewhere along or close to the pareto frontier. There's some model families that are outdated or more niche that we don't necessarily want to put our focus into.

  • teabee892 hours ago
    How does this compare to ZML's llmd https://zml.ai/llmd/ ?
    • anerli44 minutes ago
      From the looks of it, this seems focused on datacenter/batch inference, and doesn't tune its kernels to the specific hardware and workload where inference is being run like Magnitude does.

      Magnitude is optimized for maximum single-session performance and memory efficiency - so we should be more performant for local inference use cases.

  • hypercube33an hour ago
    From your description looks like this isn't for AMD or Strix Halo at all? Also one of the things I'm not sure of but definitely plays a huge factor is the variant of the model you download - how does this help select the fastest version for your specific hardware / context size?
    • anerlian hour ago
      We support Vulkan as well, we just didn't mention it in the benchmark. When AMD or Strix Halo is detected the engine will use Vulkan.

      Regarding model variants - our catalog includes different quantizations, and automatically assesses these against your hardware to determine which ones will fit in your memory and how fast they will run. This lets you pick a model to download based on your desired speed/intelligence tradeoff.

      • skohanan hour ago
        Do you have any plans to support ROCm?
        • anerlian hour ago
          We are actively benchmarking our Vulkan kernels to ROCm implementations in other engines to ensure that we can reach the performance ceiling with them. Vulkan is much more portable and also works on non-AMD hardware even though it can be more awkward to write kernels for. If we find that Vulkan is not sufficient for reaching the same performance as ROCm, we'll consider adding it as a backend
  • paulgerhardtan hour ago
    Trying to run this but keep hitting bugs. Can you open up issues reporting on your repo?
  • amirhesham2 hours ago
    Oh this is so cool. Curious about the business model, too.
  • sgtwompwomp2 hours ago
    This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
    • anerlian hour ago
      I would say the overall idea of trying to achieve performant inference for agent workloads is the strongest commonality with Wafer.

      It's not a coding agent running on your device optimizing the kernels, we have a system for writing kernels that can be tuned on the target device automatically. So we write the efficient high level kernel structure with tunable parameters, then it fits to whatever hardware it's actually running on.

  • kenzic2 hours ago
    How long does tuning take (on an M3 MacBook Pro for example)?
    • anerli2 hours ago
      Tuning is a one-time process that takes around ~1 minute whenever you download a new model. This is generally enough time to tune all the kernels' parameters to the point where tuning any longer asymptotes. Time can vary a little based on the hardware though.
      • kenzic2 hours ago
        Wow, that's impressive.
  • yolandac2 hours ago
    does it allow us to run larger models that weren't possible before?
    • anerlian hour ago
      Right now, since we use less memory for KV, you have more room for model weights when you're running longer sessions.

      However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.

  • p-e-w2 hours ago
    What is the business model?
    • anerli2 hours ago
      We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.