2 pointsby soltanov4 hours ago1 comment
  • spottedmarleyan hour ago
    Been waiting for MTP in llamacpp for a about a week now so I can speed up 3.8:flash-next but this might be just what I'm looking for. :)

    Update: Looks good so far

    Flash-Next on llama.cpp with no MTP: 28.8t/s decode, ~620t/s prefill

    on llama.cpp with MTP (froze twice): 44.1t/s decode, ~593t/s prefill

    on TensorFold + MTP: 61.0t/s decode, 2,455t/s prefill

    That's about 2.1 to 2.4x on decode and about 4x on prefill (!!)

    Need to do some long context testing but Im stoked.