Update: Looks good so far
Flash-Next on llama.cpp with no MTP: 28.8t/s decode, ~620t/s prefill
on llama.cpp with MTP (froze twice): 44.1t/s decode, ~593t/s prefill
on TensorFold + MTP: 61.0t/s decode, 2,455t/s prefill
That's about 2.1 to 2.4x on decode and about 4x on prefill (!!)
Need to do some long context testing but Im stoked.