several years ago, using 6 threads and vectorized C++, we saw about 240-270 MP/s decode speed. Hence this should be possible in < 100ms. Not sure how the current implementation differs from that state :)