I initially built it to see how much I could get the harness to make small models, especially at low quantization and context to not feel terrible to use. So while Reika doesn't solve the intelligence side (it never will), it tries to solve the overall experience when using small models at the absolute scale.
A lot of the testing and pain came through working on my M2 MacBook Air 16GB trying to run models like Qwen3.6 35B A3B and Qwen3.8 27B all day in agentic coding, maxing out the RAM and limits of my own machine. So the base of Reika comes from a legitimate source of truth.
You can also plug in an API key for those with hybrid setups too.