I was inspired by the really neat "Jev in 25 lines of Python" post a few days ago (and System One models in general), and after playing with the idea I tried whipping up something a little more complex, and Lichen is what popped out: https://github.com/Mushroom-Systems/lichen
My question after reading that post (and I know I wasn't the only one; antirez came right out and said it in the comments) was whether Google Research's prompt repetition paper for non-reasoning models (https://arxiv.org/abs/2512.14982) would improve the outcomes. Turns out it did a fair bit. After some experimentation I added two more tricks: the options are listed a second time in a rotated order to wash out the model's preference for position, and a letter remapping sums the probabilities over each option's letters. Both help on the harder and more ambiguous questions.
After testing several models I had lying around, I settled on gemma-4-26B-A4B as my daily driver. On public JevBench, Google's QAT build scores 207/231 against Jev 1.13's 200 (not statistically significant yet; I'll be submitting it for evaluation on the withheld items soon). Served from vLLM with the NVFP4 build, it scores 204 with a median response time of about 56 ms on my laptop 5090. For folks without that much VRAM to spare, the smaller Gemma 4 models also perform well: E4B fits in about 5 GB, though it does suffer on the harder queries. Qwen also works (my colleague has reported good results with Qwen3.8), but on my setup I found it a bit slow for my use case, and the benefits weren't strong enough to outweigh the slowdown.
Because Gemma 4 is multimodal, Lichen also extends the query format so images can be part of the state (served from vLLM), with answers in about 190 ms. The results are good but not perfect. On 125 Wikimedia Commons photos that I checked by hand, it got every one right. On 98 road-sign photos drawn at random from Commons, it named the right sign 92 times when it had to pick one. When I added "another sign, or no road sign" as an option, it dropped to 73: it tends to decide that a small, distant or vandalized sign isn't there, and it's confident about those misses. The confidence on images isn't well calibrated in either direction, and a separate temperature for images didn't fix it. The README has the numbers and a couple of the failures.
I've been using this as a partial replacement for the backbone of our company's context graph for a few days now, and it's performed excellently so far. Hopefully folks will find some use for it.
(Sidenote: Big thanks to my friend and colleague Jacob M for his vLLM support patch)