For the past two years, I've run the Little Dorrit Editor Benchmark. Typesetting is my hobby and I wanted to see how well LLMs could extract editorial marks from a printed page.
The initial results were not encouraging, but performance has risen rapidly since the summer of 2025. From F1 scores in the low 0.2s in 2024, we are now at 0.78 with GPT 6 Astra! Fable 5.1 scores 0.73 (high thinking mode for both).
The best thing about this benchmark is that there is still plenty of room to climb.