3 pointsby svcrunch7 hours ago1 comment
  • svcrunch7 hours ago
    The [Pelican Benchmark](https://github.com/simonw/pelican-bicycle) is in the LLM's training data and probably not a useful indicator of improving LLM capabilities any more.

    For the past two years, I've run the Little Dorrit Editor Benchmark. Typesetting is my hobby and I wanted to see how well LLMs could extract editorial marks from a printed page.

    The initial results were not encouraging, but performance has risen rapidly since the summer of 2025. From F1 scores in the low 0.2s in 2024, we are now at 0.78 with GPT 6 Astra! Fable 5.1 scores 0.73 (high thinking mode for both).

    The best thing about this benchmark is that there is still plenty of room to climb.