1 pointby Zeruxean hour ago1 comment
  • Zeruxean hour ago
    Jev and Laya answer typed questions about text in one forward pass.

    Their confidence scores aren't error rates: a model can say "critical, 78%" and be wrong, and nothing says which answers to trust.

    Zet sits on top and marks each answer sure or unsure. Sure answers are meant to stay within an error budget you choose (5% here), set on labeled examples and tested on held-back ones.

    Benchmark: MASSIVE 1.1 (human-labeled), six-topic task, 350 held-back requests. Both rows use the same Laya predictions.

    - Laya alone: 350 used automatically, 45 wrong. - Laya with Zet: 227 used automatically, 4 wrong (1.8%), 123 sent to review.

    Zet flagged 41 of Laya's 45 mistakes (91%) while leaving about two thirds of requests automated. The catch: 82 of the 123 reviews were for answers that were already right.

    What it doesn't show: - The six topics were chosen as an easier task. With 18 topics, Laya made 228 mistakes on 602 requests and Zet sent all 602 to review: no automation, but no automatic mistakes. - The cautious 95% upper bound on the automatic error rate was slightly above 5% (5.2% and 5.7% per language), even though the observed rate was lower.

    This project is under development, I just would like to share it for those of you who want to test it or help me contribute, there is still a lot to be done!

    Site: https://yoosseph.github.io/Zet/

    Feedback welcome.