> Stated findings
> Findings derived from two curated layers: which model cards mention each benchmark, and which scores could be read verbatim from those documents. Each finding names the evidence behind it.
This is 100% AI generated but the problem is it's difficult to understand - what layers is it talking about, is it the llm model layers or what.
It means two “layers” of information access/validation.
Sounds like it first made a little map of model <-> benchmark, then went and filled in the score boxes.
Definitely not LLM layers
What the hell is the point of this page? Can you put in a single bit of human prose explaining why it exists and what we are supposed to learn?