4 pointsby sthottingal15 hours ago1 comment
  • svcrunch2 hours ago
    Thanks for sharing this writeup.

    At my previous company, I led the development of a retrieval embedding model called Boomerang, which focused on minimizing same-language bias. Briefly, the way most embedders are trained, they will prefer worse answers in the language that the question was asked, versus better answers in a different language. I call that same-language bias.

    Anyway, an investor at a16z was quizzing me on why this particular model wasn't top-ranked on MTEB. I tried explaining that those benchmarks are gamed and they don't reflect real-world performance well. I could tell he wasn't buying it.

    When I saw Jev raise 700m, I really felt that the VC funding model is creating an unserious environment for development of AI/ML. What's rewarded is slick marketing, insider connections, and confidently expounding unrealistic visions of what your team will achieve in the next 3 months.

    Under these conditions there isn't much incentive to fix benchmarks.

    I remember my manager in Google Research poring through a German-English translation dataset and realizing there was a ton of invalid data inside it. When he filtered out the cruft, models showed big improvements. But the filtering was laborious and not glamorous.