15 pointsby codyznash5 hours ago3 comments
  • prasadvara5 hours ago
    This is great writeup, can we do a cross comparison with "human" experts?? whether models perform better OR worse??
    • LambdaComplex4 hours ago
      This writeup reads like it was written by Claude, which makes me immediately question its accuracy.

      > Each dot is one code sample; bars mark the median. The split between malicious and benign packages, perfect before the prune, was perfect after.

      People do not write like this.

  • cortesoft4 hours ago
    I feel like LLMs really highlight the ambiguity and imprecision of the english language in normal use. LLMs are getting really good at guessing what we mean, but it is still a guess.
  • 5 hours ago
    undefined