1 pointby marojejian5 hours ago1 comment
  • marojejian5 hours ago
    >On a composite of all three dimensions, a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009)

    note this was GPT-5.2