29 pointsby MiguelG7195 hours ago15 comments
  • sharath394 minutes ago
    Nicely built.
  • MiguelG7195 hours ago
    Hi HN,

    Over the last 2 years, we observed computer use models improving at a rapid pace and saturating benchmarks. This new benchmark replaces Online-Mind2Web with our own Browserbase Benchmark v2 that better represents the complex tasks that browser agents face in the real world. It runs against 23 models (frontier and open-weight) and 9 harnesses (Claude Code to LangChain Deep Agents) on accuracy, speed, and cost.

    This new benchmark confirmed our belief that the choice of an harness is becoming as important as the choice of a model. For example: claude-opus-5 runs 74% at $1.50/task on LangChain deep agents but 71% at ~$10/task on fx.

    The eval harness is a CLI you can run yourself (pick harness + tools/mcps + model, pass high-level tasks, grades with LLM verifiers, has trials/concurrency/OTEL tracing): https://github.com/browserbase/stagehand/tree/main/packages/...

    Happy to get into methodology, and if you want your model or harness added, just let me know.

    • alyssamaru5 hours ago
      Yes, would love to hear the methodology!
  • devk034 hours ago
    Hyper personalized harnesses are the edge that the labs cannot beat startups on. There will be a whole entire era of new harnesses coming out soon.
    • dericdinudaniel4 hours ago
      a whole new era of custom harnesses with unnecessary stuff trimmed out sounds like the future
      • MiguelG7194 hours ago
        It still feels like the harness and the model need to co-evolve together
  • smpandya5 hours ago
    Are there results comparing agents running different tools (agent-browser, playwright MCP, browse CLI), or is this mostly Stagehand focused?
    • MiguelG7194 hours ago
      You can swap the driver/tool yourself with the evals CLI! Just use `—tool-surface <one-of-the-supported-tools>`. Or define your own using the interface
  • pranaygup124 hours ago
    What model family do you find is the best for browser use overall? or does it change pretty regularly
    • MiguelG7194 hours ago
      It changes so frequently, and the world wild web is vast so it’s dependent on your use case. In general the leaderboard reflects what we see working across a broad range of domains, but the best way to tell is to define your tasks and run the evals yourself; that’s what this is for
  • mhykim5 hours ago
    Why use Stagehand when agents can write CDP / Playwright on the fly for browser use
    • MiguelG7194 hours ago
      Token efficiency, performance, observability, and most importantly: permissions/security policies
  • peeet5 hours ago
    I love the website, it would be 10/10 if I could go to chrome://dino
  • jay_sahnan3 hours ago
    Why invest in browser agents if computer use like Astra is already so good at solving tasks?
  • dericdinudaniel4 hours ago
    Cost difference across harnesses is interesting to me. Would love to see more info about more optimizations in the harnesses to trim down costs.
  • starlightttt5 hours ago
    Love the design
  • 5 hours ago
    undefined
  • 4 hours ago
    undefined
  • vishalanton4 hours ago
    [flagged]
  • 4 hours ago
    undefined
  • vishalanton4 hours ago
    [flagged]