1 pointby matildagh5 hours ago1 comment
  • matildagh5 hours ago
    At Sonata, we are creating sandboxes of real companies. We build custom digital twins of your Slack, Gmail, and all of your internal systems to mock what the agent would do. This is very similar to methods used by DeepMind/Anthropic/OpenAI researchers for predeployment testing (SOTA evals).

    1. Create digital twins with internal systems. 2. Sync them with the real production data. 3. Give the agents read access to the production data. 4. Based on the production data, use LLMs to do scenario testing.

    If you are a business that is buying or deploying the agent internally, you can stress test first and get a benchmarking report. Then can identify which use cases the agent is reliable enough to do.