The basic idea is that model-facing tools are categorized into safe and risky tools, and every risky tool requires human in the loop. The categorization is done when you create a service integration, so it doesn't require a model to follow instructions well or be aligned for the guardrails to work. We built it carefully so that we were comfortable having the model read our emails and knowing that there's nowhere it can send the information out without user approval.
Beyond that core idea we did a lot of work to make the harness legible to non-technical users. We never ask people to approve a scary looking shell command (not like most people know what to do with them anyways).
And then of course its open source. Plug in your own models (local inference supported). Write your own plugins to connect to your own private services. And its self-hosted, so you get to own all the history and know how that accumulates in your harness.
Would love any feedback from the HN community.