Some related work I recommend checking out: - https://arxiv.org/abs/2609.33023 (new last week!) - https://github.com/SREGym/SREGym - https://github.com/hyperdxio/hyperdx/tree/main/packages/hdx-... (Clickstack's version) - https://github.com/grafana/o11y-bench (Grafana's version)
Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?
These are all becoming tablestakes, imo for any AI SRE. The real moat is to build the data layer in an efficient manner so that the investigation (and mitigation) is fast and cost-efficient. The models are making it easier for any company at the same time.