We have Jev running as shadow calls alongside with the control calls (LLM models). Agreement rate differ across use cases. We then use this data to tune the confidence and probabilities threshold for each case to achieve around high enough agreement rate with the current model.