We use a routing approach where simpler tasks go to smaller, cheaper models and more complex tasks go to frontier models. We are also working on having our own model for some of these workloads.
Document processing is probably one of our biggest agent use cases for tokens at this scale.