How are you managing rate limits and what model are you using? if you would like to disclose that.
Like I alluded to in my initial comment, I've never understood why the inference providers of today charge per token. Whichever marketing department got our industry to be ok with being charged for the output of a REST call, I applaud.
I'm not sure it will all hold up at scale, but I am still waiting for the test to break. It helps to be a curious and optimistic person in this endeavor.
Not currently implementing any rate limits. Using an assortment of open models.
Would love feedback, and testers on the site!
It may not turn out to be a perfect science, but in the name of shipping something and getting feedback, it's out there
As for how it works, its a pretty standard inference provider, where we offer a chatbot and an API, at unlimited usage for a fixed cost. Since its very new, we're still wondering if there's something we're missing as for why fixed cost AI is not the norm.
Hoping it is similar to the music industry, where you once had to pay $1.99 to download a song, and now you can stream all you want for $12.99. We want to begin that new trend.