Fair Usage Policy
How we keep Osiris fast and affordable for everyone.
Osiris provides access to a wide catalog of AI models through a single API key. To ensure reliable performance and fair access across all users, certain high-demand models may have a token usage limit per billing period. This policy explains how it works and what to expect.
What is a token limit?
Some models have disproportionately high upstream costs or limited capacity. To keep pricing accessible without restricting model availability entirely, we apply a per-period token budget for specific models or model groups.
- Token limit = maximum total tokens (input + output) you can use on the specified models within one billing period.
- The limit resets automatically at the start of each billing cycle (aligned with your plan's longest quota period — typically monthly).
- Multiple models may share a single token pool. For example, if Models A and B share a 100M token limit, usage of either model draws from the same budget.
Which models are affected?
Only a small subset of models have token limits. The specific models and limits depend on your plan tier. You can always check your current usage and remaining allowance in your Dashboard.
What happens when the limit is reached?
When you exhaust the token budget for a limited model group:
- Requests to the affected models will return a
402 model_quota_exhaustederror until the next billing period. - All other models remain fully accessible. Your request quota, rate limit, and access to non-limited models are unaffected.
- You can switch to any model that is not in the limited group and continue working without interruption.
How to check your usage
| Where | What you'll see |
|---|---|
| Dashboard | Per-model token usage with progress bar, remaining allowance, and percentage used. |
| Telegram Bot | Use /balance or /mykeys to see remaining tokens per model group. |
| API Error | A 402 model_quota_exhausted response includes the model name and current/max token count. |
Need more tokens?
Higher plan tiers come with larger token budgets for limited models. You can upgrade your plan at any time from the Top Up page.
| Plan | Token Budget (limited models) |
|---|---|
| Lite | 50M tokens / period |
| Starter | 100M tokens / period |
| Pro | 400M tokens / period |
| Max | 1B tokens / period |
| Enterprise | 2B tokens / period |
Token budgets apply to the shared model group. Non-limited models are always unlimited within your plan's standard request quota.
Why do we have this policy?
- Sustainability: Some models cost significantly more upstream. Token limits allow us to offer them at accessible prices without restricting access entirely.
- Fairness: Prevents a small number of users from consuming disproportionate resources, ensuring consistent performance for everyone.
- Availability: By managing demand on high-cost models, we can maintain low latency and high availability across the platform.
Questions?
If you have questions about this policy or need a custom arrangement for high-volume usage, reach out on Telegram.