Osiris Docs

Fair Usage Policy

How we keep Osiris fast and affordable for everyone.

Osiris provides access to a wide catalog of AI models through a single API key. To ensure reliable performance and fair access across all users, certain high-demand models may have a token usage limit per billing period. This policy explains how it works and what to expect.

What is a token limit?

Some models have disproportionately high upstream costs or limited capacity. To keep pricing accessible without restricting model availability entirely, we apply a per-period token budget for specific models or model groups.

  • Token limit = maximum total tokens (input + output) you can use on the specified models within one billing period.
  • The limit resets automatically at the start of each billing cycle (aligned with your plan's longest quota period — typically monthly).
  • Multiple models may share a single token pool. For example, if Models A and B share a 100M token limit, usage of either model draws from the same budget.

Which models are affected?

Only a small subset of models have token limits. The specific models and limits depend on your plan tier. You can always check your current usage and remaining allowance in your Dashboard.

Note: The majority of models on Osiris have no token limit— you can use them freely within your plan's standard request quota and rate limits.
⚠️ Models may change: The list of limited models can be updated at any time — models may be added to or removed from the limited group as we adjust capacity and costs. Your dashboard always reflects the current configuration. Token usage from newly added models will count toward your existing budget from the moment they are included.

What happens when the limit is reached?

When you exhaust the token budget for a limited model group:

  1. Requests to the affected models will return a 402 model_quota_exhausted error until the next billing period.
  2. All other models remain fully accessible. Your request quota, rate limit, and access to non-limited models are unaffected.
  3. You can switch to any model that is not in the limited group and continue working without interruption.
Example: If your plan has a shared token limit on Models A, B, and C, and you reach that limit — simply switch to Models D, E, F (or any other available model) to keep working. Your request quota is separate.

How to check your usage

WhereWhat you'll see
DashboardPer-model token usage with progress bar, remaining allowance, and percentage used.
Telegram BotUse /balance or /mykeys to see remaining tokens per model group.
API ErrorA 402 model_quota_exhausted response includes the model name and current/max token count.

Need more tokens?

Higher plan tiers come with larger token budgets for limited models. You can upgrade your plan at any time from the Top Up page.

PlanToken Budget (limited models)
Lite50M tokens / period
Starter100M tokens / period
Pro400M tokens / period
Max1B tokens / period
Enterprise2B tokens / period

Token budgets apply to the shared model group. Non-limited models are always unlimited within your plan's standard request quota.

Why do we have this policy?

  • Sustainability: Some models cost significantly more upstream. Token limits allow us to offer them at accessible prices without restricting access entirely.
  • Fairness: Prevents a small number of users from consuming disproportionate resources, ensuring consistent performance for everyone.
  • Availability: By managing demand on high-cost models, we can maintain low latency and high availability across the platform.

Questions?

If you have questions about this policy or need a custom arrangement for high-volume usage, reach out on Telegram.