- RPM: Requests per minute
- RPD: Requests per day
- A request is defined by a call to our API
- You can hit either limit type (RPM or RPD) depending on which one you reach first
- You will be notified in every request response what the status of your rate limits are (see rate limit response headers for more information)
- If you hit a rate limit, you will be sent an error message in your response (see API error codes)
- Check the Infercom Status Page for real-time platform and model availability
Infercom Inference Service rate limit tiers
There are a few different rate limit tier offerings we provide:- Free Tier: Applied when there is no payment method linked with your account
- Developer Tier: Applied when a payment method is linked with your account
- Enterprise Tier: Higher limits for production workloads
Please see the Billing page to link a payment method to your account.
Model rate limits
- Free Tier
- Developer Tier
- Enterprise Tier
EU-hosted models (sovereign)
Global Model Catalog (non-sovereign)
Need custom limits? Enterprise customers can request custom rate limits based on workload requirements. Contact sales to discuss your needs.
Rate limit response headers
These headers are found in every request response and give information about the current status of rate limit usage. RPM (Requests per minute):x-ratelimit-limit-requests- The maximum number of requests allowed per minute.
x-ratelimit-remaining-requests- The number of requests remaining in the current minute before hitting the rate limit.
x-ratelimit-reset-requests- Time in epoch time until the per-minute request quota resets.
x-ratelimit-limit-requests-day- The maximum number of requests allowed per day.
x-ratelimit-remaining-requests-day- The number of requests remaining in the current day before hitting the rate limit.
x-ratelimit-reset-requests-day- Time in epoch time until the per-day request quota resets.