API
Rate Limits
Rate limits for the Inference.net API
Generation Rate Limits
Rate limits for model inference requests are based on your account tier:
| Tier | Requests per minute (RPM) |
|---|---|
| Free | 30 |
| Paid (Growth, Enterprise) | 1,000 |
| Custom teams | Operator-managed per-team overrides |
Free-tier teams granted free credits by an operator are raised to 200 RPM for the duration of the grant. The pay-as-you-go credit-purchase floor is 250 RPM — a real purchase always buys more headroom than a free grant. Per-team overrides for custom teams are configured by an operator; contact us to request one.
Batch API Rate Limits
- Batch file upload: 1 per minute
- Batch processing rate limits are separate from generation rate limits. See the Batch API docs for details.
Deployed models share the team's serverless inference RPM bucket rather than having their own per-instance limit.
Increasing Your Limits
If you need higher rate limits, contact us or use the support chat to request a custom tier.