Check out the newest way to compare different models for a task/agent harness: AutoEvals

API

Rate Limits

Rate limits for the Inference.net API

Generation Rate Limits

Rate limits for model inference requests are based on your account tier:

TierRequests per minute (RPM)
Free30
Paid (Growth, Enterprise)1,000
Custom teamsOperator-managed per-team overrides

Free-tier teams granted free credits by an operator are raised to 200 RPM for the duration of the grant. The pay-as-you-go credit-purchase floor is 250 RPM — a real purchase always buys more headroom than a free grant. Per-team overrides for custom teams are configured by an operator; contact us to request one.

Batch API Rate Limits

  • Batch file upload: 1 per minute
  • Batch processing rate limits are separate from generation rate limits. See the Batch API docs for details.

Deployed models share the team's serverless inference RPM bucket rather than having their own per-instance limit.

Increasing Your Limits

If you need higher rate limits, contact us or use the support chat to request a custom tier.

On this page