Check out the newest way to compare different models for a task/agent harness: AutoEvals

Command Reference

Endpoints

Inspect and update an endpoint's model and settings, shared by every deployment serving it.

An endpoint is the public model identifier callers use. One or more deployments serve it. The endpoint owns the model, the serverless pricing, the request-routing settings, the aliases, and the allowlist of extra teams, so these change on the endpoint and apply to every deployment serving it. A deployment's endpointId (from inf deployment get) names its endpoint.

Alias: inf endpoints

inf endpoint get

Get an endpoint's settings and allowed teams.

inf endpoint get --id <endpoint-id>

inf endpoint update

Update an endpoint's settings. Omitted flags leave a setting unchanged; pass null to restore a nullable setting's inherited default.

# Switch the model every deployment on the endpoint serves
inf endpoint update --id <endpoint-id> --model-id <model-id>

# Disable queueing and cap the whole request at 10 minutes (super-admin)
inf endpoint update --id <endpoint-id> \
  --queue-timeout-ms -1 \
  --max-queue-size -1 \
  --total-req-timeout-ms 600000

Changing the model rolls every deployment on the endpoint. It resets each deployment's model-config pin, parallelism, and shard overrides.

FlagAccessDescription
--id <id>AllEndpoint to update (required)
--team-id <id>AllTeam that owns the endpoint (defaults to the active team)
--model-id <id>AllModel every deployment on the endpoint serves
--public-model-identifier <id>Super-adminRename the identifier callers use
--allowed-team-ids <ids>Super-adminReplace the extra teams allowed to call the endpoint
--is-serverless <bool>Super-adminExpose as a public serverless endpoint; going dedicated makes every active deployment's GPU hours billable
--serverless-cost-per-million-in <cents>Super-adminInput price, cents per 1M tokens
--serverless-cost-per-million-cached-in <cents>Super-adminCached-input price, cents per 1M tokens; empty or 0 bills at the input price
--serverless-cost-per-million-out <cents>Super-adminOutput price, cents per 1M tokens
--queue-timeout-ms <n>Super-adminMax wait for a free slot before a 429; -1 disables queueing
--max-queue-size <n>Super-adminWaiting requests before a 429; -1 disables queueing
--total-req-timeout-ms <n>Super-adminWall-clock cap on the whole stream
--time-to-next-token-timeout-ms <n>Super-adminMax gap between stream events
--chunk-buffer-timeout-ms <n>Super-adminEngine wait for an out-of-order chunk (1000–60000)
--missing-chunk-cache-ttl-ms <n>Super-adminWorker chunk retention after a generation ends (5000–300000)
--replay-buffer-max-chunks <n>Super-adminRecent chunks the worker can replay (50–5000)
--enable-sticky-routing <bool>Super-adminRoute repeat conversations to the instance holding their KV cache
--enable-incremental-usage-billing <bool>Super-adminBill cancelled generations for tokens generated before the disconnect

Super-admin commands

inf admin endpoints adds admin conveniences over the same settings:

# Serverless pricing in dollars per 1M tokens, and the allowlist
inf admin endpoints update <endpoint-id> \
  --serverless \
  --serverless-cost-input 1.30 \
  --serverless-cost-output 4.40 \
  --allowed-team <team-id>

# Make the endpoint dedicated again (clears its serverless rates)
inf admin endpoints update <endpoint-id> --private

It also owns the endpoint's aliases (generated add-alias, remove-alias, aliases, and alias-availability) and its pricing rules.

On this page