Command Reference
Endpoints
Inspect and update an endpoint's model and settings, shared by every deployment serving it.
An endpoint is the public model identifier callers use. One or more deployments serve it. The endpoint owns the model, the serverless pricing, the request-routing settings, the aliases, and the allowlist of extra teams, so these change on the endpoint and apply to every deployment serving it. A deployment's endpointId (from inf deployment get) names its endpoint.
Alias: inf endpoints
inf endpoint get
Get an endpoint's settings and allowed teams.
inf endpoint get --id <endpoint-id>inf endpoint update
Update an endpoint's settings. Omitted flags leave a setting unchanged; pass null to restore a nullable setting's inherited default.
# Switch the model every deployment on the endpoint serves
inf endpoint update --id <endpoint-id> --model-id <model-id>
# Disable queueing and cap the whole request at 10 minutes (super-admin)
inf endpoint update --id <endpoint-id> \
--queue-timeout-ms -1 \
--max-queue-size -1 \
--total-req-timeout-ms 600000Changing the model rolls every deployment on the endpoint. It resets each deployment's model-config pin, parallelism, and shard overrides.
| Flag | Access | Description |
|---|---|---|
--id <id> | All | Endpoint to update (required) |
--team-id <id> | All | Team that owns the endpoint (defaults to the active team) |
--model-id <id> | All | Model every deployment on the endpoint serves |
--public-model-identifier <id> | Super-admin | Rename the identifier callers use |
--allowed-team-ids <ids> | Super-admin | Replace the extra teams allowed to call the endpoint |
--is-serverless <bool> | Super-admin | Expose as a public serverless endpoint; going dedicated makes every active deployment's GPU hours billable |
--serverless-cost-per-million-in <cents> | Super-admin | Input price, cents per 1M tokens |
--serverless-cost-per-million-cached-in <cents> | Super-admin | Cached-input price, cents per 1M tokens; empty or 0 bills at the input price |
--serverless-cost-per-million-out <cents> | Super-admin | Output price, cents per 1M tokens |
--queue-timeout-ms <n> | Super-admin | Max wait for a free slot before a 429; -1 disables queueing |
--max-queue-size <n> | Super-admin | Waiting requests before a 429; -1 disables queueing |
--total-req-timeout-ms <n> | Super-admin | Wall-clock cap on the whole stream |
--time-to-next-token-timeout-ms <n> | Super-admin | Max gap between stream events |
--chunk-buffer-timeout-ms <n> | Super-admin | Engine wait for an out-of-order chunk (1000–60000) |
--missing-chunk-cache-ttl-ms <n> | Super-admin | Worker chunk retention after a generation ends (5000–300000) |
--replay-buffer-max-chunks <n> | Super-admin | Recent chunks the worker can replay (50–5000) |
--enable-sticky-routing <bool> | Super-admin | Route repeat conversations to the instance holding their KV cache |
--enable-incremental-usage-billing <bool> | Super-admin | Bill cancelled generations for tokens generated before the disconnect |
Super-admin commands
inf admin endpoints adds admin conveniences over the same settings:
# Serverless pricing in dollars per 1M tokens, and the allowlist
inf admin endpoints update <endpoint-id> \
--serverless \
--serverless-cost-input 1.30 \
--serverless-cost-output 4.40 \
--allowed-team <team-id>
# Make the endpoint dedicated again (clears its serverless rates)
inf admin endpoints update <endpoint-id> --privateIt also owns the endpoint's aliases (generated add-alias, remove-alias, aliases, and alias-availability) and its pricing rules.
Related commands
inf deployment create --endpoint-id <id>adds another deployment to an endpoint.inf deployment updatechanges one deployment's own settings: name, GPUs, instances, engine overrides.