Skip to main content
Datasets can be created from captured traffic or uploaded as JSONL files. This page covers the supported formats.

Supported formats

The system auto-detects the format from the first valid line. All rows in a file must use the same format.
You cannot mix source-backed and Hugging Face rows in the same file. Mixed-format files fail validation.

Source-backed format

Each line has a top-level request and optional response object containing raw provider bodies. Validation notes:
  • The request must include a usable model value.
  • response may be omitted or set to null if you only have request-side data.

Hugging Face format

Each line has a top-level messages array with role/content objects. Valid roles: system, user, assistant, tool. Additional supported fields:
  • content may be a string or an array of content parts for multimodal rows.
  • Assistant messages may include tool_calls.
  • Tool messages must include tool_call_id.
  • Top-level tools are preserved on import.
When importing Hugging Face-format rows, the system:
  • Treats the last assistant turn as the imported response and earlier turns as request context
  • Synthesizes request/response payloads so evals and detail views work
  • Sets request_model to unknown-imported-model
  • Sets token usage and costs to zero
  • Stores the original row id in metadata as importOriginalRowId

Validation behavior

  • Files must be valid JSONL.
  • Invalid rows are reported with line numbers in the upload status details.
  • Uploads can complete with partial failures if at least one row imports successfully.
  • If every row fails validation, the upload status is failed.

Upload limits

Download formats

Datasets can be downloaded in two formats: In the UI, click Download and choose the format. In the CLI:
Hugging Face exports skip rows with empty message arrays. Source-backed exports include all rows with a valid request payload. Row counts may differ between formats.