Skip to main content
Perplexity

Search documentation

Type to search this documentation.

Create Chat Completion

POST

/

router

/

v1

/

chat

/

completions

Create Chat Completion

curl --request POST \

  --url https://api.perplexity.ai/router/v1/chat/completions \

  --header 'Authorization: Bearer <token>' \

  --header 'Content-Type: application/json' \

  --data '

{

  "messages": [

    {

      "content": "<string>",

      "role": "developer",

      "name": "<string>"

    }

  ],

  "model": "perplexity/kimi-k3",

  "tools": [

    {

      "type": "function",

      "function": {

        "name": "<string>",

        "parameters": {},

        "strict": false

      }

    }

  ]

}

'
title="200"
{

  "id": "<string>",

  "choices": [

    {

      "finish_reason": "stop",

      "index": 123,

      "message": {

        "content": "<string>",

        "refusal": "<string>",

        "role": "assistant",

        "reasoning_content": "<string>",

        "tool_calls": [

          {

            "id": "<string>",

            "type": "function",

            "function": {

              "name": "<string>",

              "arguments": "<string>"

            }

          }

        ],

        "annotations": [

          {

            "type": "url_citation",

            "url_citation": {

              "end_index": 123,

              "start_index": 123,

              "url": "<string>",

              "title": "<string>"

            }

          }

        ],

        "function_call": {

          "arguments": "<string>",

          "name": "<string>"

        },

        "audio": {

          "id": "<string>",

          "expires_at": 123,

          "data": "<string>",

          "transcript": "<string>"

        }

      },

      "logprobs": {

        "content": [

          {

            "token": "<string>",

            "logprob": 123,

            "bytes": [

              123

            ],

            "top_logprobs": [

              {

                "token": "<string>",

                "logprob": 123,

                "bytes": [

                  123

                ]

              }

            ]

          }

        ],

        "refusal": [

          {

            "token": "<string>",

            "logprob": 123,

            "bytes": [

              123

            ],

            "top_logprobs": [

              {

                "token": "<string>",

                "logprob": 123,

                "bytes": [

                  123

                ]

              }

            ]

          }

        ]

      }

    }

  ],

  "created": 123,

  "model": "<string>",

  "object": "chat.completion",

  "service_tier": "auto",

  "system_fingerprint": "<string>",

  "usage": {

    "completion_tokens": 0,

    "prompt_tokens": 0,

    "total_tokens": 0,

    "completion_tokens_details": {

      "accepted_prediction_tokens": 0,

      "audio_tokens": 0,

      "reasoning_tokens": 0,

      "rejected_prediction_tokens": 0

    },

    "prompt_tokens_details": {

      "audio_tokens": 0,

      "cached_tokens": 0,

      "cache_write_tokens": 0

    }

  },

  "moderation": {

    "input": {

      "type": "moderation_results",

      "model": "<string>",

      "results": [

        {

          "type": "moderation_result",

          "model": "<string>",

          "flagged": true,

          "categories": {},

          "category_scores": {},

          "category_applied_input_types": {}

        }

      ]

    },

    "output": {

      "type": "moderation_results",

      "model": "<string>",

      "results": [

        {

          "type": "moderation_result",

          "model": "<string>",

          "flagged": true,

          "categories": {},

          "category_scores": {},

          "category_applied_input_types": {}

        }

      ]

    }

  }

}
title="400"
{
  "error": {
    "code": "<string>",
    "message": "<string>",
    "param": "<string>",
    "type": "<string>"
  }
}
title="429"
{
  "error": {
    "code": "<string>",
    "message": "<string>",
    "param": "<string>",
    "type": "<string>"
  }
}
Parameter support

Honored: model, messages, stream, stream_options.include_usage, max_completion_tokens, max_tokens, temperature, top_p, stop, reasoning_effort, service_tier (auto, default, flex, priority), response_format (text, json_schema, json_object), tools, tool_choice (including allowed_tools), parallel_tool_calls, prompt_cache_key, prompt_cache_options, prompt_cache_retention, and message-part prompt_cache_breakpoint fields.Accepted but not forwarded to the model: user, safety_identifier, metadata.Accepted only at their default values: n (1), logprobs (false), store (false), presence_penalty (0), frequency_penalty (0).Rejected with a 400: seed, logit_bias, top_logprobs, functions (use tools), function_call (use tool_choice), modalities, audio, prediction, web_search_options, moderation, verbosity, stream_options.include_obfuscation set to true, plus any unrecognized top-level field.Notes: When reasoning_effort is omitted, the provider’s default for that model applies; defaults vary by provider and model. response_format.json_schema always runs in strict mode. Assistant messages and streamed deltas can include reasoning_content; replay it in an assistant message to preserve reasoning context across requests. Function tools require a description.

Errors

Errors use the OpenAI envelope:

JSON
{

  "error": {

    "code": null,

    "message": "Invalid model 'example/does-not-exist'. Permitted models can be found in the documentation at https://docs.perplexity.ai/docs/getting-started/models.",

    "param": null,

    "type": "invalid_request_error"

  }

}

type is invalid_request_error for 4xx errors, rate_limit_error for 429 (rate limited or model overloaded — retry after the Retry-After interval), and server_error for 5xx. Requests that fail before producing output are not billed.

Authorization

string

header

required

Your Perplexity API key.

application/json

messages

(Developer message · object | System message · object | User message · object | Assistant message · object | Tool message · object | Function message · object)[]

required

The conversation so far, as an array of message objects with role and content. The leading system or developer message is applied as the system prompt.

Minimum array length: 1

Show child attributes

model

string

required

The model that will process the request, as a creator/model-name id from the catalog — for example perplexity/kimi-k3. Any cataloged model works with this endpoint.

Example:

"perplexity/kimi-k3"

metadata

object | null

Accepted for SDK compatibility; not forwarded to the model.

Show child attributes

temperature

number | null

default:1

Sampling temperature, 0 to 2. Higher values produce more varied output. Adjust this or top_p, not both.

Required range: 0 <= x <= 2

Example:

1

top_p

number | null

default:1

Nucleus-sampling threshold, 0 to 1. Adjust this or temperature, not both.

Required range: 0 <= x <= 1

Example:

1

user

string

deprecated

Accepted for SDK compatibility; not forwarded to the model.

Example:

"user-1234"

safety_identifier

string | null

Accepted for SDK compatibility; not forwarded to the model.

Maximum string length: 64

Example:

"safety-identifier-1234"

prompt_cache_key

string | null

Stable key grouping related requests to improve prompt-cache hit rates.

Example:

"prompt-cache-key-1234"

service_tier

enum<string> | null

default:auto

Processing tier for the request. flex halves token rates on supported models; auto and default are equivalent.

Available options:

auto,

default,

flex,

scale,

priority

prompt_cache_retention

enum<string> | null

deprecated

Deprecated cache-retention setting. Use prompt_cache_options.ttl for new integrations.

Available options:

in_memory,

24h

prompt_cache_options

Prompt cache options · object

Prompt-cache settings. Set ttl to 5m, 1h, or 24h to control cache retention.

Show child attributes

reasoning_effort

enum<string> | null

How much reasoning the model performs before answering, on models that support it. When omitted, the provider's default for that model applies; defaults vary by provider and model.

Available options:

none,

minimal,

low,

medium,

high,

xhigh,

max

max_completion_tokens

integer | null

Upper bound on generated tokens for this request, including reasoning tokens.

response_format

Text · object

Output format. Use type json_schema for strict structured output or json_object for JSON mode.

Show child attributes

stream

boolean | null

default:false

When true, tokens are sent as server-sent events as they are generated, terminated by data: [DONE].

stop

default:<|endoftext|>

Up to 4 sequences at which generation stops.

Example:

"\n"

max_tokens

integer | null

deprecated

Legacy alias of max_completion_tokens.

stream_options

object | null

Streaming options. Only allowed when stream is true.

Show child attributes

tools

(Function tool · object | Custom tool · object)[]

Function tools the model may call. Every tool requires a description.

Show child attributes

tool_choice

Controls whether the model may call tools, requires a specific tool, or limits selection with allowed_tools.

Available options:

none,

auto,

required

parallel_tool_calls

boolean

default:true

Whether the model may request more than one tool call in a single turn.

Successful response. JSON for non-streaming requests; a text/event-stream of chat completion chunks terminated by data: [DONE] when stream is true.

id

string

required

Unique identifier for the completion.

choices

object[]

required

The completion. Contains exactly one choice.

Show child attributes

created

integer<unixtime>

required

Unix timestamp of when the completion was created.

model

string

required

The model id you requested. Billing always uses this model's published rates.

object

enum<string>

required

Always "chat.completion".

Available options:

chat.completion

service_tier

enum<string> | null

default:auto

The processing tier that served the request.

Available options:

auto,

default,

flex,

scale,

priority

system_fingerprint

string

deprecated

usage

object

Token accounting for the request.

Show child attributes

moderation

object

Show child attributes

Was this page helpful?

⌘I

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu