Create Chat Completion
POST
/
router
/
v1
/
chat
/
completions
Create Chat Completion
curl --request POST \
--url https://api.perplexity.ai/router/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"messages": [
{
"content": "<string>",
"role": "developer",
"name": "<string>"
}
],
"model": "perplexity/kimi-k3",
"tools": [
{
"type": "function",
"function": {
"name": "<string>",
"parameters": {},
"strict": false
}
}
]
}
'{
"id": "<string>",
"choices": [
{
"finish_reason": "stop",
"index": 123,
"message": {
"content": "<string>",
"refusal": "<string>",
"role": "assistant",
"reasoning_content": "<string>",
"tool_calls": [
{
"id": "<string>",
"type": "function",
"function": {
"name": "<string>",
"arguments": "<string>"
}
}
],
"annotations": [
{
"type": "url_citation",
"url_citation": {
"end_index": 123,
"start_index": 123,
"url": "<string>",
"title": "<string>"
}
}
],
"function_call": {
"arguments": "<string>",
"name": "<string>"
},
"audio": {
"id": "<string>",
"expires_at": 123,
"data": "<string>",
"transcript": "<string>"
}
},
"logprobs": {
"content": [
{
"token": "<string>",
"logprob": 123,
"bytes": [
123
],
"top_logprobs": [
{
"token": "<string>",
"logprob": 123,
"bytes": [
123
]
}
]
}
],
"refusal": [
{
"token": "<string>",
"logprob": 123,
"bytes": [
123
],
"top_logprobs": [
{
"token": "<string>",
"logprob": 123,
"bytes": [
123
]
}
]
}
]
}
}
],
"created": 123,
"model": "<string>",
"object": "chat.completion",
"service_tier": "auto",
"system_fingerprint": "<string>",
"usage": {
"completion_tokens": 0,
"prompt_tokens": 0,
"total_tokens": 0,
"completion_tokens_details": {
"accepted_prediction_tokens": 0,
"audio_tokens": 0,
"reasoning_tokens": 0,
"rejected_prediction_tokens": 0
},
"prompt_tokens_details": {
"audio_tokens": 0,
"cached_tokens": 0,
"cache_write_tokens": 0
}
},
"moderation": {
"input": {
"type": "moderation_results",
"model": "<string>",
"results": [
{
"type": "moderation_result",
"model": "<string>",
"flagged": true,
"categories": {},
"category_scores": {},
"category_applied_input_types": {}
}
]
},
"output": {
"type": "moderation_results",
"model": "<string>",
"results": [
{
"type": "moderation_result",
"model": "<string>",
"flagged": true,
"categories": {},
"category_scores": {},
"category_applied_input_types": {}
}
]
}
}
}{
"error": {
"code": "<string>",
"message": "<string>",
"param": "<string>",
"type": "<string>"
}
}{
"error": {
"code": "<string>",
"message": "<string>",
"param": "<string>",
"type": "<string>"
}
}Parameter support
Honored: model, messages, stream, stream_options.include_usage, max_completion_tokens, max_tokens, temperature, top_p, stop, reasoning_effort, service_tier (auto, default, flex, priority), response_format (text, json_schema, json_object), tools, tool_choice (including allowed_tools), parallel_tool_calls, prompt_cache_key, prompt_cache_options, prompt_cache_retention, and message-part prompt_cache_breakpoint fields.Accepted but not forwarded to the model: user, safety_identifier, metadata.Accepted only at their default values: n (1), logprobs (false), store (false), presence_penalty (0), frequency_penalty (0).Rejected with a 400: seed, logit_bias, top_logprobs, functions (use tools), function_call (use tool_choice), modalities, audio, prediction, web_search_options, moderation, verbosity, stream_options.include_obfuscation set to true, plus any unrecognized top-level field.Notes: When reasoning_effort is omitted, the provider’s default for that model applies; defaults vary by provider and model. response_format.json_schema always runs in strict mode. Assistant messages and streamed deltas can include reasoning_content; replay it in an assistant message to preserve reasoning context across requests. Function tools require a description.
Errors
Errors use the OpenAI envelope:
{
"error": {
"code": null,
"message": "Invalid model 'example/does-not-exist'. Permitted models can be found in the documentation at https://docs.perplexity.ai/docs/getting-started/models.",
"param": null,
"type": "invalid_request_error"
}
}type is invalid_request_error for 4xx errors, rate_limit_error for 429 (rate limited or model overloaded — retry after the Retry-After interval), and server_error for 5xx. Requests that fail before producing output are not billed.
Authorizations
Section titled “Authorizations”Authorization
string
header
required
Your Perplexity API key.
application/json
messages
(Developer message · object | System message · object | User message · object | Assistant message · object | Tool message · object | Function message · object)[]
required
The conversation so far, as an array of message objects with role and content. The leading system or developer message is applied as the system prompt.
Minimum array length: 1
Show child attributes
model
string
required
The model that will process the request, as a creator/model-name id from the catalog — for example perplexity/kimi-k3. Any cataloged model works with this endpoint.
Example:
"perplexity/kimi-k3"
metadata
object | null
Accepted for SDK compatibility; not forwarded to the model.
Show child attributes
temperature
number | null
default:1
Sampling temperature, 0 to 2. Higher values produce more varied output. Adjust this or top_p, not both.
Required range: 0 <= x <= 2
Example:
1
top_p
number | null
default:1
Nucleus-sampling threshold, 0 to 1. Adjust this or temperature, not both.
Required range: 0 <= x <= 1
Example:
1
user
string
deprecated
Accepted for SDK compatibility; not forwarded to the model.
Example:
"user-1234"
safety_identifier
string | null
Accepted for SDK compatibility; not forwarded to the model.
Maximum string length: 64
Example:
"safety-identifier-1234"
prompt_cache_key
string | null
Stable key grouping related requests to improve prompt-cache hit rates.
Example:
"prompt-cache-key-1234"
service_tier
enum<string> | null
default:auto
Processing tier for the request. flex halves token rates on supported models; auto and default are equivalent.
Available options:
auto,
default,
flex,
scale,
priority
prompt_cache_retention
enum<string> | null
deprecated
Deprecated cache-retention setting. Use prompt_cache_options.ttl for new integrations.
Available options:
in_memory,
24h
prompt_cache_options
Prompt cache options · object
Prompt-cache settings. Set ttl to 5m, 1h, or 24h to control cache retention.
Show child attributes
reasoning_effort
enum<string> | null
How much reasoning the model performs before answering, on models that support it. When omitted, the provider's default for that model applies; defaults vary by provider and model.
Available options:
none,
minimal,
low,
medium,
high,
xhigh,
max
max_completion_tokens
integer | null
Upper bound on generated tokens for this request, including reasoning tokens.
response_format
Text · object
Output format. Use type json_schema for strict structured output or json_object for JSON mode.
Show child attributes
stream
boolean | null
default:false
When true, tokens are sent as server-sent events as they are generated, terminated by data: [DONE].
stop
default:<|endoftext|>
Up to 4 sequences at which generation stops.
Example:
"\n"
max_tokens
integer | null
deprecated
Legacy alias of max_completion_tokens.
stream_options
object | null
Streaming options. Only allowed when stream is true.
Show child attributes
tools
(Function tool · object | Custom tool · object)[]
Function tools the model may call. Every tool requires a description.
Show child attributes
tool_choice
Controls whether the model may call tools, requires a specific tool, or limits selection with allowed_tools.
Available options:
none,
auto,
required
parallel_tool_calls
boolean
default:true
Whether the model may request more than one tool call in a single turn.
Response
Section titled “Response”Successful response. JSON for non-streaming requests; a text/event-stream of chat completion chunks terminated by data: [DONE] when stream is true.
id
string
required
Unique identifier for the completion.
choices
object[]
required
The completion. Contains exactly one choice.
Show child attributes
created
integer<unixtime>
required
Unix timestamp of when the completion was created.
model
string
required
The model id you requested. Billing always uses this model's published rates.
object
enum<string>
required
Always "chat.completion".
Available options:
chat.completion
service_tier
enum<string> | null
default:auto
The processing tier that served the request.
Available options:
auto,
default,
flex,
scale,
priority
system_fingerprint
string
deprecated
usage
object
Token accounting for the request.
Show child attributes
moderation
object
Show child attributes
Was this page helpful?
⌘I