Create Message
POST
/
router
/
v1
/
messages
Create Message
curl --request POST \
--url https://api.perplexity.ai/router/v1/messages \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "<string>",
"max_tokens": 1073741824,
"messages": [
{
"content": "<string>"
}
],
"thinking": {
"type": "enabled",
"budget_tokens": 123
}
}
'{
"id": "<string>",
"type": "message",
"role": "assistant",
"model": "<string>",
"content": [
{
"type": "text",
"text": "<string>",
"cache_control": {
"type": "ephemeral",
"ttl": "5m"
}
}
],
"stop_reason": "end_turn",
"stop_sequence": "<string>",
"usage": {
"input_tokens": 123,
"output_tokens": 123,
"cache_creation_input_tokens": 123,
"cache_read_input_tokens": 123,
"cache_creation": {
"ephemeral_5m_input_tokens": 123,
"ephemeral_1h_input_tokens": 123
},
"service_tier": "<string>"
}
}{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "<string>"
}
}{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "<string>"
}
}Parameter support
Honored: model, max_tokens (required), messages, system, stream, temperature, top_p, top_k, stop_sequences, thinking (only {"type": "disabled"}), tools, tool_choice. Message content and tool_result.content accept text, image, document, and search_result blocks. Document sources can be base64, plain text, or URLs and may include title and context.Accepted but not forwarded to the model: metadata.Rejected with a 400: service_tier, thinking with {"type": "enabled"}, cache_control on content blocks or tools, plus any unrecognized top-level field. A tool’s description is optional.
Errors
Errors use the Anthropic envelope:
{
"type": "error",
"error": {
"type": "overloaded_error",
"message": "upstream model is overloaded, please try again later"
}
}An overloaded model returns HTTP 429 with type overloaded_error and a Retry-After header. Requests that fail before producing output are not billed.
Authorizations
Section titled “Authorizations”Authorization
string
header
required
Your Perplexity API key.
application/json
Request body for POST /router/v1/messages.
model
string
required
Public model slug, e.g. perplexity/kimi-k3.
max_tokens
integer
required
Maximum number of tokens to generate. Required by the Messages API.
Required range: 1 <= x <= 2147483647
messages
object[]
required
Show child attributes
system
System prompt as plain text or an array of text blocks.
temperature
number<float>
Required range: 0 <= x <= 1
top_p
number<float>
Required range: 0 <= x <= 1
top_k
integer
Only sample from the top K options for each subsequent token.
stop_sequences
string[]
stream
boolean
When true, respond with server-sent events.
tools
object[]
Show child attributes
tool_choice
object
Show child attributes
metadata
object
Show child attributes
thinking
object
Legacy budget-based thinking; budget_tokens maps to the nearest llm-api reasoning-effort tier.
Show child attributes
service_tier
enum<string>
Accepted and ignored — llm-api's service_tier is OpenAI-only.
Available options:
auto,
standard_only
cache_control
object
Ephemeral cache breakpoint. ttl "1h" is rejected until 1h writes are billed distinctly (they price 2x the 5m rate).
Show child attributes
output_config
object
Show child attributes
context_management
object
Accepted and ignored — context edits are not applied. Kept raw so evolving strategies decode.
container
any
Rejected by validation — no code-execution containers.
inference_geo
string
Rejected by validation — geo pinning cannot be honored, silently ignoring it would break residency expectations.
mcp_servers
any
Rejected by validation — no server-side MCP execution.
speed
enum<string>
"standard" is inert; "fast" is rejected (no fast-mode routing).
Available options:
standard,
fast
fallbacks
any
Rejected by validation — fallback models would bill as the requested slug.
fallback_credit_token
string
Rejected by validation together with fallbacks.
Response
Section titled “Response”Successful response. JSON for non-streaming requests; a text/event-stream of typed events (message_start through message_stop) when stream is true.
Non-streaming response body and the message_start payload.
id
string
required
type
enum<string>
required
Available options:
message
role
enum<string>
required
Available options:
assistant
model
string
required
content
object[]
required
Show child attributes
stop_reason
enum<string> | null
required
Available options:
end_turn,
max_tokens,
stop_sequence,
tool_use,
pause_turn,
refusal,
model_context_window_exceeded
stop_sequence
string | null
required
usage
object
required
Show child attributes
Was this page helpful?
⌘I