OpenAI Chat Completions
Field-for-field compatible with OpenAI /v1/chat/completions, including streaming, tool calls, reasoning parameters and structured output.
Endpoint
POST/v1/chat/completions
| Field | Type | Description | |
|---|---|---|---|
| model | string | required | Model name or alias, see the model catalogue. |
| messages | array | required | Conversation messages, role ∈ system / user / assistant / tool; content supports text and images (image_url). |
| stream | boolean | optional | When true, returns SSE chunks as data: {...}; the last frame carries usage, followed by data: [DONE]. |
| stream_options.include_usage | boolean | optional | Attach usage to the final frame when streaming. |
| max_completion_tokens | integer | optional | Output token cap (max_tokens is accepted too). |
| temperature / top_p | number | optional | Sampling parameters, passed through upstream. |
| tools / tool_choice | array / string | optional | Function-calling definitions and policy, passed through upstream. |
| response_format | object | optional | json_object or json_schema structured output (where the model supports it). |
| reasoning_effort | string | optional | Reasoning effort low / medium / high (reasoning models). |
curl https://www.link42.ai/v1/chat/completions \
-H "Authorization: Bearer $LINK42_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.2",
"stream": true,
"stream_options": {"include_usage": true},
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "List three benefits of using an API gateway"}
]
}'{
"id": "chatcmpl-8f1c…",
"object": "chat.completion",
"created": 1788300000,
"model": "glm-5.2",
"choices": [{"index": 0, "message": {"role": "assistant", "content": "…"}, "finish_reason": "stop"}],
"usage": {"prompt_tokens": 38, "completion_tokens": 120, "total_tokens": 158}
}- The token counts in usage are the billing basis; cached input is charged at the cache_read price.
- If the client drops the connection mid-stream, the part already generated is still settled at actual usage.
Model list
GET/v1/models
Returns the models available to the key's group in the OpenAI shape {object:"list", data:[{id, object:"model", owned_by}]}.
Model reference
The first column is the catalogue id you put in `model`. Deployment-configured aliases such as seed-2.0-pro also work; the model plaza shows the ones your deployment defines.
This table includes migration references and models not yet activated; it does not promise every row is callable. /anthropic/v1/messages and /v1/responses support also depends on the selected model and your key group's current capabilities. Check GET /v1/models and test the endpoint before use. Unspecified fields follow the endpoint reference.
| Model id | Name | Upstream | Billing | Fields that differ | Limits and gotchas |
|---|---|---|---|---|---|
| seed-2-0-pro-260328 | Dola-Seed-2.0-pro | BytePlus Ark | Per token | messages[].content accepts text, image_url and input_audio; the reply message carries reasoning_content; usage carries prompt_tokens_details.cached_tokens / audio_tokens and completion_tokens_details.reasoning_tokens | Above 128,000 input tokens in one request the whole call is priced at double rate; image input counts inside prompt_tokens at the text input price; audio input tokens are counted twice under the legacy_v1 profile |
| seed-2-0-lite-260428 | Dola-Seed-2.0-lite | BytePlus Ark | Per token | Same as 2.0-pro but without image input; identical usage detail | Same long-context threshold and legacy_v1 audio double count as 2.0-pro |
| seed-2-0-mini-260428 | Dola-Seed-2.0-mini | BytePlus Ark | Per token | Same as 2.0-lite; accepts input_audio audio input | Same long-context threshold and legacy_v1 audio double count as 2.0-pro |
| dola-seed-2-1-turbo-260628 | Dola-Seed-2.1-turbo | BytePlus Ark | Per token | Plain chat fields, no extensions; usage carries only prompt / completion plus cached_tokens and reasoning_tokens detail | No long-context surcharge tier |
| deepseek-v4-pro-260425 | DeepSeek-V4-Pro | BytePlus Ark | Per token | Plain chat fields, no extensions | No long-context surcharge tier |
| deepseek-v4-flash-260425 | DeepSeek-V4-Flash | BytePlus Ark | Per token | Plain chat fields, no extensions | No long-context surcharge tier |
| glm-5-2-260617 | GLM-5.2 | BytePlus Ark | Per token | Plain chat fields, no extensions | No long-context surcharge tier |
| deepseek-v4-1-flash | DeepSeek-V4.1-Flash | Volcengine Ark, Beijing | Native CNY list price converted to USD using today's checked SAFE quote, then the effective discount (user if set, otherwise model) | Standard chat / Responses; text, vision and tools; cache hits have their own rate | Upstream deepseek-v4-1-flash-260910; check the live catalogue and routing state for availability, not this document. Weekday peak pricing is 09:00–12:00 and 14:00–18:00 Beijing time. |
| deepseek-v4-pro | DeepSeek-V4-Pro GA | Volcengine Ark, Beijing | Native CNY list price converted to USD using today's checked SAFE quote, then the effective discount (user if set, otherwise model) | Standard chat / Responses; cache hits are priced separately | Upstream deepseek-v4-pro-ga-260813; check the live catalogue and routing state for availability, not this document. |
| deepseek-v4-flash | DeepSeek-V4-Flash GA | Volcengine Ark, Beijing | Native CNY list price converted to USD using today's checked SAFE quote, then the effective discount (user if set, otherwise model) | Standard chat / Responses; cache hits are priced separately | Currently maps upstream to deepseek-v4-flash-ga-260731; see the live marketplace for the USD rate. |
| glm-5-3-flash | GLM-5.3-Flash | Volcengine Ark, Beijing | Native CNY list price converted to USD using today's checked SAFE quote, then the effective discount (user if set, otherwise model) | Standard chat / Responses; cache hits are priced separately | Upstream glm-5-3-flash-260828; check the live catalogue and routing state for availability, not this document. |
| glm-5.2 | Alibaba GLM-5.2 | Alibaba Bailian (OpenAI-compatible endpoint) | Per token | Same single-turn prompt shape on the private protocol as Alibaba DeepSeek-V4-Pro | cached_tokens are billed at the cache-read price; no long-context surcharge tier |
- The long-context surcharge applies to the whole call: once one request exceeds the threshold, that call's input and output are both settled at the raised rate — the surcharge is not limited to the tokens above the threshold.
- Counting audio input tokens twice is historical behaviour kept only in the legacy_v1 profile (migrated ak- keys). Newly created sk- keys are standard_v1 and count them once.
- The unsuffixed deepseek-v4-pro and deepseek-v4-flash models are bound to domestic Volcengine Ark; the dated -260425 rows are separate historical BytePlus entries. Parameter documentation does not imply current activation. Check the live marketplace and your API key's group catalogue.