本文へスキップ
Zhipu確認中

glm-5v-turbo

glm/glm-5v-turbo

chat completionsContext: 205KツールVisionReasoning
料金を読み込み中…
Official price$0.736 / $3.236 / 100 万トークン

Token prices

USD per 1M tokens. The tier is chosen by the prompt size of each request.

料金を読み込み中…

How token billing works

  1. When a request is accepted, Hermes reserves credits for the estimated prompt plus max_tokens of output (8,192 when you set none). Set max_tokens to keep the reservation small.
  2. When the response finishes, the reservation is replaced by the exact charge from the usage the model reports: prompt tokens at the input rate, cached prompt tokens at the cache-read rate, completion tokens at the output rate. The remainder is released immediately.
  3. If the model rejects the request nothing is charged. If a stream is interrupted after tokens were generated, the generated text is estimated and charged.