Google확인 중
gemini-3.1-flash-lite-preview
google/gemini-3.1-flash-lite-preview
chat completionsContext: 1.05M도구JSON
가격 불러오는 중…
Official price$0.25 / $1.5 / 100만 토큰
Conversation
Send a message to run gemini-3.1-flash-lite-preview with your account.
Settings
Last response
Usage and charge appear here after a reply.
About this model
gemini-3.1-flash-lite-preview is served through the Hermes AI OpenAI-compatible chat completions endpoint. Send the standard messages array; streaming, function tools, JSON output and reasoning settings pass through unchanged. You pay per token from prepaid credits, at the rates on the Pricing tab.
| Property | 값 |
|---|---|
| Model id | google/gemini-3.1-flash-lite-preview |
| Vendor | |
| Context window | 1,048,576 tokens |
| 기능 | 도구, JSON, Streaming |
| 엔드포인트 | https://api.hermes-ai.net/api/v1/chat/completions |
| Hermes AI 가격 | 가격 불러오는 중… |
Request
POST https://api.hermes-ai.net/api/v1/chat/completions · API key with the chat:create scope, or your signed-in session.
| 필드 | 유형 | 참고 |
|---|---|---|
| model | string | google/gemini-3.1-flash-lite-preview |
| messages | array | OpenAI chat messages: system, user, assistant and tool roles; text and image_url parts. |
| stream | boolean | Server-sent events with a final usage chunk. |
| max_tokens | integer | Output cap; also sizes the credit reservation. |
| tools / tool_choice | array / object | Function tools only. |
| response_format | object | JSON mode and JSON schema where the model supports it. |
| temperature, top_p, stop, seed, reasoning_effort | Passed through unchanged. |
응답
| 필드 | 참고 |
|---|---|
| choices[].message / delta | The OpenAI chat completion shape. |
| usage | prompt_tokens, completion_tokens and prompt_tokens_details.cached_tokens as reported by the model. |
| hermes.charged_micros | Exact charge in USD micro-units (JSON responses). Streams expose it on GET /api/v1/chat/requests. |
| x-hermes-request | Request id header, also used in the billing ledger. |
| error.code | Documented Hermes error code; branch on it, not on the message. |
Token prices
USD per 1M tokens. The tier is chosen by the prompt size of each request.
가격 불러오는 중…
How token billing works
- When a request is accepted, Hermes reserves credits for the estimated prompt plus max_tokens of output (8,192 when you set none). Set max_tokens to keep the reservation small.
- When the response finishes, the reservation is replaced by the exact charge from the usage the model reports: prompt tokens at the input rate, cached prompt tokens at the cache-read rate, completion tokens at the output rate. The remainder is released immediately.
- If the model rejects the request nothing is charged. If a stream is interrupted after tokens were generated, the generated text is estimated and charged.
코드 예시
Any OpenAI SDK works: set the base URL and your Hermes key.
shell
export HERMES_API_KEY=your_api_keycurl https://api.hermes-ai.net/api/v1/chat/completions \
-H "Authorization: Bearer $HERMES_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemini-3.1-flash-lite-preview",
"messages": [{"role": "user", "content": "Say hello in one sentence."}],
"max_tokens": 200
}'