gemini-3.5-flash-lite-saver
google/gemini-3.5-flash-lite-saver
Conversation
Send a message to run gemini-3.5-flash-lite-saver with your account.
Settings
Last response
Usage and charge appear here after a reply.
About this model
gemini-3.5-flash-lite-saver is served through the Hermes AI OpenAI-compatible chat completions endpoint. Send the standard messages array; streaming, function tools, JSON output and reasoning settings pass through unchanged. You pay per token from prepaid credits, at the rates on the Pricing tab. The saver channel is the same model on a lower-priority upstream route, sold well below the manufacturer list price.
| Property | 값 |
|---|---|
| Model id | google/gemini-3.5-flash-lite-saver |
| Vendor | |
| Context window | 1,048,576 tokens |
| 기능 | 도구, Vision, 오디오, Streaming |
| 엔드포인트 | https://api.hermes-ai.net/api/v1/chat/completions |
| Hermes AI 가격 | 가격 불러오는 중… |
Request
POST https://api.hermes-ai.net/api/v1/chat/completions · API key with the chat:create scope, or your signed-in session.
| 필드 | 유형 | 참고 |
|---|---|---|
| model | string | google/gemini-3.5-flash-lite-saver |
| messages | array | OpenAI chat messages: system, user, assistant and tool roles; text and image_url parts. |
| stream | boolean | Server-sent events with a final usage chunk. |
| max_tokens | integer | Output cap; also sizes the credit reservation. |
| tools / tool_choice | array / object | Function tools only. |
| response_format | object | JSON mode and JSON schema where the model supports it. |
| temperature, top_p, stop, seed, reasoning_effort | Passed through unchanged. |
응답
| 필드 | 참고 |
|---|---|
| choices[].message / delta | The OpenAI chat completion shape. |
| usage | prompt_tokens, completion_tokens and prompt_tokens_details.cached_tokens as reported by the model. |
| hermes.charged_micros | Exact charge in USD micro-units (JSON responses). Streams expose it on GET /api/v1/chat/requests. |
| x-hermes-request | Request id header, also used in the billing ledger. |
| error.code | Documented Hermes error code; branch on it, not on the message. |
Token prices
USD per 1M tokens. The tier is chosen by the prompt size of each request.
가격 불러오는 중…
How token billing works
- When a request is accepted, Hermes reserves credits for the estimated prompt plus max_tokens of output (8,192 when you set none). Set max_tokens to keep the reservation small.
- When the response finishes, the reservation is replaced by the exact charge from the usage the model reports: prompt tokens at the input rate, cached prompt tokens at the cache-read rate, completion tokens at the output rate. The remainder is released immediately.
- If the model rejects the request nothing is charged. If a stream is interrupted after tokens were generated, the generated text is estimated and charged.
코드 예시
Any OpenAI SDK works: set the base URL and your Hermes key.
export HERMES_API_KEY=your_api_keycurl https://api.hermes-ai.net/api/v1/chat/completions \
-H "Authorization: Bearer $HERMES_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemini-3.5-flash-lite-saver",
"messages": [{"role": "user", "content": "Say hello in one sentence."}],
"max_tokens": 200
}'