Chat completions (LLM)
Call language models through an OpenAI-compatible endpoint with per-token billing from prepaid credits.
OpenAI-compatible endpoint
Language models use a synchronous, OpenAI-compatible endpoint. Any OpenAI SDK works with the base URL https://api.hermes-ai.net/api/v1 and your Hermes key. Streaming, function tools, JSON mode and reasoning settings pass through.
curl https://api.hermes-ai.net/api/v1/chat/completions \
-H "Authorization: Bearer $HERMES_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4.5",
"messages": [{"role": "user", "content": "Say hello in one sentence."}],
"max_tokens": 200
}'Models and token billing
GET /api/v1/chat/models lists callable models with prices in USD per 1M tokens. Billing reserves credits for the estimated prompt plus max_tokens, then charges the exact usage the model reports and releases the rest; set max_tokens to keep reservations small. Errors use the OpenAI error object with a documented code.