Aller au contenu

Chat completions (LLM)

Call language models through an OpenAI-compatible endpoint with per-token billing from prepaid credits.

OpenAI-compatible endpoint

Language models use a synchronous, OpenAI-compatible endpoint. Any OpenAI SDK works with the base URL https://api.hermes-ai.net/api/v1 and your Hermes key. Streaming, function tools, JSON mode and reasoning settings pass through.

curl https://api.hermes-ai.net/api/v1/chat/completions \
  -H "Authorization: Bearer $HERMES_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-4.5",
    "messages": [{"role": "user", "content": "Say hello in one sentence."}],
    "max_tokens": 200
  }'

Models and token billing

GET /api/v1/chat/models lists callable models with prices in USD per 1M tokens. Billing reserves credits for the estimated prompt plus max_tokens, then charges the exact usage the model reports and releases the rest; set max_tokens to keep reservations small. Errors use the OpenAI error object with a documented code.

Browse language models