# Hermes AI API > One API for image, video and music AI models. Submit an asynchronous task with one request > shape for every model, poll for the result, and pay a per-task quote from prepaid credits. Base URL: https://api.hermes-ai.net. Authenticate with `Authorization: Bearer ` (keys start with `hermes_`), or connect an MCP client to https://api.hermes-ai.net/mcp and sign in with OAuth. ## Docs - [API guide](https://api.hermes-ai.net/docs): quickstart, authentication, tasks, results, errors and pricing - [MCP and Agent Skill](https://api.hermes-ai.net/agents): connect Claude Code, Cursor, Codex and other agents - [Agent Skill](https://api.hermes-ai.net/skills/hermes-ai-api/SKILL.md): SKILL.md with the full task workflow for agents - [OpenAPI 3.1](https://api.hermes-ai.net/api/openapi.json): machine-readable API, including each model's input schema - [Pricing](https://api.hermes-ai.net/pricing): current per-task prices ## Models ### Image - [nano-banana-pro/text-to-image](https://api.hermes-ai.net/models/nano-banana-pro-text-to-image-economy): `nano-banana-pro-text-to-image-economy` — Google's Nano Banana Pro (Gemini 3.0 Pro Image) is an industry-leading cutting-edge text-to-image model, which can generate high-definition - [nano-banana-pro/edit](https://api.hermes-ai.net/models/nano-banana-pro-edit-economy): `nano-banana-pro-edit-economy` — Google Nano Banana Pro (Gemini 3.0 Pro Image) Edit supports professional image editing with high-quality 4K-capable ultra-clear output, driv - [nano-banana-2-lite/image-to-image](https://api.hermes-ai.net/models/nano-banana-2-lite-image-to-image-economy): `nano-banana-2-lite-image-to-image-economy` — It delivers ultra-low-latency and cost-effective image generation & editing capabilities. Optimized to achieve latency under 2 seconds and d - [nano-banana-2-lite/text-to-image](https://api.hermes-ai.net/models/nano-banana-2-lite-text-to-image-economy): `nano-banana-2-lite-text-to-image-economy` — Lightweight text-to-image API built for high concurrency & instant response, with low-latency, budget-friendly image generation and editing. - [nano-banana-2/image-to-image](https://api.hermes-ai.net/models/nano-banana-2-image-to-image-economy): `nano-banana-2-image-to-image-economy` — An image-to-image and editing endpoint powered by a highly efficient visual engine. It enables rapid style transfer, inpainting, or backgrou - [nano-banana-2/text-to-image](https://api.hermes-ai.net/models/nano-banana-2-text-to-image-economy): `nano-banana-2-text-to-image-economy` — A lightweight text-to-image endpoint engineered for high-concurrency and rapid response. As the core of Nano Banana 2, it balances visual fi - [gpt-image-2.5/sunburst/image-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-sunburst-image-to-image-economy): `gpt-image-2-5-sunburst-image-to-image-economy` — GPT Image 2.5 Sunburst Edit provides precision-first revisions while preserving subjects, composition, and visual identity. - [gpt-image-2.5/sunburst/text-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-sunburst-text-to-image-economy): `gpt-image-2-5-sunburst-text-to-image-economy` — GPT Image 2.5 Sunburst focuses on intricate detail, reliable typography, and controlled composition. - [gpt-image-2.5/flare/image-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-flare-image-to-image-economy): `gpt-image-2-5-flare-image-to-image-economy` — GPT Image 2.5 Flare Edit supports fast everyday revisions while preserving subjects, composition, and background. - [gpt-image-2.5/flare/text-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-flare-text-to-image-economy): `gpt-image-2-5-flare-text-to-image-economy` — GPT Image 2.5 Flare is a fast, versatile image generator for social content, commerce assets, ideation, and high-volume creative work. - [xai/grok-imagine-2.0/text-to-image](https://api.hermes-ai.net/models/xai-grok-imagine-2-0-text-to-image-economy): `xai-grok-imagine-2-0-text-to-image-economy` — Grok Imagine 2.0 low-price channel edition provides async text-to-image generation with seven common aspect ratios and 1–12 images per reque - [gpt-image-2.0/text-to-image](https://api.hermes-ai.net/models/gpt-image-2-0-text-to-image-economy): `gpt-image-2-0-text-to-image-economy` — The GPT-Image-2 Text-to-Image is a state-of-the-art generation foundation designed for high-standard commercial scenarios. It features revol - [gpt-image-2.0/edit](https://api.hermes-ai.net/models/gpt-image-2-0-edit-economy): `gpt-image-2-0-edit-economy` — The GPT-Image-2 Image-to-Image provides professional developers and designers with unprecedented image control. Powered by robust semantic c - [grok-imagine-image/text-to-image](https://api.hermes-ai.net/models/grok-imagine-image-text-to-image-economy): `grok-imagine-image-text-to-image-economy` — Grok 4.2's text-to-image mode empowers creators to build magnificent visual worlds entirely from scratch. By simply inputting natural langua - [grok-imagine-image/image-to-image](https://api.hermes-ai.net/models/grok-imagine-image-image-to-image-economy): `grok-imagine-image-image-to-image-economy` — In the image-to-image mode, Grok 4.2 transforms into a highly controllable visual design engine. By uploading basic line art, composition sk ### Video - [gemini-omni-flash/image-to-video](https://api.hermes-ai.net/models/gemini-omni-flash-image-to-video-economy): `gemini-omni-flash-image-to-video-economy` — Gemini Omni Flash Image-to-Video turns reference images into dynamic short videos. It supports single-image video generation with 1 image an - [gemini-omni-flash/text-to-video](https://api.hermes-ai.net/models/gemini-omni-flash-text-to-video-economy): `gemini-omni-flash-text-to-video-economy` — Gemini Omni Flash is a unified video generation model that creates high-quality short videos from text prompts. It supports 720p, 1080p, and - [gemini-omni-flash/video-edit](https://api.hermes-ai.net/models/gemini-omni-flash-video-edit-economy): `gemini-omni-flash-video-edit-economy` — Omni Flash - All-in-One Video Image-to-video generation powered by reference images and videos. Compatible with 720p/1080p/4K resolutions . - [xai/grok-imagine-v1.5/text-to-video](https://api.hermes-ai.net/models/xai-grok-imagine-v1-5-text-to-video-economy): `xai-grok-imagine-v1-5-text-to-video-economy` — Leveraging xAI’s advanced reasoning capabilities, it excels at translating intricate prompts into visually stunning, logically coherent narr - [xai/grok-imagine-v1.5/image-to-video](https://api.hermes-ai.net/models/xai-grok-imagine-v1-5-image-to-video-economy): `xai-grok-imagine-v1-5-image-to-video-economy` — It is engineered to breathe life into static concepts while maintaining absolute subject identity. By deeply analyzing the geometric structu - [google/veo3.1-pro/start-end-to-video](https://api.hermes-ai.net/models/google-veo3-1-pro-start-end-to-video-channel-low-price): `google-veo3-1-pro-start-end-to-video-channel-low-price` — It represents the next evolution in cinematic video synthesis from DeepMind. It transforms still images or start-and-end frame pairs into hi - [google/veo3.1-pro/image-to-video](https://api.hermes-ai.net/models/google-veo3-1-pro-image-to-video-channel-low-price): `google-veo3-1-pro-image-to-video-channel-low-price` — the latest flagship video generation engine from Google, engineered to transform static concepts into high-fidelity cinematic visuals. This - [google/veo3.1-fast/start-end-to-video](https://api.hermes-ai.net/models/google-veo3-1-fast-start-end-to-video-channel-low-price): `google-veo3-1-fast-start-end-to-video-channel-low-price` — Veo 3.1 Fast is engineered for creators who prioritize speed and rapid iteration without sacrificing structural control. In Start & End Fram - [google/veo3.1-fast/image-to-video](https://api.hermes-ai.net/models/google-veo3-1-fast-image-to-video-channel-low-price): `google-veo3-1-fast-image-to-video-channel-low-price` — Google Veo3.1 I2V converts static images into cinematic dynamic videos with smooth realistic motion and natural lighting, delivering results - [google/veo3.1-pro/text-to-video](https://api.hermes-ai.net/models/google-veo3-1-pro-text-to-video-channel-low-price): `google-veo3-1-pro-text-to-video-channel-low-price` — Google's flagship advanced AI Text-to-Video model, Veo3.1 Premium mode. Enables native text-to-video with fully synchronized ambient sound, - [google/veo3.1-fast/text-to-video](https://api.hermes-ai.net/models/google-veo3-1-fast-text-to-video-channel-low-price): `google-veo3-1-fast-text-to-video-channel-low-price` — Google's latest advanced AI Text-to-Video model, Veo3.1 Fast mode. Features native text-to-video with synchronized audio & video generation, ### Audio - [suno-single-v5.5](https://api.hermes-ai.net/models/suno-single-v5-5): `suno-single-v5-5` — A single sentence describing the song's theme or mood. Output: Two alternative complete songs automatically generated, including melody, lyr - [suno-custom-v5.5](https://api.hermes-ai.net/models/suno-custom-v5-5): `suno-custom-v5-5` — Custom lyrics and style tags. Output: Two alternative songs generated following the instructions. v5.5 delivers Suno's highest-ever audio qu - [suno-single-v5](https://api.hermes-ai.net/models/suno-single-v5): `suno-single-v5` — A single sentence describing the song content. Output: Two complete songs automatically generated. The core upgrade of v5 is studio-grade au - [suno-custom-v5](https://api.hermes-ai.net/models/suno-custom-v5): `suno-custom-v5` — Custom lyrics and style tags. Output: Two alternative songs generated following the instructions. v5 enhances response precision for complex - [suno-single-v4.5](https://api.hermes-ai.net/models/suno-single-v4-5): `suno-single-v4-5` — Input a single sentence describing the song (style/mood/scene). The model automatically generates two alternative complete tracks. Compared - [suno-custom-v4.5](https://api.hermes-ai.net/models/suno-custom-v4-5): `suno-custom-v4-5` — Input custom song title, lyrics, and style tags. The model generates two alternative song audios following the instructions. The core upgrad ### Text - [suno-lyrics](https://api.hermes-ai.net/models/suno-lyrics): `suno-lyrics` — Theme prompt describing desired lyric topic. Output: Pure text lyrics with song structure (Verse/Chorus labels), song title, and style tags. ## Optional - [Model catalog JSON](https://api.hermes-ai.net/api/v1/models) - [Price list JSON](https://api.hermes-ai.net/api/v1/prices) - [Terms of Service](https://api.hermes-ai.net/terms) - [Privacy Policy](https://api.hermes-ai.net/privacy) --- # Hermes AI API Hermes AI API runs hosted image, video, music and text models through one asynchronous job API. Every call is a task: submit it, poll it, download the result. Tasks are paid from the user's prepaid credits, and only successful tasks are charged. ## Setup - If the `hermes-ai-api` MCP server is connected (https://api.hermes-ai.net/mcp), prefer its tools: `search_models`, `get_model`, `create_task`, `wait_for_task`, `get_task`, `list_tasks`, `get_balance`. - Otherwise use the REST API below with the key in `$HERMES_API_KEY`. If it is not set, ask the user to create one at https://api.hermes-ai.net/dashboard/api-keys and export it. Never print the key or write it into files. ## Workflow 1. **Pick a model.** `GET https://api.hermes-ai.net/api/v1/models` lists every model with its `slug`, `category` and input `fields`. Current prices: `GET https://api.hermes-ai.net/api/v1/prices`. A model without a price, or with `enabled: false`, cannot run. 2. **Build the input** from that model's `fields` only. Respect `required`, `options`, `default`, `min`/`max` and length limits. Unknown fields are rejected. Media fields take a public HTTPS URL, or upload a local file first and pass the returned `reference` (`media:`): `curl -X POST https://api.hermes-ai.net/api/v1/media -H "Authorization: Bearer $HERMES_API_KEY" -H "Content-Type: image/png" --data-binary @photo.png` 3. **Submit** with a fresh Idempotency-Key: ```bash curl -X POST https://api.hermes-ai.net/api/v1/jobs \ -H "Authorization: Bearer $HERMES_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{"model": "", "input": { ... }}' ``` The response (202) contains the task `id`. After a network error, retry with the same Idempotency-Key and body; it never charges twice. 4. **Poll** `GET https://api.hermes-ai.net/api/v1/jobs/` every 4–5 seconds until the status is terminal. 5. **Download** each `assets[].url` with the same Authorization header and save it with an extension that matches `assets[].type`. Text results are in `assets[].text`. ## Task statuses - `queued`: Accepted and waiting for a worker. - `submitting`: Being sent to the model provider. - `running`: The model is generating. - `succeeded` (terminal): Finished. Results are in assets. - `failed` (terminal): Finished without a result. The hold is released. - `reconciliation_required` (terminal): The result is being verified. Keep the task ID and do not resubmit automatically. Never resubmit a `reconciliation_required` task automatically; report its ID to the user. ## Errors Error bodies are `{ "code": "...", "error": "..." }`. Branch on `code`. | HTTP | Code | Meaning | | --- | --- | --- | | 401 | `AUTH_REQUIRED` | Missing or invalid API key or session. | | 403 | `INSUFFICIENT_SCOPE` | The API key lacks the scope this endpoint needs. | | 403 | `PERMISSION_DENIED` | This credential cannot perform the operation. | | 429 | `API_KEY_RATE_LIMITED` | The API key's rate limit was reached. Honor Retry-After. | | 429 | `API_KEY_USAGE_EXCEEDED` | The API key reached its usage limit. Create a new key. | | 400 | `INVALID_REQUEST` | The body is not a JSON object or a required value is missing. | | 400 | `INVALID_IDEMPOTENCY_KEY` | Idempotency-Key must be 1–128 printable ASCII characters. | | 400 | `INVALID_MODEL_INPUT` | input does not match the model schema (unknown field, missing required field, or invalid value). | | 400 / 413 | `INPUT_TOO_LARGE` | The request body is larger than 256 KB. Upload media first and send a media reference. | | 404 | `REQUEST_NOT_FOUND` | No chat request with this ID belongs to your account. | | 404 | `MODEL_NOT_FOUND` | No model has this slug. | | 404 | `JOB_NOT_FOUND` | No task with this ID belongs to your account. | | 409 | `IDEMPOTENCY_CONFLICT` | This Idempotency-Key was used with different input. Use a new key for a new request. | | 409 | `PRICE_CHANGED` | The price differs from expectedPriceMicros. Fetch the price and confirm again. | | 409 | `MODEL_PRICE_UNAVAILABLE` | No price is configured for this model yet, so it cannot run. | | 409 | `PAYMENT_REVIEW_REQUIRED` | A refund left an uncovered balance. Contact support or cover the balance before submitting tasks. | | 409 | `INSUFFICIENT_CREDITS` | Available credits do not cover the task price. | | 429 | `RATE_LIMITED` | Too many task submissions per minute. Retry later. | | 429 | `CONCURRENCY_LIMITED` | Too many unfinished tasks. Wait for one to finish. | | 503 | `EXECUTION_UNAVAILABLE` | Execution is not configured for this model right now. | | 429 | `UPSTREAM_RATE_LIMITED` | The model provider is rate limiting chat requests. Retry with backoff; nothing was charged. | | 502 | `UPSTREAM_ERROR` | The model provider failed the chat request. Nothing was charged; retry later. | | 503 | `PLATFORM_BUDGET_EXCEEDED` | This model is temporarily unavailable. Try again later. | | 503 | `EXECUTION_TEMPORARILY_UNAVAILABLE` | Execution is temporarily unavailable. Retry with the same Idempotency-Key. | | 503 | `AUTH_UNAVAILABLE` | Authentication is temporarily unavailable. Retry later. | | 503 | `API_UNAVAILABLE` | The API is temporarily unavailable. Retry later. | ## Guidance - Tell the user the model and price before running an expensive task, and check the balance with `GET https://api.hermes-ai.net/api/v1/usage` when unsure. - Video and music tasks can take minutes; keep polling and report progress. - Full reference: https://api.hermes-ai.net/docs · Models: https://api.hermes-ai.net/models · OpenAPI: https://api.hermes-ai.net/api/openapi.json