# Hermes AI API > One API for image, video and music AI models. Submit an asynchronous task with one request > shape for every model, poll for the result, and pay a per-task quote from prepaid credits. Base URL: https://api.hermes-ai.net. Authenticate with `Authorization: Bearer ` (keys start with `hermes_`), or connect an MCP client to https://api.hermes-ai.net/mcp and sign in with OAuth. ## Docs - [API guide](https://api.hermes-ai.net/docs): quickstart, authentication, tasks, results, errors and pricing - [MCP and Agent Skill](https://api.hermes-ai.net/agents): connect Claude Code, Cursor, Codex and other agents - [Agent Skill](https://api.hermes-ai.net/skills/hermes-ai-api/SKILL.md): SKILL.md with the full task workflow for agents - [OpenAPI 3.1](https://api.hermes-ai.net/api/openapi.json): machine-readable API, including each model's input schema - [Pricing](https://api.hermes-ai.net/pricing): current per-task prices ## Models ### Image - [nano-banana-pro/text-to-image](https://api.hermes-ai.net/models/nano-banana-pro-text-to-image-economy): `nano-banana-pro-text-to-image-economy` — Google's Nano Banana Pro (Gemini 3.0 Pro Image) is an industry-leading cutting-edge text-to-image model, which can generate high-definition - [nano-banana-pro/edit](https://api.hermes-ai.net/models/nano-banana-pro-edit-economy): `nano-banana-pro-edit-economy` — Google Nano Banana Pro (Gemini 3.0 Pro Image) Edit supports professional image editing with high-quality 4K-capable ultra-clear output, driv - [nano-banana-2-lite/image-to-image](https://api.hermes-ai.net/models/nano-banana-2-lite-image-to-image-economy): `nano-banana-2-lite-image-to-image-economy` — It delivers ultra-low-latency and cost-effective image generation & editing capabilities. Optimized to achieve latency under 2 seconds and d - [nano-banana-2-lite/text-to-image](https://api.hermes-ai.net/models/nano-banana-2-lite-text-to-image-economy): `nano-banana-2-lite-text-to-image-economy` — Lightweight text-to-image API built for high concurrency & instant response, with low-latency, budget-friendly image generation and editing. - [nano-banana-2/image-to-image](https://api.hermes-ai.net/models/nano-banana-2-image-to-image-economy): `nano-banana-2-image-to-image-economy` — An image-to-image and editing endpoint powered by a highly efficient visual engine. It enables rapid style transfer, inpainting, or backgrou - [nano-banana-2/text-to-image](https://api.hermes-ai.net/models/nano-banana-2-text-to-image-economy): `nano-banana-2-text-to-image-economy` — A lightweight text-to-image endpoint engineered for high-concurrency and rapid response. As the core of Nano Banana 2, it balances visual fi - [gpt-image-2.5/sunburst/image-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-sunburst-image-to-image-economy): `gpt-image-2-5-sunburst-image-to-image-economy` — GPT Image 2.5 Sunburst Edit provides precision-first revisions while preserving subjects, composition, and visual identity. - [gpt-image-2.5/sunburst/text-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-sunburst-text-to-image-economy): `gpt-image-2-5-sunburst-text-to-image-economy` — GPT Image 2.5 Sunburst focuses on intricate detail, reliable typography, and controlled composition. - [gpt-image-2.5/flare/image-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-flare-image-to-image-economy): `gpt-image-2-5-flare-image-to-image-economy` — GPT Image 2.5 Flare Edit supports fast everyday revisions while preserving subjects, composition, and background. - [gpt-image-2.5/flare/text-to-image](https://api.hermes-ai.net/models/gpt-image-2-5-flare-text-to-image-economy): `gpt-image-2-5-flare-text-to-image-economy` — GPT Image 2.5 Flare is a fast, versatile image generator for social content, commerce assets, ideation, and high-volume creative work. - [xai/grok-imagine-2.0/text-to-image](https://api.hermes-ai.net/models/xai-grok-imagine-2-0-text-to-image-economy): `xai-grok-imagine-2-0-text-to-image-economy` — Grok Imagine 2.0 low-price channel edition provides async text-to-image generation with seven common aspect ratios and 1–12 images per reque - [gpt-image-2.0/text-to-image](https://api.hermes-ai.net/models/gpt-image-2-0-text-to-image-economy): `gpt-image-2-0-text-to-image-economy` — The GPT-Image-2 Text-to-Image is a state-of-the-art generation foundation designed for high-standard commercial scenarios. It features revol - [gpt-image-2.0/edit](https://api.hermes-ai.net/models/gpt-image-2-0-edit-economy): `gpt-image-2-0-edit-economy` — The GPT-Image-2 Image-to-Image provides professional developers and designers with unprecedented image control. Powered by robust semantic c - [grok-imagine-image/text-to-image](https://api.hermes-ai.net/models/grok-imagine-image-text-to-image-economy): `grok-imagine-image-text-to-image-economy` — Grok 4.2's text-to-image mode empowers creators to build magnificent visual worlds entirely from scratch. By simply inputting natural langua - [grok-imagine-image/image-to-image](https://api.hermes-ai.net/models/grok-imagine-image-image-to-image-economy): `grok-imagine-image-image-to-image-economy` — In the image-to-image mode, Grok 4.2 transforms into a highly controllable visual design engine. By uploading basic line art, composition sk ### Video - [gemini-omni-flash/image-to-video](https://api.hermes-ai.net/models/gemini-omni-flash-image-to-video-economy): `gemini-omni-flash-image-to-video-economy` — Gemini Omni Flash Image-to-Video turns reference images into dynamic short videos. It supports single-image video generation with 1 image an - [gemini-omni-flash/text-to-video](https://api.hermes-ai.net/models/gemini-omni-flash-text-to-video-economy): `gemini-omni-flash-text-to-video-economy` — Gemini Omni Flash is a unified video generation model that creates high-quality short videos from text prompts. It supports 720p, 1080p, and - [gemini-omni-flash/video-edit](https://api.hermes-ai.net/models/gemini-omni-flash-video-edit-economy): `gemini-omni-flash-video-edit-economy` — Omni Flash - All-in-One Video Image-to-video generation powered by reference images and videos. Compatible with 720p/1080p/4K resolutions . - [xai/grok-imagine-v1.5/text-to-video](https://api.hermes-ai.net/models/xai-grok-imagine-v1-5-text-to-video-economy): `xai-grok-imagine-v1-5-text-to-video-economy` — Leveraging xAI’s advanced reasoning capabilities, it excels at translating intricate prompts into visually stunning, logically coherent narr - [xai/grok-imagine-v1.5/image-to-video](https://api.hermes-ai.net/models/xai-grok-imagine-v1-5-image-to-video-economy): `xai-grok-imagine-v1-5-image-to-video-economy` — It is engineered to breathe life into static concepts while maintaining absolute subject identity. By deeply analyzing the geometric structu - [google/veo3.1-pro/start-end-to-video](https://api.hermes-ai.net/models/google-veo3-1-pro-start-end-to-video-channel-low-price): `google-veo3-1-pro-start-end-to-video-channel-low-price` — It represents the next evolution in cinematic video synthesis from DeepMind. It transforms still images or start-and-end frame pairs into hi - [google/veo3.1-pro/image-to-video](https://api.hermes-ai.net/models/google-veo3-1-pro-image-to-video-channel-low-price): `google-veo3-1-pro-image-to-video-channel-low-price` — the latest flagship video generation engine from Google, engineered to transform static concepts into high-fidelity cinematic visuals. This - [google/veo3.1-fast/start-end-to-video](https://api.hermes-ai.net/models/google-veo3-1-fast-start-end-to-video-channel-low-price): `google-veo3-1-fast-start-end-to-video-channel-low-price` — Veo 3.1 Fast is engineered for creators who prioritize speed and rapid iteration without sacrificing structural control. In Start & End Fram - [google/veo3.1-fast/image-to-video](https://api.hermes-ai.net/models/google-veo3-1-fast-image-to-video-channel-low-price): `google-veo3-1-fast-image-to-video-channel-low-price` — Google Veo3.1 I2V converts static images into cinematic dynamic videos with smooth realistic motion and natural lighting, delivering results - [google/veo3.1-pro/text-to-video](https://api.hermes-ai.net/models/google-veo3-1-pro-text-to-video-channel-low-price): `google-veo3-1-pro-text-to-video-channel-low-price` — Google's flagship advanced AI Text-to-Video model, Veo3.1 Premium mode. Enables native text-to-video with fully synchronized ambient sound, - [google/veo3.1-fast/text-to-video](https://api.hermes-ai.net/models/google-veo3-1-fast-text-to-video-channel-low-price): `google-veo3-1-fast-text-to-video-channel-low-price` — Google's latest advanced AI Text-to-Video model, Veo3.1 Fast mode. Features native text-to-video with synchronized audio & video generation, ### Audio - [suno-single-v5.5](https://api.hermes-ai.net/models/suno-single-v5-5): `suno-single-v5-5` — A single sentence describing the song's theme or mood. Output: Two alternative complete songs automatically generated, including melody, lyr - [suno-custom-v5.5](https://api.hermes-ai.net/models/suno-custom-v5-5): `suno-custom-v5-5` — Custom lyrics and style tags. Output: Two alternative songs generated following the instructions. v5.5 delivers Suno's highest-ever audio qu - [suno-single-v5](https://api.hermes-ai.net/models/suno-single-v5): `suno-single-v5` — A single sentence describing the song content. Output: Two complete songs automatically generated. The core upgrade of v5 is studio-grade au - [suno-custom-v5](https://api.hermes-ai.net/models/suno-custom-v5): `suno-custom-v5` — Custom lyrics and style tags. Output: Two alternative songs generated following the instructions. v5 enhances response precision for complex - [suno-single-v4.5](https://api.hermes-ai.net/models/suno-single-v4-5): `suno-single-v4-5` — Input a single sentence describing the song (style/mood/scene). The model automatically generates two alternative complete tracks. Compared - [suno-custom-v4.5](https://api.hermes-ai.net/models/suno-custom-v4-5): `suno-custom-v4-5` — Input custom song title, lyrics, and style tags. The model generates two alternative song audios following the instructions. The core upgrad ### Text - [suno-lyrics](https://api.hermes-ai.net/models/suno-lyrics): `suno-lyrics` — Theme prompt describing desired lyric topic. Output: Pure text lyrics with song structure (Verse/Chorus labels), song title, and style tags. ## Optional - [Model catalog JSON](https://api.hermes-ai.net/api/v1/models) - [Price list JSON](https://api.hermes-ai.net/api/v1/prices) - [Terms of Service](https://api.hermes-ai.net/terms) - [Privacy Policy](https://api.hermes-ai.net/privacy)