跳到主要内容

// 模型目录

为你的任务找到合适的模型

图像、视频和音乐模型走任务 API,语言模型走 OpenAI 兼容的对话 API。按产出类型筛选,查看输入和价格,然后打开 Playground 或复制请求。

11 个模型

清除筛选
Google检查中

gemini-omni-flash/image-to-video

Gemini Omni Flash Image-to-Video turns reference images into dynamic short videos. It supports single-image video generation with 1 image and reference fusion with 3 images, offering 720p, 1080p, and 4k outputs with 4, 6, 8, and 10-second durations for character animation, product showcases, creative shots, and multi-reference video creation.

图生视频/文本 + 图像 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

gemini-omni-flash/text-to-video

Gemini Omni Flash is a unified video generation model that creates high-quality short videos from text prompts. It supports 720p, 1080p, and 4k outputs with 4, 6, 8, and 10-second durations, making it suitable for creative clips, advertising assets, social videos, and visual concept demos.

文生视频/文本 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

gemini-omni-flash/video-edit

Omni Flash - All-in-One Video Image-to-video generation powered by reference images and videos. Compatible with 720p/1080p/4K resolutions . Great for character motion, product display, creative shots and multi-image referenced videos.

视频编辑/文本 + 图像 + 视频 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
xAI检查中

xai/grok-imagine-v1.5/text-to-video

Leveraging xAI’s advanced reasoning capabilities, it excels at translating intricate prompts into visually stunning, logically coherent narratives. The model demonstrates superior performance in managing cinematic camera movements and complex physical interactions, ensuring that every frame adheres to realistic physics and lighting. It is designed for creators who demand high-precision storytelling and unparalleled visual fidelity without semantic loss.

文生视频/文本 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
xAI检查中

xai/grok-imagine-v1.5/image-to-video

It is engineered to breathe life into static concepts while maintaining absolute subject identity. By deeply analyzing the geometric structure and material properties of a reference image, it synthesizes motion that aligns perfectly with physical intuition. The model excels in temporal stability and lighting inheritance, ensuring that core elements remain undistorted during intense transitions. It provides a seamless bridge from a single visual reference to a high-tension, cinematic sequence characterized by fluid motion and professional-grade rendering.

图生视频/文本 + 图像 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

google/veo3.1-pro/start-end-to-video

It represents the next evolution in cinematic video synthesis from DeepMind. It transforms still images or start-and-end frame pairs into high-fidelity 1080p motion sequences with stunning temporal continuity. Standing out with its native audio generation, the model automatically crafts synchronized soundscapes that breathe life into the visuals. Whether executing complex camera dollies or seamless scene morphing, Veo 3.1 delivers professional-grade consistency and narrative depth for storyboarding and creative production.

图生视频/文本 + 图像 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

google/veo3.1-pro/image-to-video

the latest flagship video generation engine from Google, engineered to transform static concepts into high-fidelity cinematic visuals. This model creates stunning videos at up to 4K resolution. It features outstanding subject consistency and built-in audio generation, accurately reproducing the textures and lighting of reference images while producing realistic ambient sounds in sync. Additionally, it supports first-and-last frame guidance and video extension, delivering director-level transitions and spatial stability for clips up to 8 seconds long.

图生视频/文本 + 图像 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

google/veo3.1-fast/start-end-to-video

Veo 3.1 Fast is engineered for creators who prioritize speed and rapid iteration without sacrificing structural control. In Start & End Frame mode, it delivers near-instantaneous interpolation, bridging two visual anchors with fluid motion in seconds. This model is optimized for low-latency workflows and high-volume prototyping. While maximizing throughput, it maintains impressive consistency in geometry and motion logic.

图生视频/文本 + 图像 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

google/veo3.1-fast/image-to-video

Google Veo3.1 I2V converts static images into cinematic dynamic videos with smooth realistic motion and natural lighting, delivering results 30% faster than the standard version. Preserves original image composition and visual style, enables native synchronized audio generation, supports dialogue & lip-sync, ideal for social content creation, concept visualization and casual creative storytelling with high cost performance.

图生视频/文本 + 图像 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

google/veo3.1-pro/text-to-video

Google's flagship advanced AI Text-to-Video model, Veo3.1 Premium mode. Enables native text-to-video with fully synchronized ambient sound, dialogue and music, supports dialogue lip-sync, subject consistency and video interpolation. Generates top-tier cinematic videos with natural lighting, smooth camera transitions and strong narrative consistency, full flagship functions for professional storytelling and marketing, ultra-premium quality with ultra-high pricing.

文生视频/文本 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证
Google检查中

google/veo3.1-fast/text-to-video

Google's latest advanced AI Text-to-Video model, Veo3.1 Fast mode. Features native text-to-video with synchronized audio & video generation, delivers high-quality videos with basic cinematic realism and smooth motion. Equipped with natural scene presentation and accurate audio-visual synchronization, outstanding quality at an ultra-low price, the best cost-effective choice for daily creative needs and casual video generation scenarios.

文生视频/文本 → 视频

Hermes AI 价格
正在加载价格…
官方原价 · 未验证