grok-imagine-video
via InferenceSaver Gateway
200K
16K
$1.00
$1.00
About grok-imagine-video
Grok Imagine Video is a video generation model available through InferenceSaver that turns a text prompt or a reference image into a short video clip.
It generates video from a text prompt, or from a source image where image-conditioned generation is supported.
Output resolution and clip length are configurable per request, so you can trade off render time and cost against visual fidelity and duration.
Key Features
Video Generation
Generates video clips from a text prompt or reference image.
Capabilities
| Feature | Supported |
|---|---|
Vision Support Process and analyze images, charts, and visual content | No |
Function Calling Execute custom functions and integrate with external tools | No |
Streaming Real-time token-by-token response generation | No |
Structured Output Generate responses conforming to JSON schemas | No |
Video Generation Generate and process video content | Yes |
Context and Token Limits
| Property | Value | Description |
|---|---|---|
| Context Window | 200K | Maximum input tokens the model can process at once |
| Max Output Tokens | 16K | Maximum tokens the model can generate in a single response |
Token Usage Note
Tokens can be words or parts of words. On average, 1 token is approximately 4 characters or 0.75 words in English. The actual token count depends on the specific text and language.
Best Practices
Prompt Engineering
Describe the scene, subject motion, and camera behavior explicitly. Specific, concrete prompts produce more predictable results than abstract ones.
Choosing Resolution and Duration
Lower resolutions and shorter durations render faster and cost less, useful for quick iteration. Reserve 1080p and longer durations for final output.
Image-to-Video Inputs
For Image-to-Video, use a clear, well-composed source image — the model animates from it, so composition and framing carry through to the output clip.