
Grok Imagine is xAI's image and video model family: $0.02 fast images, an Imagine Quality tier, natural-language editing, and video (1.0 and 1.5) generating 1-15 second clips at up to 1080p with one-pass audio for $0.08/second. Its flagship skill is image-to-video source fidelity. All eight variants run inside invideo alongside 200+ other models.
Updated August 2026
Grok Imagine's standout skill is image-to-video — animating an existing still while preserving its detail and lighting rather than reinterpreting it — and it delivers that as one of the cheapest frontier video models on the market, at $0.08 per second for Video 1.5 (per xAI's official pricing, August 2026). The family, from xAI, spans a fast image model at $0.02 per image, a higher-fidelity Imagine Quality tier, natural-language image editing, and a video branch (Video 1.0 and 1.5) that generates clips of 1–15 seconds with sound effects, ambience, and dialogue produced in the same pass — no separate audio step. Below is the full family — every variant, its specs, what each is for, and how to prompt it.
The full Grok Imagine spec sheet
| Spec | Image branch | Video branch |
|---|---|---|
| Models | Grok Imagine ($0.02/image), Imagine Quality ($0.05/image) | Video 1.0 ($0.05/sec), Video 1.5 ($0.08/sec) |
| Resolution | 1k and 2k presets (exact pixel dimensions not published) |
480p (default), 720p, 1080p |
| Duration | — | 1–15 seconds per generation; longer clips via Extend |
| Aspect ratios | 14 options including 1:1, 16:9, 9:16, 2:1 — plus phone-screen ratios 19.5:9, 9:19.5, 20:9, 9:20 that few other image APIs expose | 16:9 (default), 9:16, 1:1, 4:3, 3:4, 3:2, 2:3 |
| Audio | — | Native one-pass generation: SFX, ambience, and dialogue synced to the action |
| Frame rate | — | Not officially documented by xAI (third-party "24 fps" claims are unverified) |
| Editing | Text-prompted edits; up to 3 source images per multi-image edit | Video Edit runs on the 1.0 model, capped at 8.7 seconds and 720p |
| Character references | — | Up to 7 reference images; up to 3 preset voices; output capped at 720p |
| Batch | Up to 10 images per request | Async generation; batch-result URLs expire after 1 hour |
Sources: xAI's image generation, video generation, and reference-to-video documentation, as of August 2026.
Read the table for its asymmetries: generation got the 1.5 upgrade to 1080p, while editing and character references stay fenced at 1.0-era caps — 8.7 seconds and 720p — a pattern the rest of this page keeps running into.
Which Grok Imagine variant should you pick?
The family shows up as eight entries in invideo's model picker. Here is what each one actually is:
| Picker entry | What it does | Pick it when |
|---|---|---|
| Grok Imagine | Fast/standard image generation, $0.02/image | Volume image work, drafts, first frames for video |
| Grok Imagine Pro | Deprecated alias — since roughly May 2026 it resolves to Imagine Quality ($0.05/image), per the official model card | You want the quality tier; "Pro" and "Quality" are now the same model |
| Grok Imagine Image Edit | Natural-language edits to existing images; up to 3 source images per request | Retouching, compositing, restyling without regenerating |
| Grok Imagine Video | Video 1.0 — text- or image-to-video at $0.05/sec, 480/720p | Budget clips, and it is still the engine behind Video Edit |
| Video 1.5 | The flagship: better motion and physics, clearer synced speech, native 1080p, $0.08/sec | Almost everything — final-quality clips, text-to-video, 1080p |
| Video Edit | Prompt-based edits to an existing clip. Runs on the 1.0 model: output capped at 8.7 seconds and 720p | Small changes to a clip you already like |
| Video Extend | Adds seconds onto an existing generation | Building sequences past the 15-second single-generation limit |
| Reference to Video | Video 1.5 with image and preset-voice references for consistent characters | Serialized content where the same face (and voice) recurs |
One constraint that catches people: a first-frame image and reference_images cannot be combined in the same request — you animate a still or you use character references, not both (xAI docs).
What Grok Imagine does well
Image-to-video source fidelity. This is the family's flagship skill. xAI's own model card says the model "holds detail and lighting from the input frame, so the result continues the original image rather than reinterpreting it," and that it "works well for sequences — stage each frame, animate it, chain shots together." The blind-test record backed this up at launch: Video 1.5 debuted at #1 on Artificial Analysis' Image-to-Video Arena in June 2026 at roughly 1473 Elo, a ~52-point jump over Video 1.0, passing Seedance 2.0, Sora 2, Veo 3.1, and Kling (Blockchain.News, June 2026).
Speed and cost. Video 1.5 Fast produces a 6-second 720p clip in around 25 seconds (TechTimes, June 2026), and at $0.08/second with audio included, it undercuts most frontier video models by a wide margin — xAI's January 2026 API launch marketed it as roughly 86% cheaper per minute than Sora 2 Pro.
One-pass audio. Sound effects, ambience, and dialogue are generated with the video, landing on the action, rather than bolted on afterward — per xAI's Video 1.5 announcement, speech is "clearer and better synced" than 1.0.
Where Grok Imagine falls short
The "#1 video model" claim has expired. By August 2026, Artificial Analysis' Image-to-Video Arena shows Video 1.5 at #4 in the no-audio bracket (~1328 Elo, behind Gemini Omni Flash, MiniMax H3, and Seedance 2.0), and notably weaker in the with-audio bracket (~1114). Any ranking you read about this model needs a date attached.
Feature asymmetries. Video Edit never got the 1.5 upgrade — it still runs the 1.0 model with an 8.7-second/720p ceiling. Reference-to-video caps at 720p even though plain generation goes to 1080p. Voice references use a preset roster only, capped at 3 voices, with API access limited to US trusted partners; you cannot upload your own audio (reference docs). That tight fencing on voice features reflects xAI's response to criticism of Grok Imagine's content moderation at its August 2025 consumer launch.
Undocumented frame rate. xAI publishes no fps figure anywhere in its docs — an unusual omission for a video model, and every "24 fps" figure circulating in third-party coverage is unsourced.
How to prompt Grok Imagine video
- Lead with subject + action. Put who is doing what in the first sentence, then camera, then atmosphere and lighting, then audio.
- One camera move per clip. The model understands cinematography vocabulary — dolly, push-in, orbit, pan, handheld, crane, tracking, rack focus — but stacking moves produces mush. To lock the frame entirely, write "camera not moving"; "stable camera" still tends to drift.
- Name concrete sounds. "Product click and low bass" or "cloth movement and footsteps" gets synced foley; vague cues like "cinematic sound" trigger generic background music.
- For image-to-video, start from a strong still. Generate the frame with the image model first, keep the motion prompt short, then chain shots with Extend — this staged workflow is what xAI itself recommends for sequences.
- Tag references inline. Reference-to-video uses explicit tags in the prompt —
<IMAGE_1>for a reference image,<AUDIO_0>for a preset voice — so the model knows which entity does what (reference docs).
What can you make with Grok Imagine?
- Stills, thumbnails, and wallpapers. The $0.02 price and the phone-screen 19.5:9/20:9 ratios make the image branch a natural fit for high-volume work in an AI image generator workflow.
- Animating existing art and photos. Source fidelity is the family's arena-proven strength, which is exactly the image-to-video job: stage a frame, animate it, extend it.
- Product and e-commerce clips. xAI's own listed reference-to-video use cases are "virtual try-on, product placement, character-consistent storytelling, and voice identity" — a direct match for product video work where the same item must appear accurately in every shot.
- Fast social-first video. Sub-30-second generation with native audio makes it one of the lowest-friction models for short-form AI video generation.
Grok Imagine FAQ
How long can a Grok Imagine video be?
Up to 15 seconds per generation (the duration parameter accepts 1–15). Longer clips are built with Video Extend — and note the semantics: the extension's duration parameter specifies only the added portion, so a 10-second input extended by 5 seconds yields a 15-second output. xAI documents no hard cap on total chained length (extension docs).
Can I use Grok Imagine output commercially?
Yes — xAI's terms of service assign output ownership to the user, with commercial use permitted on paid tiers and no revenue thresholds. Two caveats: xAI may use prompts and outputs for training, and it offers no IP indemnification, so infringement risk sits with you. Review the current terms before relying on this for client work, as clauses change.
Do Grok Imagine videos have a watermark?
Reports indicate a visible Grok watermark on lower consumer tiers, removed on higher subscription tiers; xAI has not published a definitive watermark policy page, and its terms reserve the right to apply AI-generated-content disclosures. There is no published C2PA or cryptographic provenance commitment as of August 2026.
Does Grok Imagine generate audio, and can I use my own voice?
Audio is native and one-pass — SFX, ambience, and dialogue arrive with the video. Voices, however, come only from xAI's built-in preset roster (up to 3 per request, tagged <AUDIO_0>–<AUDIO_2>); custom audio uploads are not supported, and API voice references are limited to US trusted partners (reference docs).
Why is my edited video capped at 8.7 seconds?
Video Edit still runs on the original 1.0 model, which hard-caps edit output at 8.7 seconds and 720p regardless of the input clip's length or any duration you request (editing docs). To change a longer clip, regenerate or re-extend instead.
What is the difference between Grok Imagine Pro and Imagine Quality?
Nothing anymore. The grok-imagine-image-pro model was deprecated around May 2026 and now exists only as an alias that resolves to grok-imagine-image-quality ($0.05/image), per the official model card. Coverage that lists "Pro" as a separate live model is out of date.
What frame rate does Grok Imagine output?
xAI does not document it. The 24 fps figure that circulates in third-party writeups appears nowhere in official specs, so treat it as unconfirmed.
Where can you use Grok Imagine?
The whole family — both image tiers, edit, and every video variant through Reference to Video — is available inside invideo, where the agent sits it alongside the other 200+ models (Veo 3.1, Sora 2, Kling 3.0, Seedance 2.0, and the rest) and can sequence shots across them: a Grok Imagine still animated by Video 1.5, say, with other models picking up the shots they are better at. The Grok Imagine hub on invideo covers the family in the context of the platform, and the image and image-to-video workflows above are the fastest ways to put its two strongest skills to work without touching the API.
Version history: Grok Imagine consumer launch Aug 2025 (480p) → API launch with Video 1.0 and revamped image generation Jan 2026 → Video 1.5 preview May 31, 2026, GA Jun 16, 2026 → References, true text-to-video, and native 1080p added to 1.5 late Jul 2026. Imagine Pro deprecated to an alias of Imagine Quality ~May 2026.