Grok Imagine Image 2.0: complete guide, features, prompts, and editing

Grok Imagine Image 2.0: complete guide, features, prompts, and editing

The complete Grok Imagine Image 2.0 guide: designer-grade typography, magic wand region edits, multi-reference generation, Smart Resize, and prompt examples.

Grok Imagine Image 2.0 features and capabilities

Grok Imagine Image 2.0 is xAI's precision image model. The recurring theme is control: instructions land as written, typography gets planned rather than painted, edits touch one region at a time, and one finished image recomposes into every format you ship.

Image 2.0 is generally available as the Quality Mode on grok.com/imagine and in the Grok iOS and Android apps. API access is coming soon.

FeatureWhat it doesBest for
Instruction-faithful generationFollows a dense, multi-part brief down to the detailsExact layouts, art direction
Typography and layout planningPlans type like a designer, keeps small text sharpPosters, infographics, packaging
Magic wand and segmentationEdits only the region you point at, rest untouchedRetouching, product swaps
Multi-reference editingBlends up to 5 input images in one generationComposites, style transfers
Smart ResizeRecomposes one image into 9 aspect ratiosCampaign formats, banners, stories

Instruction-faithful generation

Image 2.0 was built around one goal: images you can use in real work. That starts with fidelity to the brief. A prompt with a subject, a layout, exact wording, and a lighting note comes back with all four honored, so complex multi-element directions land on the first pass more often, and revisions become refinements rather than re-rolls.

Typography and layout planning

Most image models paint letters; Image 2.0 plans them. It works out the type hierarchy and the layout the way a designer would before rendering, so dense, multi-part visuals like tutorial sheets, infographics, and title screens hold their structure, and small text comes out sharp instead of smudged. Spell the exact words in quotes and say where they sit, and they render as written.

Magic wand, segmentation, and background removal

Real work is iterative, so editing is region-precise rather than a full re-roll. The magic wand edits the region you point at and leaves the rest untouched. Segmentation selects precise areas of the image to change, one garment, one label, one surface. Background removal exports any subject with a transparent background, ready to drop into other work.

Multi-reference editing

A single generation accepts up to 5 input images. That turns compositing into a prompt: one reference carries the product, another the style, another the scene, and the model fuses them in one pass instead of you assembling the result by hand. It preserves what you put in, so the product in the output is recognizably the product you supplied.

Smart Resize

One image, any size. Pick a ratio and the model fills in the frame, recomposing rather than cropping, across 9 formats from a 1:2 tall banner through 9:16, square, and 16:9 out to a 2:1 wide banner. A finished hero visual becomes a full campaign set without regenerating the concept from scratch.

Templates and world building

Templates package common jobs, photo edits, product color changes, e-commerce photos, headshots, icons, mascots, game assets, and merch, into ready-made starting points where you supply the inputs. And because the model holds a look across generations, you can build a world image by image: a character, her locations, and the props she carries, generated separately and holding one style, ready to feed a video pipeline.

Grok Imagine Image 2.0 use cases

Posters with working type

Set the headline, the subline, and the hierarchy in the prompt and the poster comes back laid out, with the words rendered as written. Small print stays legible instead of dissolving into texture.

A farmers market poster with the headline SUNDAY HARVEST, date and venue lines, and legible small print

Infographics and tutorial sheets

Dense, multi-part visuals hold their logic: labeled steps, arrows that point at the right things, and captions that read. A process diagram or a step-by-step sheet arrives structured in one pass.

A field-guide infographic of six alpine wildflowers, each labeled with a short caption

Product and e-commerce edits

Point at the label, the colorway, or the background and change just that. The lighting, shadows, and the rest of the frame stay put, which is what makes catalog work practical.

Before and after of the same perfume bottle, cream paper label changed to matte black

One visual, every format

Smart Resize recomposes a finished image into stories, banners, squares, and widescreen slides. The model fills each new frame rather than cropping into it, so every format reads composed.

The same harbor scene recomposed as a tall banner, a square, and a widescreen frame

A world that holds one style

A character, her locations, and her props, generated separately across prompts and holding one look. The consistency carries into edits, so a world sheet stays recognizable as it grows.

The same red-cloaked courier as a portrait, in a market street, and on a mountain pass, one painted style

Photo fixes and headshots

Magic wand retouching, background removal, and headshot templates turn an existing photo into a finished asset. The edit touches what you asked for and nothing else.

Before and after of the same woman, a dim casual snapshot beside a clean studio headshot

How to get the best out of Grok Imagine Image 2.0

Image 2.0 rewards a structured brief and a habit of editing in place rather than re-rolling. A few practices carry most of the quality:

  • Write a brief, not a caption. Subject, layout, exact words, style, light. The model follows details, so give it details.
  • Put on-image text in quotes. The exact words render as written, which is the whole point for posters, packaging, and labels.
  • Describe the layout structurally. "Headline across the top third, product lower right, caption under it" beats "a nice poster layout".
  • Point, don't re-roll. For a change to an existing image, magic wand the region and describe only the change.
  • One edit at a time. Scoped changes keep the rest of the frame stable and stack more predictably than one sweeping instruction.
  • Name what each reference contributes. With up to 5 inputs, say which carries the subject, which the style, and which the scene.
  • Finish first, then resize. Lock the hero image, then let Smart Resize recompose it into the other formats you ship.
  • Reach for a template on recurring jobs. Headshots, product shots, and icons have the workflow pre-configured.

For the full capability list and specifications, see the Grok Imagine Image 2.0 model page.

Grok Imagine Image 2.0 prompt guide

A strong Image 2.0 prompt reads like a short design brief. The model plans typography and layout before it renders, so the prompt's job is to give it something to plan: what is in the frame, how it is arranged, what the words say, and how it is lit.

ElementWhat it coversExample
SubjectWho or what is in frame, described concretelyA ceramic teapot on a linen cloth
LayoutWhere things sit and how they relateProduct centered, caption band along the bottom
Exact textThe on-image words, in quotesThe label reads 'HARVEST NO. 3'
StyleMedium, palette, and moodFlat editorial illustration, muted greens
LightDirection and quality of lightSoft window light from the left

A worked example:

A recipe card for lemon shortbread. Title "LEMON SHORTBREAD" across the top
in a serif, five numbered steps down the left side with short captions, a
finished biscuit photo lower right. Cream background, thin rule lines,
soft even light. Flat, printable, no clutter.

Editing prompts

An edit is a scoped instruction, not a new brief. Select the region with the magic wand or segmentation, describe only the change, and say what stays:

Change the armchair's upholstery to forest green corduroy. Keep the light,
the shadows, the floor, and everything else exactly as it is.

The model preserves what you put in, so the discipline is on your side: one region, one change, and the rest of the frame carries over. Chain small edits rather than bundling five changes into one instruction, and you keep every version usable along the way.

Multi-reference prompts

With up to 5 input images, assign each a role instead of letting the model guess:

Image 1 is the product: keep its shape, label, and color exact.
Image 2 sets the style: match its palette and grain.
Image 3 is the scene: place the product on this counter, in this light.

The naming does the work. A reference without a role bleeds everything into the result, its background, its palette, its composition, while a reference with a role contributes just the attribute you asked for.

Smart Resize workflow

Treat resizing as a finishing step. Generate and refine the hero image at the ratio you designed for, then recompose the keeper into each format you ship. Check the type after each resize: the model re-plans the layout to fill the new frame, and a tall 9:16 recomposition may re-stack elements that sat side by side in 16:9. That re-stacking is the feature, but headlines deserve a glance before export.

Weak vs strong prompts

Name the layout, the exact words, and the scope of an edit rather than leaving them to chance.

FocusWeakStrong
TypographyA poster for a jazz nightA jazz night poster, headline 'BLUE HOURS' top center in a wide serif, date and venue in a small caps line beneath
Edit scopeMake the photo look betterBrighten only the subject's face, keep the background exposure and color as they are
ReferencesUse these imagesImage 1 is the jacket to keep exact, image 2 sets the pose and framing, image 3 sets the palette

Common mistakes

  • Vague text instructions: name the exact words in quotes, or the model invents its own copy.
  • Re-rolling instead of editing: a full regeneration throws away a composition the magic wand could have kept.
  • Unassigned references: five inputs without roles blend into mush; five inputs with roles compose.
  • Bundled edits: "change the label, the background, and the light" in one pass gives you three half-changes. Scope one at a time.
  • Resizing before finishing: recompose the locked keeper, not a draft you are still iterating on.

Simple pricing

Get started for free today, with the option to upgrade or cancel anytime.

Basic

$9/ month
billed as $0 per year

900 monthly credits

1 user only

All models

Workflows

Standard

$24/ month
billed as $0 per year

3200 monthly credits

1 user only

All models

Workflows

Pro

$45/ month
billed as $0 per year

6200 shared monthly credits

1 user

+ up to 4 more at extra cost

All models

Workflows

Pro Max

$170/ month
billed as $0 per year

24000 shared monthly credits

1 user

+ up to 9 more at extra cost

All models

Workflows

Enterprise

For higher limits

Custom

pricing and billing terms

High-volume credits
Custom seat limits
All models
Workflows
Pricing Gradient

Free

For playing around

$0

forever free

Up to 20 credits
1 user only
Limited models
Workflows

FAQs

How do I write a good Grok Imagine Image 2.0 prompt?
Write it like a design brief, not a caption. Name the subject, the layout, the exact on-image words in quotes, the style, and the light, in that order. Image 2.0 follows instructions down to the details, so the more structure you give the brief, the more the result reads as designed rather than improvised.
How does the magic wand work in Grok Imagine Image 2.0?
You point at the region you want to change, describe the change, and the model edits that region while leaving the rest of the image untouched. Segmentation works the same way for precise areas, like one garment or one product, and background removal exports the subject with a transparent background.
How many reference images can I use in Grok Imagine Image 2.0?
Up to 5 input images in a single generation in the Grok apps. Use them as roles rather than a pile: one for the subject, one for the style, one for the layout or scene. Naming what each reference contributes is what turns multi-reference editing into a controlled composite instead of a blend.
What is Smart Resize in Grok Imagine Image 2.0?
Smart Resize recomposes one finished image into any of 9 aspect ratios, from a 1:2 tall banner to a 2:1 wide one. The model fills in the new frame instead of cropping, so the same visual can ship as a story, a square post, a banner, and a widescreen slide without regenerating from scratch.
How good is Grok Imagine Image 2.0 at rendering text?
Text is one of its strongest suits. The model plans typography and layout the way a designer would, so dense, multi-part visuals like infographics and posters hold together and small text comes out sharp. Put the exact words in quotes and describe where they sit, and they render as written.
What are templates in Grok Imagine Image 2.0?
Templates package common image jobs into ready-made starting points, with the workflow already configured. They cover photo editing, product shots, e-commerce photos, professional headshots, icons, mascots, game assets, emoji, and merch, so you supply the inputs and get a finished result without writing the workflow yourself.
How do I use Grok Imagine Image 2.0?
Open grok.com/imagine or the Grok iOS or Android app and pick Quality Mode, which is Image 2.0. Write the brief, generate, then refine with the editing tools: point the magic wand at a region to change it, select areas with segmentation, remove the background, or resize the keeper into other formats.
How is Image 2.0 different from the original Grok Imagine?
The original Grok Imagine is known for bold, stylized output across images and video. Image 2.0 is the precision release: it follows dense briefs closely, plans typography like a designer, and adds region-precise editing, up to 5 reference images, and Smart Resize, aimed at work where the details have to be exactly right.