AI Model

Grok Comfy Workflows

Run xAI's Grok image and video generation in ComfyUI: photorealistic generation, image editing, and image-to-video, on Comfy Cloud.

8 workflows

About This AI Model

What is the Grok model?

The starting point for Grok comfyui workflows is that Grok is a cloud model from xAI, so these workflows reach it through ComfyUI's API nodes rather than loading weights. There is no local download, which means you trade self-hosting for access to xAI's image and video stack without managing a GPU. The workflows on this page wire those calls into editable graphs, so the prompt, the reference images, and the output flow through nodes you can adjust and chain.

The image side is built on Aurora, xAI's autoregressive generator, which leans toward photorealistic results and renders text and logos in-frame more accurately than many models. You can generate from a prompt, request several variations in one run, or edit an existing photo by uploading it with a text instruction. The Grok Image Edit node accepts up to three source images, and xAI quotes outputs in roughly four seconds per image, so iterating on framing stays fast. On the video side, Grok Imagine turns a prompt into a clip and can animate a still image you provide, with native audio generated alongside the picture and clip length set in the workflow.

Because these are open ComfyUI graphs around a closed model, you keep control of the inputs that shape the result: the prompt wording, the reference images for editing, the number of outputs, and the aspect ratio. Partner nodes also add reference-to-video, which takes up to seven reference images for a consistent clip, and a video-extend node that lengthens an existing clip while keeping motion and sound continuous. The model itself is fixed and hosted by xAI, but everything around it in the graph is yours to tune, and the output drops straight into the rest of your ComfyUI pipeline.

To run one, open a Grok workflow on Comfy Cloud with nothing to install, since there is no local option. Write your prompt or upload the photo you want to edit, set the options the workflow exposes, and generate. Re-run with a revised prompt or different references until the framing and detail are right. xAI notes that Grok Imagine is still a beta under active development, so it pays to layer your own controls and combine its output with other nodes rather than relying on a single pass.

HY 3D: Image to Model
Hunyuan3D

Hunyuan3D

Cloud via API

Cloud model from xAI; no public weights, run through ComfyUI API nodes.

Realistic 2k Images - Quick Variations
Grok

Grok

Aurora model

Aurora autoregressive image model tuned for photorealistic results.

Reproducible, tunable, yours

Every step is exposed and adjustable, and the exact graph is yours to reuse across projects.

Start free. Upgrade when you’re ready.

See pricing plans

Capabilities

What You Can Create

Grok is xAI's cloud image and video stack, accessed through ComfyUI API nodes. The Aurora image model favors photorealism and accurate in-frame text, while Grok Imagine handles video generation.

Text and logos

Accurate rendering of text and logos inside generated images.

Try workflow
Ovis-Image Text to Image

Photo editing

Image editing from a photo plus a text instruction, up to three source images.

Try workflow
Seedream 4.0: Image Edit
ByteDance

Video with audio

Grok Imagine generates video with native audio and animates a still image.

Try workflow
Kling2.6: Animate Images with Audio
Kling 2.6
HOWComfyWORKS

Connect models, processing steps, and outputs on a canvas where every decision is visible and every step is inspectable.

Start from a community template or build from scratch.

Comparison

Why Comfy Workflows

  • Image and video in one graph

    Aurora generation, photo editing, and Grok Imagine video all run from editable ComfyUI nodes, so the output of one step feeds the next.

  • Pick your variant and settings

    ONLY ONComfyCLOUD

    Choose the Grok task, set the source images, output count, aspect ratio, and clip length, then re-run until the framing and detail are right.

  • Cloud access without a GPU

    Grok has no public weights, so the workflows call xAI through ComfyUI API nodes on Comfy Cloud with no install and nothing to download.

  • Accurate text and native audio

    Aurora renders legible text and logos in-frame, and Grok Imagine generates sound together with the video in a single pass.

Applications

What people use it for

What people generate and edit through Grok's cloud image and video models.

Nano Banana 2 Lite: Text to Image

Text to image

Nano Banana 2

Generate photorealistic images from a prompt with a grok text to image workflow

Hunyuan Video 1.5 Image to Video

Image to video

Hunyuan Video

Animate a still into a short clip with a grok image to video workflow

Frequently Asked Questions

Grok FAQ

Ready to create?

Start generating with this workflow in seconds

Try Comfy Cloud