Venice keys Venice Learn · Guides Open Venice ↗
Guide · Video

How to generate private AI video with Grok Imagine on Venice

This guide shows three ways to create video with the private Grok Imagine models on Venice: agentic chat, classic chat, and Venice Studio. You will learn how each workflow handles images, credits, and aspect ratios, and how the advanced reference-to-video feature lets you build scenes from up to seven tagged images.

Watch the full walkthrough
What you'll learn
  • Why the private Grok Imagine models keep your prompts and generations from being stored or trained on
  • How to generate video in three different Venice interfaces and when to use each
  • How to prep an image for 16:9 widescreen video using a private image-edit model
  • The difference between private models and anonymized models on Venice
  • How reference-to-video lets you combine and tag multiple images in a single prompt

Why private video generation matters

On Venice, the Grok Imagine video models run privately. Your prompts and generations are not stored anywhere and are not fed back into any AI training set. That difference is the whole point of this guide: it lets you work with images you would never want to hand to a public AI service.

The clearest example is personal photos. If you value your privacy and would rather not upload pictures of yourself or your family to the open internet, the private models let you experiment with those images and know they stay yours. The guide uses a personal headshot throughout to demonstrate exactly this.

Method 1: Agentic chat in one prompt

The fastest route is the agentic chat. You describe the whole result in plain language and the agent handles the intermediate steps. In the video the prompt is:

I want a video created with Grok Imagine of me at the beach with my dogs. Create this image to start out that we can then turn into a 16:9 5 second video.

The agent first generates the base image, then prepares a prompt and offers to feed that image into the Grok Imagine image-to-video model. You confirm with "yes, generate video" and it produces the clip without further clicks.

Video generation costs credits, and Venice shows you the cost before committing. In the example the video costs 44 credits. A confirmation warning appears so you can review the base image before spending; you can toggle that warning off to skip straight to generation, but leaving it on lets you catch anything you want fixed first.

Method 2: Classic chat for more control

The classic chat does the same job in manual steps, which gives you finer control. Click the plus button, choose edit image, and upload your photo. Pick a private model so the workflow stays private, then prompt the edit, for example "remove my shirt and put me at the beach with two dogs." Editing an image costs about 4 credits, and you can regenerate or switch to a higher-quality variant if the result is off.

Venice can automatically enhance your prompt. That toggle lives in the advanced settings and improves results by default, though you can turn it off when you want your exact wording respected.

Before turning the image into video, set the aspect ratio. Not every edit model can change it: Grok Imagine keeps the aspect ratio automatic, so the video switches to a private model that can reframe, sets 16:9, toggles prompt enhancement off, and prompts "extend the scene to 16:9" to get a widescreen base frame. Clicking create video loads that image as the first frame. Note the chat may default to an anonymized model; choose Grok Imagine private to keep the run fully private. Set the clip length and resolution (5 seconds, 720p or 480p in the demo) and generate.

Private models versus anonymized models

Venice offers two kinds of protection, and the distinction matters when you pick a model. Private models, like Grok Imagine private, never store your prompts or generations at all.

Anonymized models are different. The default video model in the example is uncensored but anonymous, meaning your request is routed through Venice to the upstream provider with your identity stripped, rather than kept entirely private. Both protect you, but if your goal is full privacy, confirm you are on a private model rather than an anonymized one before generating.

Method 3: Venice Studio and running multiple models

Venice Studio, linked in the sidebar, is the third workflow and the most flexible. It has tabs for generating images, editing images, generating audio, voice, and sound effects, generating video, and a movie editor for assembling clips.

In the video tab you can select several models at once and send the same image and prompt to all of them, instead of re-prompting each one separately as you would in chat. A filter controls which models appear, so you can narrow the list to image-to-video models or private models only. The demo runs Grok Imagine alongside two other models for comparison.

You do not have to wait for one generation to finish before starting another. Queue as many as you like and review them together, which makes it easy to give every model the same starting image, prompt, and references and then pick the best result.

Reference-to-video: building scenes from multiple images

Reference-to-video is the advanced feature. Instead of a single starting frame, you add multiple reference images and tag them in your prompt to control what appears and when. Among the private image-to-video options, Grok Imagine R2V private is the one that supports this; other private image-to-video models accept only a single uploaded image.

You can add images from your device or pull from assets already saved in Venice, up to seven references. Tag each one in the prompt and direct the action. The demo builds a cyberpunk chase by combining a personal photo, a cyberpunk scene, and a generated alien:

Put @image one into the cyberpunk world of image two. He is being chased by image three.

Set length, resolution, and aspect ratio in the settings (8 seconds, 720p, 16:9 in the example) and generate. Because it is video, not every subject needs to be in frame at the same time, so reference-to-video lets you stage detail that a single image edit cannot. Finished clips can be dragged into the Studio movie editor, trimmed, and scored with music for a complete piece of content.

Key takeaways

Adapted from the @askvenice video on YouTube. Models and prices change fast; verify current details in Venice before production use.