Browse documentation
All documentation

Media generation

Choose a model for the job

Select models by required capability, creative constraint, access, speed, and estimated cost instead of name alone.

Last reviewed September 25, 2026

Choose a model by eliminating routes that cannot perform the job, then compare the remaining options with a controlled test. Model names, release dates, and popularity are secondary to the capability the work requires.

1. Define the non-negotiables

Write a one-line requirement before opening the picker. Include:

  • output type: image, video, or audio;
  • operation: generate, edit, animate, transform, extend, lip sync, or assemble;
  • required source media and number of references;
  • aspect ratio, resolution, and duration;
  • whether video must include or respond to audio;
  • fidelity requirements for a product, person, voice, label, or existing composition;
  • the maximum acceptable estimate for this attempt.

These constraints create a shortlist. If a model cannot accept the required source video or dialogue audio, its visual quality is irrelevant to that workflow.

2. Read the model picker

The workspace groups models by the selected media type. Search by model or provider, then read the states beside each option.

SignalWhat it meansWhat it does not mean
SelectedThe model will be used for this node with the visible settingsThe model is best for every task
PopularA curated discovery signal in the current catalogGuaranteed quality or lowest cost
NewThe release is recent according to the catalog dateMature behavior for every workflow
Premium or unavailableThe route is visible but not available under the current access stateThe route will silently fall back to another model
Credit estimateExpected usage for the current route and settingsA fixed final charge for every multi-step run

After you select a model, the composer reconciles its controls. If resolution, duration, ratio, voice, or reference controls disappear, that route does not declare those choices for the current mode.

3. Understand the workflow modes

ModeUse it whenRequired thinking
GenerateA prompt should create a new image or soundDescribe the intended result and destination
EditExisting media must remain the foundationState what must change and what must remain
Text to videoNo starting frame is requiredDescribe subject, action, camera, pace, and ending state
Image to videoA still image anchors the opening compositionProtect identity and describe only the intended motion
References to videoSeveral compatible sources define subject, style, motion, or audioGive every reference a clear role and priority
First and last frameThe transition between two controlled states mattersMake both frames compatible and describe the path between them
Extend videoExisting motion should continuePreserve direction, camera logic, lighting, and tempo
Speech or dialogueText should become a performed voice trackSpecify voice, language, pronunciation, pace, and speaker turns
Sound scene, music, or effectThe result is environmental or non-verbal audioDescribe audible events, structure, intensity, and ending

Mode names describe capability, not quality. Two routes can support the same mode while producing different results or exposing different limits.

4. Select for image art direction

For a new image, compare aspect ratios, available resolutions, prompt limits, and whether the route supports later editing. For controlled art, check how many references the edit route accepts.

Match the test to the job:

  • Product fidelity: compare label shape, materials, color, reflections, and small construction details.
  • Layout work: compare adherence to negative space, crop, hierarchy, and camera position.
  • Typography-sensitive work: inspect every generated character; plan to add final campaign copy outside the generation when exact text matters.
  • Style exploration: hold the subject and composition stable while varying only the visual treatment.
  • Iteration: prefer a route with a compatible edit mode when the first output should become the source for controlled revisions.

Do not choose a route only because it exposes the largest resolution. A strong composition at a practical resolution is more useful than a large file with the wrong structure.

5. Select for video production

Video selection begins with the source workflow. Decide whether the shot is text-led, image-led, reference-led, framed by first and last images, an extension, or an edit.

Then compare:

  • allowed duration for the active mode;
  • aspect-ratio and resolution choices;
  • number and type of accepted references;
  • end-frame support;
  • native or driving-audio support;
  • continuity requirements for identity, product detail, camera, and motion.

A route that generates native audio may still be the wrong choice when you already have final dialogue. For a visible speaker driven by a finished voice track, use the dedicated lip-sync workflow. For several approved clips, use Assembly instead of asking one generation to create an entire sequence.

6. Select for audio production

Start with the required mode. Speech models expose voice choices and may differ in languages, delivery controls, timestamps, or output formats. A broader sound model can support dialogue, sound scenes, music, or effects and may accept image or audio references.

For voice comparison, use the same short script and review pronunciation, pacing, emotional fit, consistency, and noise. For music or effects, compare structure, usable ending, unwanted voices, and how well the result fits the intended edit.

7. Match the model to the production stage

Use the lowest-commitment supported route while discovering the concept. Move to the route and settings needed for final fidelity only after the brief, references, and destination format are stable.

  1. Explore: test the core idea and prompt structure.
  2. Compare: run the same controlled brief through a small shortlist.
  3. Refine: use the strongest output as a source when editing is supported.
  4. Finish: select the required resolution, duration, or voice and review at delivery size.

This is not a rule that “fast” models are always for drafts or that expensive models are always better. The picker and the result of a controlled test are the evidence.

8. Run a fair comparison

Keep the prompt, references, aspect ratio, resolution target, and review criteria fixed. Change only the model. Record the model and settings with each result.

Score what the job actually needs: prompt adherence, brand fit, subject fidelity, composition, text, anatomy, motion continuity, audio quality, speed, editability, and estimate. A model that wins a cinematic landscape test may lose a packaging-fidelity test.

9. Confirm access and cost

Availability depends on the current account, environment, and service configuration. Confirm that the route is enabled before building a workflow around it. Review the estimate again after changing resolution, duration, mode, or references.

If no single model supports every constraint, split the work into connected nodes. Generate or edit the source image first, animate it in a compatible video route, create audio separately, then use lip sync or Assembly when appropriate.

Generated from the product catalog

Current model catalog

These tables are generated from Fuzzbucket's current product catalog, so model names and declared capabilities stay aligned with the workspace. Access can vary by account. Open the model picker to confirm availability and see the estimate for your exact settings.

Image models15 models in the current catalogView

Compare generation and editing support, reference capacity, and the output controls declared by each image route.

ModelModesInputsOutput controls
GPT Image 2PopularGenerate, EditPrompt · up to 16 refs6 aspect ratios
Grok Imagine ProPopularGenerate, EditPrompt · up to 3 refs1K, 2K
Nano Banana 2PopularGenerate, EditPrompt · up to 14 refs0.5K, 1K, 2K, 4K
FLUX 2 ProGenerate, EditPrompt · up to 10 refs5 aspect ratios
GPT Image 2.5 SunburstGenerate, EditPrompt · up to 16 refs1K, 2K, 4K
Grok Imagine ImageGenerate, EditPrompt · up to 3 refs1K, 2K
Ideogram V4Generate, EditPrompt · up to 1 reference5 aspect ratios
Krea 2 LargeGeneratePrompt1K
Krea 2 MediumGeneratePrompt1K
Nano Banana ProGenerate, EditPrompt · up to 14 refs1K, 2K, 4K
Qwen Image 2Generate, EditPrompt · up to 3 refs5 aspect ratios
Recraft V4.1GeneratePrompt5 aspect ratios
Reve 2.1Generate, EditPrompt · up to 1 reference18 aspect ratios
Seedream 4.5Generate, EditPrompt · up to 14 refs2K, 4K
Wan 2.7 ImageGenerate, EditPrompt · up to 4 refs5 aspect ratios
Video models18 models in the current catalogView

Compare source workflows, reference capacity, duration, resolution, and whether a route can generate or use audio.

ModelModesInputsOutput controls
Grok Imagine VideoPopularText to video, Image to video, References to video, Edit, Extend videoPrompt · up to 7 refs480p, 720p · 1–15s
Seedance 2.5PopularText to video, Image to video, References to videoPrompt · up to 30 refs480p, 720p · 4–30s · audio-capable
Veo 3.1 FastPopularText to video, Image to video, References to video, Extend video, First and last framePrompt · up to 3 refs720p, 1080p · 4–8s · audio-capable
FLUX 3Text to video, Image to video, First and last frame, Extend videoPrompt720p, 1080p · 5–20s · audio-capable
Gemini Omni FlashText to video, Image to video, References to video, EditPrompt · up to 10 refs2 aspect ratios · 3–10s
Kling Video O3 StandardText to video, Image to video, References to videoPrompt · up to 4 refs3 aspect ratios · 3–15s · audio-capable
Kling Video V3 ProText to video, Image to videoPrompt3 aspect ratios · 3–15s · audio-capable
LTX 2.3Text to video, Image to video, First and last frame, Extend videoPrompt, Image, Video, end frame1080p, 1440p, 2160p · 2–20s · audio-capable
MiniMax H3Text to video, Image to video, References to videoPrompt, Image, Video +3480P, 768P, 2K, 4K · 5–15s · audio-capable
MiniMax H3 MaxText to video, Image to video, References to videoPrompt, Image, Video +3480P, 768P · 5–15s · audio-capable
MiniMax H3 Max TurboText to video, Image to videoPrompt, Image, end frame480P, 768P · 5–15s · audio-capable
PixVerse V6Text to video, Image to video, First and last frame, Extend videoPrompt, Image, Video, end frame360p, 540p, 720p, 1080p · 1–15s · audio-capable
Seedance 2.0Text to video, Image to video, References to videoPrompt · up to 9 refs480p, 720p, 1080p, 4k · 4–15s · audio-capable
Seedance 2.0 FastText to video, Image to video, References to videoPrompt · up to 9 refs480p, 720p · 4–15s · audio-capable
Sora 2Text to video, Image to video, EditPrompt4–20s
Veo 3.1Text to video, Image to video, References to video, Extend video, First and last framePrompt · up to 3 refs720p, 1080p, 4k · 4–8s · audio-capable
Vidu Q3Text to video, Image to video, First and last frame, References to videoPrompt, Image, up to 4 refs, end frame360p, 540p, 720p, 1080p · 1–16s · audio-capable
Wan 2.7Text to video, Image to video, First and last frame, Extend video, References to video, EditPrompt, Image, Video +3720p, 1080p · 2–15s · audio-capable
Audio models4 models in the current catalogView

Compare speech and sound-generation modes, accepted source types, formats, duration, and voice options.

ModelModesInputsOutput controls
ElevenLabs Turbo v2.5SpeechPromptMP3 · 8 voices
ElevenLabs v3SpeechPromptMP3 · 8 voices
Seed Audio 1.0Speech, Dialogue, Sound scene, Music, Sound effectPrompt, Image, Audio, up to 3 refsup to 120s · WAV, MP3, PCM, OGG_OPUS
Seed Speech TTS 2.0SpeechPromptMP3, OPUS · 10 voices