Browse documentation
Media generation
Choose a model for the job
Select models by required capability, creative constraint, access, speed, and estimated cost instead of name alone.
Last reviewed September 25, 2026
Choose a model by eliminating routes that cannot perform the job, then compare the remaining options with a controlled test. Model names, release dates, and popularity are secondary to the capability the work requires.
1. Define the non-negotiables
Write a one-line requirement before opening the picker. Include:
- output type: image, video, or audio;
- operation: generate, edit, animate, transform, extend, lip sync, or assemble;
- required source media and number of references;
- aspect ratio, resolution, and duration;
- whether video must include or respond to audio;
- fidelity requirements for a product, person, voice, label, or existing composition;
- the maximum acceptable estimate for this attempt.
These constraints create a shortlist. If a model cannot accept the required source video or dialogue audio, its visual quality is irrelevant to that workflow.
2. Read the model picker
The workspace groups models by the selected media type. Search by model or provider, then read the states beside each option.
| Signal | What it means | What it does not mean |
|---|---|---|
| Selected | The model will be used for this node with the visible settings | The model is best for every task |
| Popular | A curated discovery signal in the current catalog | Guaranteed quality or lowest cost |
| New | The release is recent according to the catalog date | Mature behavior for every workflow |
| Premium or unavailable | The route is visible but not available under the current access state | The route will silently fall back to another model |
| Credit estimate | Expected usage for the current route and settings | A fixed final charge for every multi-step run |
After you select a model, the composer reconciles its controls. If resolution, duration, ratio, voice, or reference controls disappear, that route does not declare those choices for the current mode.
3. Understand the workflow modes
| Mode | Use it when | Required thinking |
|---|---|---|
| Generate | A prompt should create a new image or sound | Describe the intended result and destination |
| Edit | Existing media must remain the foundation | State what must change and what must remain |
| Text to video | No starting frame is required | Describe subject, action, camera, pace, and ending state |
| Image to video | A still image anchors the opening composition | Protect identity and describe only the intended motion |
| References to video | Several compatible sources define subject, style, motion, or audio | Give every reference a clear role and priority |
| First and last frame | The transition between two controlled states matters | Make both frames compatible and describe the path between them |
| Extend video | Existing motion should continue | Preserve direction, camera logic, lighting, and tempo |
| Speech or dialogue | Text should become a performed voice track | Specify voice, language, pronunciation, pace, and speaker turns |
| Sound scene, music, or effect | The result is environmental or non-verbal audio | Describe audible events, structure, intensity, and ending |
Mode names describe capability, not quality. Two routes can support the same mode while producing different results or exposing different limits.
4. Select for image art direction
For a new image, compare aspect ratios, available resolutions, prompt limits, and whether the route supports later editing. For controlled art, check how many references the edit route accepts.
Match the test to the job:
- Product fidelity: compare label shape, materials, color, reflections, and small construction details.
- Layout work: compare adherence to negative space, crop, hierarchy, and camera position.
- Typography-sensitive work: inspect every generated character; plan to add final campaign copy outside the generation when exact text matters.
- Style exploration: hold the subject and composition stable while varying only the visual treatment.
- Iteration: prefer a route with a compatible edit mode when the first output should become the source for controlled revisions.
Do not choose a route only because it exposes the largest resolution. A strong composition at a practical resolution is more useful than a large file with the wrong structure.
5. Select for video production
Video selection begins with the source workflow. Decide whether the shot is text-led, image-led, reference-led, framed by first and last images, an extension, or an edit.
Then compare:
- allowed duration for the active mode;
- aspect-ratio and resolution choices;
- number and type of accepted references;
- end-frame support;
- native or driving-audio support;
- continuity requirements for identity, product detail, camera, and motion.
A route that generates native audio may still be the wrong choice when you already have final dialogue. For a visible speaker driven by a finished voice track, use the dedicated lip-sync workflow. For several approved clips, use Assembly instead of asking one generation to create an entire sequence.
6. Select for audio production
Start with the required mode. Speech models expose voice choices and may differ in languages, delivery controls, timestamps, or output formats. A broader sound model can support dialogue, sound scenes, music, or effects and may accept image or audio references.
For voice comparison, use the same short script and review pronunciation, pacing, emotional fit, consistency, and noise. For music or effects, compare structure, usable ending, unwanted voices, and how well the result fits the intended edit.
7. Match the model to the production stage
Use the lowest-commitment supported route while discovering the concept. Move to the route and settings needed for final fidelity only after the brief, references, and destination format are stable.
- Explore: test the core idea and prompt structure.
- Compare: run the same controlled brief through a small shortlist.
- Refine: use the strongest output as a source when editing is supported.
- Finish: select the required resolution, duration, or voice and review at delivery size.
This is not a rule that “fast” models are always for drafts or that expensive models are always better. The picker and the result of a controlled test are the evidence.
8. Run a fair comparison
Keep the prompt, references, aspect ratio, resolution target, and review criteria fixed. Change only the model. Record the model and settings with each result.
Score what the job actually needs: prompt adherence, brand fit, subject fidelity, composition, text, anatomy, motion continuity, audio quality, speed, editability, and estimate. A model that wins a cinematic landscape test may lose a packaging-fidelity test.
9. Confirm access and cost
Availability depends on the current account, environment, and service configuration. Confirm that the route is enabled before building a workflow around it. Review the estimate again after changing resolution, duration, mode, or references.
If no single model supports every constraint, split the work into connected nodes. Generate or edit the source image first, animate it in a compatible video route, create audio separately, then use lip sync or Assembly when appropriate.
Generated from the product catalog
Current model catalog
These tables are generated from Fuzzbucket's current product catalog, so model names and declared capabilities stay aligned with the workspace. Access can vary by account. Open the model picker to confirm availability and see the estimate for your exact settings.
Image models15 models in the current catalogViewHide
Compare generation and editing support, reference capacity, and the output controls declared by each image route.
| Model | Modes | Inputs | Output controls |
|---|---|---|---|
| GPT Image 2Popular | Generate, Edit | Prompt · up to 16 refs | 6 aspect ratios |
| Grok Imagine ProPopular | Generate, Edit | Prompt · up to 3 refs | 1K, 2K |
| Nano Banana 2Popular | Generate, Edit | Prompt · up to 14 refs | 0.5K, 1K, 2K, 4K |
| FLUX 2 Pro | Generate, Edit | Prompt · up to 10 refs | 5 aspect ratios |
| GPT Image 2.5 Sunburst | Generate, Edit | Prompt · up to 16 refs | 1K, 2K, 4K |
| Grok Imagine Image | Generate, Edit | Prompt · up to 3 refs | 1K, 2K |
| Ideogram V4 | Generate, Edit | Prompt · up to 1 reference | 5 aspect ratios |
| Krea 2 Large | Generate | Prompt | 1K |
| Krea 2 Medium | Generate | Prompt | 1K |
| Nano Banana Pro | Generate, Edit | Prompt · up to 14 refs | 1K, 2K, 4K |
| Qwen Image 2 | Generate, Edit | Prompt · up to 3 refs | 5 aspect ratios |
| Recraft V4.1 | Generate | Prompt | 5 aspect ratios |
| Reve 2.1 | Generate, Edit | Prompt · up to 1 reference | 18 aspect ratios |
| Seedream 4.5 | Generate, Edit | Prompt · up to 14 refs | 2K, 4K |
| Wan 2.7 Image | Generate, Edit | Prompt · up to 4 refs | 5 aspect ratios |
Video models18 models in the current catalogViewHide
Compare source workflows, reference capacity, duration, resolution, and whether a route can generate or use audio.
| Model | Modes | Inputs | Output controls |
|---|---|---|---|
| Grok Imagine VideoPopular | Text to video, Image to video, References to video, Edit, Extend video | Prompt · up to 7 refs | 480p, 720p · 1–15s |
| Seedance 2.5Popular | Text to video, Image to video, References to video | Prompt · up to 30 refs | 480p, 720p · 4–30s · audio-capable |
| Veo 3.1 FastPopular | Text to video, Image to video, References to video, Extend video, First and last frame | Prompt · up to 3 refs | 720p, 1080p · 4–8s · audio-capable |
| FLUX 3 | Text to video, Image to video, First and last frame, Extend video | Prompt | 720p, 1080p · 5–20s · audio-capable |
| Gemini Omni Flash | Text to video, Image to video, References to video, Edit | Prompt · up to 10 refs | 2 aspect ratios · 3–10s |
| Kling Video O3 Standard | Text to video, Image to video, References to video | Prompt · up to 4 refs | 3 aspect ratios · 3–15s · audio-capable |
| Kling Video V3 Pro | Text to video, Image to video | Prompt | 3 aspect ratios · 3–15s · audio-capable |
| LTX 2.3 | Text to video, Image to video, First and last frame, Extend video | Prompt, Image, Video, end frame | 1080p, 1440p, 2160p · 2–20s · audio-capable |
| MiniMax H3 | Text to video, Image to video, References to video | Prompt, Image, Video +3 | 480P, 768P, 2K, 4K · 5–15s · audio-capable |
| MiniMax H3 Max | Text to video, Image to video, References to video | Prompt, Image, Video +3 | 480P, 768P · 5–15s · audio-capable |
| MiniMax H3 Max Turbo | Text to video, Image to video | Prompt, Image, end frame | 480P, 768P · 5–15s · audio-capable |
| PixVerse V6 | Text to video, Image to video, First and last frame, Extend video | Prompt, Image, Video, end frame | 360p, 540p, 720p, 1080p · 1–15s · audio-capable |
| Seedance 2.0 | Text to video, Image to video, References to video | Prompt · up to 9 refs | 480p, 720p, 1080p, 4k · 4–15s · audio-capable |
| Seedance 2.0 Fast | Text to video, Image to video, References to video | Prompt · up to 9 refs | 480p, 720p · 4–15s · audio-capable |
| Sora 2 | Text to video, Image to video, Edit | Prompt | 4–20s |
| Veo 3.1 | Text to video, Image to video, References to video, Extend video, First and last frame | Prompt · up to 3 refs | 720p, 1080p, 4k · 4–8s · audio-capable |
| Vidu Q3 | Text to video, Image to video, First and last frame, References to video | Prompt, Image, up to 4 refs, end frame | 360p, 540p, 720p, 1080p · 1–16s · audio-capable |
| Wan 2.7 | Text to video, Image to video, First and last frame, Extend video, References to video, Edit | Prompt, Image, Video +3 | 720p, 1080p · 2–15s · audio-capable |
Audio models4 models in the current catalogViewHide
Compare speech and sound-generation modes, accepted source types, formats, duration, and voice options.
| Model | Modes | Inputs | Output controls |
|---|---|---|---|
| ElevenLabs Turbo v2.5 | Speech | Prompt | MP3 · 8 voices |
| ElevenLabs v3 | Speech | Prompt | MP3 · 8 voices |
| Seed Audio 1.0 | Speech, Dialogue, Sound scene, Music, Sound effect | Prompt, Image, Audio, up to 3 refs | up to 120s · WAV, MP3, PCM, OGG_OPUS |
| Seed Speech TTS 2.0 | Speech | Prompt | MP3, OPUS · 10 voices |