Other tools hide which model rendered your clip. Hybrig shows you the whole stack — which model runs where, what it's good at, what it's mediocre at, and what it costs. Pick the right tool for the job, or trust the recipe library to pick for you.
16 models8 pipeline tools12 recipes
Section 01 · Compare
Compare these models
Pick up to four models. Honest scores, honest costs. The 0-100 marks are our internal calibration — they won't flatter any vendor.
Pick up to 4 models to compare
Tell me what I'm making
Seedance 2.0 Standard
Wan 2.2 (local)
Gemini Omni Flash (cloud)
Cost / 5s clip
$1.51
free
$0.50
Wall clock / 5s
~2m
~3m
~3m
Face quality (0-100)
92
78
80
Motion quality (0-100)
85
72
90
Identity preservation
90
81
68
Cinematic look
82
65
94
Hardware
cloud
16GB VRAM
cloud
Free?
no
yes
no
Lip-sync
yes
no
yes
Recipes
Premium Brand Avatar, Multi-Take A/B, Brand Ad (Hermes-style)
UGC Spokesperson (Free Local), UGC Spokesperson (Polished), Quick Social Clip, Animated Explainer
Cinematic Cold Open, Impossible-shot hero beats, Any scene we were previously hand-rendering in Google Flow
Seedance 2.0 Standard
Cost / 5s clip
$1.51
Wall clock / 5s
~2m
Face quality (0-100)
92
Motion quality (0-100)
85
Identity preservation
90
Cinematic look
82
Hardware
cloud
Free?
no
Lip-sync
yes
Recipes
Premium Brand Avatar, Multi-Take A/B, Brand Ad (Hermes-style)
Wan 2.2 (local)
Cost / 5s clip
free
Wall clock / 5s
~3m
Face quality (0-100)
78
Motion quality (0-100)
72
Identity preservation
81
Cinematic look
65
Hardware
16GB VRAM
Free?
yes
Lip-sync
no
Recipes
UGC Spokesperson (Free Local), UGC Spokesperson (Polished), Quick Social Clip, Animated Explainer
Gemini Omni Flash (cloud)
Cost / 5s clip
$0.50
Wall clock / 5s
~3m
Face quality (0-100)
80
Motion quality (0-100)
90
Identity preservation
68
Cinematic look
94
Hardware
cloud
Free?
no
Lip-sync
yes
Recipes
Cinematic Cold Open, Impossible-shot hero beats, Any scene we were previously hand-rendering in Google Flow
Section 02 · Model Atlas
Every video model we support
Each card is a real review. We're especially direct about weaknesses — knowing what a model is bad at is more useful than knowing what it's good at.
Showing 16 of 16
BD
Seedance 2.5
pro · cloud
$0.47/sec
sample seedance-2.5-standard
Newest ByteDance model. Native audio, up to 30s in ONE generation. 720p ceiling — use 2.0 when you need 4K.
Strengths
+Up to 30s in a single generation — multi-shot sequences without stitching or per-cut identity drift
+Generates synchronised audio in the same latent space as the picture
+Accepts image, video AND audio references together
+Seven aspect ratios including 21:9 and 4:3
Weaknesses
−720p ceiling on fal as of 2026-08-09 — 1080p and 4k were removed from the endpoint; use Seedance 2.0 for 4K
−~1.5x the 2.0 token rate ($0.0214 vs $0.014 per 1k) — a 30s 720p clip is ~$14
−No fast tier, so there is no cheap draft option on 2.5
−Unproven here; we have not run a side-by-side against 2.0 yet
5s clip
$2.36
~Wall clock
3m
Hardware
cloud
Best for
Premium Brand AvatarBrand Ad (Hermes-style)
BD
Seedance 2.0 Standard
standard · cloud
$0.30/sec
sample seedance-2.0-standard
Top-tier quality. Best face lock + VFX. Use for finals.
Strengths
+Best face lock under heavy motion in our tests
+Strong cinematic VFX and lighting consistency
+Reliable identity preservation across takes
Weaknesses
−Mediocre on hands when subject gestures
−Backgrounds can drift on clips longer than ~6s
5s clip
$1.51
~Wall clock
2m
Hardware
cloud
Best for
Premium Brand AvatarMulti-Take A/BBrand Ad (Hermes-style)
??
MiniMax H3 (2K)
standard · cloud
$0.26/sec
sample minimax-h3
Native 2K with its own audio, up to 15s. Cheapest way to finish a shot at this size.
Strengths
+Native 2K — not an upscale of 1080p
+Generates its own stereo audio alongside the picture
+Up to 15s in one take, where Seedance stops at 12
+Cheaper per second than Seedance Standard despite the higher resolution
Weaknesses
−No seed parameter — runs cannot be pinned, so one-variable A/B testing is not available
−Reference images past the first five cost $0.08 each
−New here: we have no measured face-lock or drift numbers of our own yet
5s clip
$1.30
~Wall clock
2m
Hardware
cloud
Best for
Multi-Take A/BBrand Ad (Hermes-style)Quick Social Clip
??
MiniMax H3 (768p draft)
fast · cloud
$0.16/sec
sample minimax-h3-draft
Same model at 768p for a third less. Use it to get the shot right, then finish at 2K.
Strengths
+Cheapest cloud render available here — good for burning through takes
+Identical model and prompt behaviour as the 2K entry, so what you fix carries over
+~20% cheaper than Standard with most quality intact
+Faster turnaround — good for A/B drafting
+Same face-lock backbone as Standard
Weaknesses
−Visible quality dip on complex multi-subject scenes
−Texture detail softer than Standard at 1080p
5s clip
$1.21
~Wall clock
2m
Hardware
cloud
Best for
UGC Spokesperson (Polished)Multi-Take A/BQuick Social Clip
AB
Wan 2.2 (local)
standard · local
LOCAL FREE
sample wan-2.2-local
Alibaba's Wan 2.2 — the new default local model. Slightly better motion + identity than 2.1, same install path. Runs on your own GPU via ComfyUI + WanVideoWrapper. ~3-8 min per 5s clip on a 4090. Free forever. Requires ~14 GB on disk and 16+ GB GPU memory.
Strengths
+Free forever — runs on your own GPU, no per-second cost
+No cloud queue — starts in seconds, not hours
+Privacy: footage never leaves your machine
+Better motion + identity than Wan 2.1 in our tests
Weaknesses
−Still behind Seedance on demanding face-heavy shots
−Hands and crowd scenes weaker than top cloud models
−Requires ComfyUI install and 16+ GB GPU memory
5s clip
free
~Wall clock
3m
Hardware
16GB VRAM
Best for
UGC Spokesperson (Free Local)UGC Spokesperson (Polished)Quick Social ClipAnimated Explainer
AB
Wan VACE face swap (local)
standard · local· coming soon
LOCAL FREE
sample wan-vace-face-swap-local
Video-to-video face swap — feed an existing video clip + a reference photo of who you want to be the new face, get the same video back with the face replaced. The technique behind viral 'insert myself into Game of Thrones' or 'fake news interview' content. Locks identity via PuLID + VACE conditioning on top of Wan 2.2. Runs on your own GPU via ComfyUI + kijai/WanVideoWrapper. ~6-12 min per 5s clip on a 4090. Free forever.
Strengths
+Free forever — runs on your own GPU, no per-second cost
+Inserts you into existing footage at near-cloud quality without renting Runway/Pika
+Locks identity across all frames of the clip via PuLID face conditioning
+Source video can be anything — public-domain clips, your own footage, royalty-free B-roll
Weaknesses
−Requires Wan 2.2 base model (Q5_K_M quantized version — compressed so it fits a 4090)
−Hands and crowd geometry still slightly behind Seedance / Veo on tight close-ups
−Source video length matters — best results at 3-8 seconds; longer clips drift on subject identity
−Requires a clean reference photo of the target face (high-res, neutral expression, well-lit)
5s clip
free
~Wall clock
8m
Hardware
16GB VRAM
Best for
Insert-yourself-into-famous-scene viral content (Game of Thrones, Star Wars, etc.)Fake-news / fake-interview ad genre (the Miami pistachio chase technique)Replacing a stunt double's face with the actor in postBrand-personalized cameo videos (contractor's face on a spokesperson clip)
AB
Wan 2.1 14B FLF2V (local)
standard · local
LOCAL FREE
sample wan-2.2-flf2v-local
First/last frame interpolation — give it two images, get the video that morphs between them. The ONLY local model with native start+end keyframe conditioning (standard Wan I2V, Hunyuan, LTX do not have it). Perfect for weathering time-lapses, before/after roof shots, age progression, and any morph where both ends of the clip are locked. Runs on your own GPU via ComfyUI + kijai/WanVideoWrapper. ~5-8 min per 5s clip on a 4090. Free forever.
Strengths
+Free forever — runs on your own GPU, no per-second cost
+Only LOCAL option for first/last frame interpolation — Wan 2.1, Hunyuan, LTX do not have this capability
+Locks both ends of the clip — model can't drift away from your end keyframe
+Privacy: both source images and the output never leave your machine
Weaknesses
−Requires Wan 2.2 base model — full fp16 doesn't fit most consumer GPUs; we ship the Q5_K_M quantized build (compressed version of the model that runs on smaller GPUs at slightly lower quality)
−For single-frame I2V (one starting image only), Wan 2.1 wins on quality — swap back when you don't need the end-frame lock
−Both source images MUST match framing/composition exactly — same camera, same crop, same subject position — or the morph looks broken
−Caps at 5s output; longer morphs need to be chained
5s clip
free
~Wall clock
6m
Hardware
16GB VRAM
Best for
Weathering time-lapse (fresh shingle → 25-year-aged shingle)Before/after restoration shotsAge progression on identical-framing portraitsAny morph where both endpoints are locked
AB
Wan 2.1 (local)
standard · local
LOCAL FREE
sample wan-2.1-local
Older Wan generation — kept around as a backup for when 2.2 misbehaves on your install. Same look, slightly weaker motion + identity than 2.2. Runs on your own GPU via ComfyUI + WanVideoWrapper. ~3-8 min per 5s clip on a 4090. Free forever.
Strengths
+Free forever — runs on your own GPU, no per-second cost
+No cloud queue — starts in seconds, not hours
+Privacy: footage never leaves your machine
+Mature: well-tested workflow, lots of community fixes
Weaknesses
−~70-80% quality vs Seedance on demanding shots
−Hands and complex motion are noticeably weaker
−Requires ComfyUI install and 16+ GB GPU memory
5s clip
free
~Wall clock
3m
Hardware
16GB VRAM
Best for
UGC Spokesperson (Free Local)UGC Spokesperson (Polished)Quick Social ClipAnimated Explainer
TC
HunyuanVideo 1.5 (local)
standard · local
LOCAL FREE
sample hunyuan-video-local
Tencent's HunyuanVideo 1.5 — strong on natural motion, especially body movement and physics. Handles both text-to-video and image-to-video. Runs locally via ComfyUI's Hunyuan nodes. Use the smaller, slightly less accurate version (fp8) — about 13 GB on disk. Needs 16+ GB GPU memory. Slower than Wan: plan on 6-12 min per 5s clip on a 4090.
Strengths
+Best-in-class motion realism among free local models
+Body physics and limb movement noticeably better than Wan
+Strong on action, walking, gestures
+Different lineage from Wan — survives a broken WanVideoWrapper install
Weaknesses
−Slower than Wan — heavier model
−Identity lock weaker than Wan on tight close-ups
−Larger download (~13 GB)
5s clip
free
~Wall clock
5m
Hardware
16GB VRAM
Best for
Action DemoUGC Spokesperson (Polished)Quick Social Clip
LT
LTX-2.3 (local)
standard · local
LOCAL FREE
sample ltx-video-local
Lightricks LTX-2.3 — the only local model here that makes picture AND sound in one pass. Everything else in this catalog is silent; this one generates its own audio, synced, in the same render. Open weights, huge quantization spread (52 builds on Hugging Face) so it scales from a 3080 up to a 4090. Runs via ComfyUI. Plan on ~4-9 min per 5s clip at fp8 on a 4090.
Strengths
+Free forever — runs on your own GPU, no per-second cost
+Only local model in this catalog with native synced audio + video in a single pass
+Widest VRAM range in the catalog — fp8 runs at 12 GB, GGUF builds go lower
+Strong fine-tuning ecosystem: IC-LoRA adapters, first/last-frame and camera-control LoRAs
Weaknesses
−Native audio is impressive but not broadcast-clean — expect to replace VO on anything scripted
−Face lock still behind Seedance on tight close-ups
−The quantization spread is a trap — grab the wrong build and quality drops hard with no warning
5s clip
free
~Wall clock
6m
Hardware
12GB VRAM
Best for
Quick Social ClipAnimated ExplainerAmbient / atmosphere shots that need their own sound
KS
Kling 3.0 (cloud)
pro · cloud· coming soon
$0.11/sec
sample kling-3.0-cloud
Kuaishou's cinematic model, now with native audio. Different aesthetic than Seedance and a third of the price — often wins on establishing and camera-move shots.
Strengths
+Distinctive cinematic look — strong on establishing shots
−Pricing and slug not yet verified — treat as experimental
5s clip
$2.00
~Wall clock
2m
Hardware
cloud
Best for
Animated ExplainerCustom Workflow Builder
GG
Google VEO 3.1 (cloud)
pro · cloud· coming soon
$0.40/sec
sample veo-3.1-cloud
Google's VEO 3.1 — text-to-video with native audio. Google now points new work at Omni Flash instead; keep VEO for scene extension, last-frame control, and existing pipelines built against it.
Strengths
+Best-in-class text-to-video coherence
+Cinematic look with very good world physics
+Native synced audio on supported endpoints
Weaknesses
−Reference-image fidelity weaker than Seedance
−Identity-lock not as tight on returning characters
5s clip
$2.00
~Wall clock
3m
Hardware
cloud
Best for
Cinematic Cold OpenAnimated Explainer
GG
Gemini Omni Flash (cloud)
pro · cloud· coming soon
$0.10/sec
sample gemini-omni-flash-cloud
Google's current flagship video model and their recommended default. Cheapest premium cloud option in this catalog at $0.10/sec — about a third of Seedance Standard. Caps at 10 second generations while in preview. Needs a Gemini API key, not a fal key.
Strengths
+Cheapest premium cloud model in the catalog — $0.10/sec
+Google's own recommended default over VEO 3.1
+Native synced audio
+Built for conversational video editing, not just one-shot generation
Weaknesses
−Public preview — quotas, pricing, and region availability can change without notice
−Caps at 10-second generations for now
−Google-direct only: does not ride our existing fal.ai plumbing
−Preview model IDs get retired; expect to re-pin the slug
5s clip
$0.50
~Wall clock
3m
Hardware
cloud
Best for
Cinematic Cold OpenImpossible-shot hero beatsAny scene we were previously hand-rendering in Google Flow
Section 03 · Tool Atlas
Everything else in the pipeline
A finished render isn't just a model call. Voice cloning, identity verification, polish, captioning, brand frames — all the tools that participate, what they cost, where they run.
Pre-render
pre-render · voice
ElevenLabs
ElevenLabs Inc.
live
Voice cloning, TTS in 30+ languages, and lip-sync polish. We use it as the spokesperson's voice and as the 'translator' for multilingual dubs.
Runs on
Cloud API
Cost
~$0.30 per 1k characters (Creator tier). Cloning included in monthly plan.
pre-render · identity
Wan ClipVision
live
Face encoder that locks the spokesperson's identity into the diffusion model. Runs as part of the local Wan workflow — no cloud round trip.
Runs on
On your GPU
Cost
Free. Bundled with WanVideoWrapper.
Post-render
post-render · polish
Topaz Astra
Topaz Labs
planned
AI upscale + temporal artifact cleanup. Takes a 720p Wan render to a clean 1080p/4K. Fixes flicker on long clips and softens compression artifacts.
Runs on
Desktop app
Cost
$299 one-time desktop license. No per-render fee.
post-render · polish
Hedra Character-3
Hedra Labs
experimental
Audio-driven micro-expressions on the talking head — eyebrow raises, head tilts, breath. Layered after the base render to add humanity.
Runs on
Cloud API
Cost
Roughly $0.15/sec when integrated.
post-render · verification
Identity Verifier
live
Hybrig's own face-match scoring across takes. Flags any take where the rendered face drifts more than ~25% from the source character pack and offers an automatic re-roll.
Runs on
On your GPU
Cost
Free. Runs on CPU in the worker.
Export
export · editing
DaVinci Resolve
Blackmagic Design
live
Final color grade, audio sweetening, and timeline assembly. Hybrig exports cut-ready clips that drop straight into a Resolve project.
Runs on
Desktop app
Cost
Free (Resolve free version covers everything we need).
export · compositing
FFmpeg / CapCut
live
Caption burn-in, aspect ratio reframing (9:16 / 1:1 / 16:9), and platform-specific export presets. Runs server-side via fal.ai or locally via the worker.
Runs on
Hybrig server
Cost
Free. Server-side ffmpeg costs us only fal compute cents per export.
export · compositing
HyperFrames
live
Branded intro/outro/chyron compositor. Lets you stamp consistent brand frames onto every render with a single template.
Runs on
On your GPU
Cost
Free. Built into Hybrig.
Section 04 · Recipe Library
Named workflows, end to end
Recipes are pre-wired sequences of models + tools. Free local, full cloud, or hybrid — pick one and we'll handle the routing.
UGC Spokesperson (Free Local)
Free local
Talking-head clips on your own GPU. Zero cloud spend, ~75% Seedance quality.