Why Is AI Video Generation So Expensive?
The previous page said one image costs a few cents. Video? It's priced by the second — roughly ¥1.5 to ¥5 per second. Generating one 10-second clip costs as much as tens of thousands of chat messages. Why?
Because a video isn't "one image" — it's twenty-plus images per second that all have to stay coherent, plus synchronized sound and physically plausible motion. The output volume is dozens of times an image's, and the difficulty rises much more than dozens of times.
Frames have to "remember each other"
The character's clothing color, the cup in the background, the direction of the light — no frame is allowed to change its mind. The model has to keep hundreds of frames "in mind" and aligned with each other — far harder than painting hundreds of independent pictures, and the reason characters in early AI videos would "change faces" mid-walk.
Physics has to hold up
Apples must fall, water must flow downhill, hair must move with the wind. The model has to "intuit" the laws of physics from massive amounts of video to avoid giving itself away — this ability burns the most training cost.
Sound has to match the lips
The newest generation of models (Veo 3, Sora 2, and others) generates dialogue and sound effects directly — lip movements, footsteps, and ambient sound all synced to the picture. That's doing the video job and the audio job at once, which is why "with sound" costs noticeably more than "silent."
✅ What this page wants to share with you
- Video is priced by the second: ¥1.5–5 per second — 10 seconds ≈ tens of thousands of chats
- Why it's expensive: twenty-plus frames per second × frame-to-frame coherence × plausible physics × synced audio
- The money-saving order: polish the script in text → lock the visuals with images → generate video last
- The trend is your friend: prices drop every year — today's "expensive" is temporary