AI filmmaking glossary
Definitions of the terms AI filmmaking tools assume you know: text-to-video, character consistency, previs, animatics, generative extend, world models.
- text-to-video
- Generating a video clip directly from a written prompt describing the scene, action and style. [source]
- image-to-video
- Animating a still image (often a storyboard frame or concept art) into a moving clip, sometimes with specified camera motion. [source]
- character consistency
- Keeping a character's face, wardrobe and proportions identical across separately generated shots, typically via reference images or persistent character profiles. [source]
- camera control
- Directing virtual camera moves (pan, zoom, dolly, orbit) in generated video so shots match a director's intended blocking and framing. [source]
- temporal coherence
- The frame-to-frame stability of generated video — objects, lighting and identity staying consistent over time instead of flickering or morphing.
- previs (previsualization)
- Creating rough versions of scenes before production — storyboards, animatics or low-fidelity video — to plan shots, blocking and edits. [source]
- animatic
- A timed sequence of storyboard frames with audio, used to test pacing and story flow before shooting or final animation. [source]
- script breakdown
- Analyzing a screenplay to tag every production element (cast, props, locations, VFX, wardrobe) needed per scene; now commonly automated with AI. [source]
- generative extend
- Using generative AI inside an editor to add extra frames or room tone to the head or tail of an existing clip, e.g. to cover a transition. [source]
- voice cloning
- Synthesizing a specific person's voice from recordings, used in film for de-aging voices, dubbing, or preserving performances with consent. [source]
- world model
- An AI model that simulates an environment's dynamics, enabling explorable or interactive generated scenes rather than fixed linear clips. [source]
- text-based editing
- Editing video or audio by editing its transcript — deleting a sentence in the text removes the corresponding media. [source]