Social & content
Faceless short-form, captioned by one POST.
No camera, no on-screen presenter. Compose stock footage, royalty-free music, and your own voice-over into a JSON project, layer word-timed captions, and get back a hosted video rendered frame-accurately by FFmpeg.
Use case
A faceless channel, built from JSON
Faceless content lives or dies on two things: the right B-roll and captions that hold the eye. Zvid gives you both from a single JSON document — pull royalty-free footage, images, GIFs, and music from Zvid's stock library, then drop word-level captions over the top in any of 12 animated modes rendered through ASS/libass.
You bring the voice-over — upload your own audio track and Zvid mixes it under the visuals with per-track volume, trim, speed, and fades. There is no synthesized voice here; you keep full control of the narration. Everything resolves server-side and renders on FFmpeg, so the same JSON produces the same frame-accurate MP4 every time, returned as a hosted cdn.zvid.io URL.
What you get
Everything a faceless channel needs
Twelve caption modes
Karaoke, highlight, typewriter, pop, bounce and seven more animated modes render through ASS/libass with word-level timing, active-word styling and 8 on-screen positions.
Bring your own voice
Upload your voice-over as an audio track and mix it under the visuals with per-track volume, trim, speed and fade in-out — Zvid never synthesizes voice, so the narration stays yours.
Stock from five providers
Pull royalty-free footage, images, GIFs, and background music from Zvid's stock library — no camera required.
Captions from a transcript
Import SRT, VTT, ASS/SSA or Whisper JSON, author inline, or set up to 20 words per line — active-word styling and a caption box make them TikTok and Reels grade.
Lock in your look
Save any project as a reusable template with up to 200 {{variables}}, so every episode keeps the same fonts, caption style and pacing while the copy and clips change.
Batch a whole series
Bulk-render up to 500 videos per call with per-item validation — invalid rows come back by index while the valid ones queue under one batch id.
12
Caption modes
20
Words per line
5
Stock providers
1,200 free
Credits to start
Under the hood
The features that make it work
Subtitles & Captions
Word-timed captions with 12 animated styles — karaoke, highlight, typewriter, pop, and more. Import SRT, VTT, ASS, or Whisper JSON, or author by hand.
Audio & Music
Layer multiple audio tracks with per-track volume, trim, speed, and fades. Pull royalty-free music from Zvid's stock library or bring your own voice-over.
Templates & Library
Hundreds of ready-made video and image templates across 23 categories. Open one in the editor, or render it with your own data via the API.
Stock Media & Uploads
Search royalty-free media in Zvid's stock library, or upload your own — then drag it straight onto the canvas.
Who it's for
Teams that build this with Zvid
Content creators
Batch a whole content calendar from one JSON template, then use 12 animated caption modes and Zvid's stock library to keep every faceless short on brand. No watermark on any plan.
Explore →Social media managers
Batch a week of reels, shorts and stories from one JSON project, then resize it to every platform aspect ratio. 12 animated caption modes and 400+ templates keep the look consistent.
Explore →AI engineers & agent builders
Your agent authors the JSON; Zvid renders it deterministically with FFmpeg — video or images, never text-to-video. The MCP server ships 30 tools plus a free pre-flight validator, so you catch errors before a credit is spent.
Explore →Related use cases
More for teams like yours
News & article videos
Turn published content into video. Feed headlines, articles, or an RSS feed into a template and render a captioned clip per item — automatically, on a schedule, at scale.
Explore →Social media videos
Turn one JSON project into Reels, Shorts, Stories and TikToks — captions the feed rewards, 18 aspect-ratio presets to fit every placement, and batch renders that fill a whole posting calendar.
Explore →Slideshow & photo videos
Turn a set of photos into a hosted video from one JSON project: iterate an image array and Zvid builds a scene per photo with auto-chained transitions, Ken Burns motion, music, and captions.
Explore →FAQ
Frequently asked questions
Does Zvid generate the voice-over for me? +
No. Zvid does not synthesize voice or offer text-to-speech — you upload your own narration as an audio track, and Zvid mixes it under the visuals with per-track volume, trim, speed and fades.
Can the captions match the TikTok and Reels style? +
Yes. Twelve animated caption modes — including karaoke, highlight and typewriter — render through ASS/libass with word-level timing, active-word styling, a caption box and 8 on-screen positions.
How do I keep every episode looking the same? +
Save a project as a reusable template with up to 200 type-preserving {{variables}}, then swap the script, clips and music per episode via the API while fonts, caption style and pacing stay fixed.
Explore more: see every use cases
Ready to ship video at scale?
Start free — your first render is only a JSON payload away.