A video ad is no longer a production — it is a render job. Feed the pipeline your product images and a script template; get back a folder of finished ad variants in every aspect ratio, captioned in correct Vietnamese, voiced, scored and named for A/B testing. This article walks through the pipeline we build for Vietnamese sellers, and the traps that separate a content factory from an expensive toy.
The pipeline, input to upload
- Inputs: product images + a script template. The template is parameterised — product name, key benefit, price hook, call to action — so one template yields a video per product, per angle. Good product photography going in is the highest-leverage investment in the whole chain; if you are compositing AI backgrounds, follow the rules in our brand-consistent AI imagery guide.
- Scene assembly. The engine builds a shot list from the template: hero shot, benefit close-ups, social-proof card, CTA end-card. Each shot is defined once and rendered into every target ratio.
- Three aspect ratios from one edit. 9:16 for Reels/TikTok/Shorts, 1:1 for feeds, 16:9 for YouTube and landing pages. Design safe zones per ratio up front — text that is comfortable in 16:9 gets amputated in 9:16 if the layout is an afterthought.
- Auto-captions with correct diacritics. Captions are generated from the script (not re-transcribed from the voiceover — you already have ground truth). Font choice matters more than people expect: many display fonts render Vietnamese diacritics badly, stacking or clipping marks on ắ, ộ, ữ. Test your caption font against a full diacritic sample before it goes anywhere near production.
- AI voiceover track. Generated per variant from the same script, with accent and emotion chosen for the audience — the details deserve their own article, and they have one: Vietnamese AI voiceovers: accents, emotion and pipelines.
- Music. Licensed library tracks or generated music only. Trending commercial audio that is fine for organic posts is a rights problem in paid ads.
- Batch render with a naming convention. Every output file encodes product, template, variant, ratio and version — e.g.
serum01_tpl-benefit_v03_9x16.mp4. When fifty variants are live, the naming convention is your analytics join key. - Auto-upload with UTM discipline. The uploader pushes finished assets to ad platforms and stamps every destination link with UTMs derived from the file name. Ad-spend reporting then reconciles itself instead of relying on someone's memory. The same discipline applies to organic distribution through your auto-posting pipeline.
Traps that show up after week one
Platform AI-disclosure policies
Major ad platforms increasingly require disclosure of AI-generated or synthetic content, and the rules differ by platform and change over time. Build the disclosure flag into your upload metadata from day one rather than retrofitting it after a rejection — or a takedown.
Brand drift across generations
Generation ten looks slightly different from generation one: colours wander, layouts mutate, the logo shrinks. Lock the brand layer — logo placement, palette, font, end-card — as fixed template elements the AI never touches, and let generation vary only inside the slots you define.
Render costs balloon without a budget
Variants are combinatorial: 10 products × 3 templates × 3 ratios × 4 hook lines is 360 renders. Set a per-campaign variant budget and make the pipeline enforce it. Render the variants your test plan actually needs, not every combination that exists.
Skipping human review on new templates
Where the videos live
Rendered ads uploaded natively to each platform are the platform's problem to deliver. But the same videos on your own landing pages and product pages are your problem — and in Vietnam, single-CDN video hosting fails in ways that will make your ads land on a spinner. Read why one CDN is never enough before you put self-hosted video behind paid traffic.
What good looks like
Measure the pipeline the way you would measure a factory: cost per finished variant, time from brief to live, and the review-rejection rate per template. A healthy setup drives the first two down every month while the third stays low and stable — a rising rejection rate on an old template usually means the product data feeding it has drifted, not that the template broke. And keep the loop honest: pause the variants that lose their A/B tests and feed the reasons back into the next template revision, so the factory learns instead of merely producing.
A seller on this pipeline briefs a campaign on Monday — products, angle, budget — and by Tuesday morning has a reviewed batch of variants live across platforms, each traceable from ad-platform metrics back to a file name back to a template. The marketing factory layer of the end-to-end automation stack is exactly this: creative volume without a content team, with judgement kept where it belongs — in the brief and the review.