Blog

AI video ads: what already works and what still goes wrong

You can now generate 20 seconds of video with synchronised audio from a single sentence. After months of producing with AI, here is the honest list: where it saves weeks, where it still ruins the material, and what should never be synthetic.

August 10, 2026 · Agência Primeira Página

AI video ads: what already works and what still goes wrong

Since late July 2026 you can describe a scene in one sentence and get up to 20 seconds of video with synchronised audio — dialogue, effects and ambience — without filming anything. FLUX 3, from Black Forest Labs, put all of that into a single model, generally available through an API at 1080p since 5 August. The question that matters if you sell something isn't whether the technology impresses. It's what it already solves in your product video — and what it still gets wrong.

We have been producing AI video for months, for the agency and for clients. What follows is the honest list: where it saves weeks and where it still ruins the material.

What actually changed

Three things left the lab in 2026. First, native audio: until now sound and picture were generated in different tools and synced by hand; now they come out together — footsteps, breaking glass and voice in the same pass. Second, character consistency across scenes, historically the Achilles heel of AI video, where the same person changed face at every cut. Third, scene chaining, which lets you build a continuous sequence instead of loose clips.

Alongside it came a logistics layer: tools that pick the model for you based on priority — quality, speed or cost. For an agency producing volume that matters more than it sounds: tests on the cheap model, final delivery on the expensive one.

Where it already works

  • Atmosphere and b-roll. The supporting shot — the street, the office, the set table, the factory — comes out convincing and costs almost nothing next to a shooting day.
  • Product in a scene. Putting the product into an environment that doesn't exist or would be expensive to build. Results are strong here, especially combined with the product's real 3D model.
  • Variants for testing. Generating five different openings of the same ad and finding out which one holds the viewer. That used to be prohibitive; now it's an afternoon.
  • A short spokesperson. Eight or ten seconds of someone talking to camera works — as long as the script is short and the scene is clean.

What still goes wrong (we've been burned by all of it)

  • On-screen text turns to scribble. Any lettering the AI has to draw — a sign, a label, an app interface, a subtitle — comes out garbled. Fix: all text goes in afterwards, by compositing, never by the AI.
  • The logo gets recreated, not copied. On a wall in the background it usually comes out faithful. On a t-shirt, a badge or a cup, the AI redraws it freely — and the client spots it immediately. There we stamp the real logo on top.
  • The voice misses the intonation. Synthetic narration nails pronunciation and misses intent: stress on the wrong word, a question that lands as a statement. Regenerating is cheap; listening back is mandatory. Words the AI insists on reading in English we spell phonetically in the script.
  • Continuity between clips. Even with the new models, a character only stays the same if the description is identical and specific in every prompt. One word changed, one face changed.
  • Hands, reflections and small objects. Much improved, but still where the shot gives away that it is synthetic.

The workflow that works

  1. Approve the still before animating. Generate the static scene first, fix framing, branding and wardrobe there — and only then turn it into video. Fixing one frame is cheap; fixing 20 seconds is not.
  2. Separate what is AI from what is programmatic. Scene, environment and people from the AI; text, interface, charts and logo by compositing. That hybrid is what looks professionally finished.
  3. Music underneath, voice in front. Music goes in at the mix, well below the narration, without regenerating a take that already works.
  4. Regenerate instead of repairing. When a take comes out almost right, the temptation is to retouch it. Asking for another one is faster and better.

When not to use AI in video

Some material should not be synthetic, and it isn't about quality. Customer testimonials, before-and-afters and the real person behind the brand have to be real — if the audience finds out the "happy customer" was generated, the damage is not aesthetic. The same goes for anything that works as proof: treatment results, an image of a property that exists, a team photo. And when a piece is plainly a generated illustration, saying so in the caption costs nothing and avoids the awkward conversation later.

What changes in the budget

The cost moved. It used to sit in production: the shooting day, the crew, the gear, the location. Now the biggest line is direction and review — deciding what the scene has to say, writing the right script, picking from thirty takes the one that doesn't betray its origin, and finishing it properly. That is still human work, and it is what separates a video that sells from a video that impresses for ten seconds.

That is how we do promotional video production: AI where it is good, compositing where it fails, and a script that exists before any generation.

Sources

FLUX 3 features and dates (video up to 20 seconds with native synchronised audio, character consistency, general availability via API at 1080p from 5 August 2026) come from the Black Forest Labs announcement. The topic reached us through the Café com AI newsletter. The rest comes from our own production work.