The video that was almost perfect, except for the price written on it
Early on, we let the video model generate the on-screen text as well: price, sale date, coupon. The scene came out beautiful, the numbers came out broken. A seven instead of an eight, a letter swapped with a symbol that doesn’t exist in the language, and a date that looked right at first glance only to turn out wrong when a customer called asking why the sale was no longer valid.
This happened more than once. In a campaign with a real media budget behind it, such an error isn’t “just another bug”—it’s money going towards an ad with incorrect information. From then on, we set a non-negotiable rule: no AI model writes text that appears on screen, neither in our videos nor in our images.
The rule sounds obvious when explained. In real-time, when the model released this week promises “perfect text in every language,” it’s easy to be tempted and compromise. We don’t get tempted, because we’ve already seen how that ends.
Our test is simple: take ten versions of the same scene with the same requested text, and count how many come out completely accurate. A new model scoring nine out of ten still means one version might come out with a wrong number, and there’s no way to know which one in advance. Ninety percent accuracy is an impressive number in a demo, but unacceptable in an ad advertising a price.
Why generative models are unreliable for text
Video and image models learn to draw shapes that look like letters, not actual letters. This works reasonably well in English with short words, but falls apart in Hebrew, with exact numbers, and with long text. Even when the result looks fine at first glance, a closer inspection reveals a missing letter or a swapped digit. In Hebrew, the gap is even more prominent: final letters pop up in the middle of words, and acronym quotation marks disappear or double. A model trained primarily on English hasn’t seen enough Hebrew to know how it should look.
The problem worsens when the text must be accurate every single time: price, expiration date, phone number. There is no “almost right” with a product price. Either it’s correct, or the ad causes damage.
Added to this is the damage to trust. A customer who sees an ad with a broken date assumes the business didn’t check what it published, even when the issue is purely technical.
And there is the matter of scale. A single ad can be visually inspected before it goes live. A campaign with hundreds of versions automatically distributed by target audience, age, or region cannot be checked manually one by one. The moment the model isn’t 100% reliable, any broad distribution becomes a gamble.
The solution: separate rendering, precise overlay
The rule we set became the core principle of our production engine. The model generates only the visual scene, without text. Every piece of text and number is rendered separately, like a webpage, and then overlaid onto the video. The price, the date, and the brand look exactly as entered, because they weren’t “recreated”—they were printed.
This also solves the consistency issue. When you need ten versions of the same ad with different dates, the text is re-rendered in each version without touching the scene, instead of generating the entire video from scratch every time.
Technically, the font, size, and position of each field are predetermined in a template, and only the value changes. That way, the price always sits in the exact same spot at a legible size, even when the number itself varies from version to version, or when it’s one digit longer than expected.
The principle relies on something simple: a browser rendering engine has known how to draw precise text for twenty years. This is not a problem that needs to be reinvented with an AI model. Combining the proven accuracy of traditional rendering with the visual power of a new video model delivers both advantages without compromising on either.
What the client sees in the end
The customer doesn’t see layers. They see a video ad with a legible price, a correct date, and a logo that looks just right. Separating text rendering from video generation is the difference between a campaign that runs smoothly and one that requires pausing and fixing after already going live.
For retail chains that update prices weekly, this saves substantial production time: you change a single value in a file instead of rebuilding a new video. The same background recurs week after week, with only the text layer being replaced. A client who wants to verify for themselves receives the values file alongside the video, seeing line-by-line what was printed on screen.
What this requires when we adopt a new tool
Every plugin or tool added to our production engine is tested against this principle. If it requires an AI model to draw text on the screen, it doesn’t get in. This applies to Veo, Kling, and Seedance in video, and to Nano-Banana Pro and Gemini 3 Pro Image in images. The models generate the visual background and stop there.
It’s not always fun sticking to this rule when a new model looks impressive in a demo. But thanks to it, our clients don’t have to scrutinize every ad before it launches. The rule will change the day a model demonstrates full accuracy—not almost full—on Hebrew text across hundreds of tests.
If a previous campaign of yours got stuck with incorrect text in AI video, we’d love to share how our engine solves this without sacrificing generative production speed.