We have been generating video in Google Flow for two weeks — an ad, a set of product pieces, and now a short narrative film. The most expensive lesson so far had nothing to do with writing better prompts. It was that we were quietly using the wrong model and did not know it.

These are the notes we kept. Every one of them is something we measured on a real run, not something we read.

The watermark tells you which model actually ran

We generated a logo reveal on 18 July and it came out perfect: our gold ring, the four dot-nodes, the calligraphic glyph, all reproduced exactly. We generated the same reveal the next day and the glyph came back as an angular Latin “k”.

Same brief. Same reference file. Different result.

The cause was not the prompt. It was that the project’s video default had been set to Omni Flash rather than Veo, so every new clip was silently being generated by a different model. The tell was sitting in the corner of the frame the whole time: the good clips carry a Veo watermark, the broken ones carry a Gemini ✦ watermark.

If the output says Gemini, you were not on the model you thought you were on. For anything carrying a brand mark, select a Veo tier explicitly and confirm it by the watermark on the result. This one check would have saved us the entire second day.

Never describe the mark itself

Once we were on the right model, we broke it again — this time with words.

The clips that worked used terse prompts: “Gold ring emblem forms.” The clip that failed described the mark in full: “a clean gold circle with small dots around it and the calligraphic gold glyph centered.”

Those extra words are an instruction to re-invent the logo from a description. The model stops copying the reference and starts drawing what the sentence says, and a description of a glyph is not the glyph.

Load the real logo file as an Ingredient, then say only that the emblem from the reference forms. Put latin letter k, redrawn glyph in the negative. Terse plus a reference is faithful; verbose is drift.

The image model is faithful. The video model is not.

Even terse, video-generated logos kept drifting across multiple attempts. So we generated the identical reveal as a still image instead — and it came out correct on the first try.

That is the whole shape of the fix. Generate the “formed logo” frame as an image, where the model reproduces it faithfully. Then use that correct frame as the first or last frame of a video and let the model interpolate to or from it. The glyph cannot drift, because at least one end of the clip is already right.

The general rule underneath: image models copy a reference, video models re-draw it every frame. Anything that must survive frame-by-frame should enter the video as a frame, not as a sentence.

The set held. The actors didn’t.

This was the finding we least expected, and it changed how we plan a shoot.

We ran two clips from a single approved room still — a slow push-in and a full pan right. Every prop, the wallpaper, the drawings, the bookshelf and the printer stayed exactly where they were. Zero invented geometry, on the cheapest tier available.

The people were another matter. One prompt read, in effect, “the camera pans right while the man keeps typing.” The man stood up and walked out of frame, and a second character walked in.

One clip, second zero and second eight. The carved wall panel, the sofa, the coffee table, the ashtray and the headphones are all exactly where they started. The man at the desk is gone, and the child who was standing in the doorway is asleep on the sofa.

A camera instruction is not a performance instruction. The model reads a sentence about the camera as permission to decide what the people do. State what every person is doing for the entire duration, and put the action you do not want in the negative — the man stands up, the man leaves his chair.

Two smaller rules from the same pair of clips. Smoke compounds: “a thin ribbon of smoke” had become a floor-to-ceiling plume by second eight, so prompt a single barely-visible wisp and negative the rest. And a camera move must land on a subject — our pan ran past the doorway and spent its final two seconds on empty shelving.

A room bible fixes one frame and breaks the next

Our worst idea was the most reasonable-sounding one.

To get an alternate camera angle of an approved set, we uploaded the master still and pasted a full fact sheet of the room — back wall, left wall, doorway immediately right of the desk — then asked for a new camera position. The reverse angle grew a second door, on the wall it could actually see. The real doorway was behind the new camera, and the model still had to satisfy the stated fact, so it put one where we could see it.

A fact sheet suppresses hallucination inside a frame and causes it across frames. Describe only what is in front of the lens.

The cheaper conclusion is to stop generating angle stills altogether. The video model infers the room in three dimensions from one first frame, and will pan, dolly and push to reveal parts of the set the still never showed — consistently, for the price of a draft clip. Image models have no such spatial model. One approved master still plus motion prompts beats any number of angle stills.

The same clip at second zero and second eight. The camera has moved in and the room holds: the wallpaper, the drawings, the monitors, the printer and the bookshelf are unchanged. The thin wisp of smoke on the left has become a visible column on the right.

What this costs

Worth budgeting honestly: iterating a set design through image edits is cheap per step and expensive in aggregate. This run spent roughly ten image generations before the set locked. Stills are a real line item. Once the set is right, stop editing it and move to video.

If you are starting this week

Confirm the model by the watermark before anything else. Keep brand marks out of your sentences and in your references. Approve one master still, then get every angle from camera moves rather than new stills. Say what each person does for the full eight seconds. And expect the set to behave better than the cast.