Contents
In the first post about Marguerite I said I would publish the flat weeks along with the wins. This is the first one. Nothing failed in public. The account is fine. What failed was the way we made her third video, and I think the failure is more useful than the two videos that went smoothly.
The mermaid teaser is a departure. The first two were Marguerite alone, talking to camera or doing a mirror reveal. This one is a thrift-store comedy. She pushes a cart down an aisle, pitching the pieces for an eleven-dollar mermaid costume, and behind her a friend named Moreen slowly gets a lampshade stuck on her head. Marguerite never notices. The audience does. That is the whole joke.
Two characters, a background gag, bystanders, four clips. On paper it looked like a small step up from what we had already done. It was not a small step.
Follow along on Instagram, Facebook, or Pinterest — the next look posts there before it ever makes the blog.
The plan looked good
The morning went well. We wrote the script in the terminal, went through eight versions in about forty minutes, and then handed it to Rae, a viral-video strategist persona we run as a BMad agent. She cut it from forty seconds to thirty, moved the hook, killed a second open thread, and turned the ending into a loop. Good notes. All of that is on a local storyboard page with a shot list, timings and the four voiceover lines.
Voice went through its own false start. We’d already paid for a Fish Audio subscription for Marguerite’s voice model, and only after paying found out the subscription covers the website, not the API — API credits are billed completely separately, so the money we’d already spent didn’t buy Claude Code a way to generate voiceovers on its own. Every line meant opening Fish Audio’s site by hand, typing it in, rendering, downloading, then copy-pasting the file into the project folder. That’s not a workflow, it’s a chore, and we didn’t want to keep doing it. A day later we tried Gemini’s Flash text-to-speech instead, and it was simply better: it takes a plain-language style instruction and inline delivery tags, a short pause, a laugh, right in the same request, and it answers to a normal API call from the terminal — no browser tab, no copy-paste. We switched, and every voiceover since has come from there.
I recorded the four voice clips. We generated forty seconds of thrift-store room tone, five location stills of the store, a character still for Moreen, one for a guy who walks past, four start frames of Marguerite, and composites that put Moreen into each of them. I looked at all six start frames on the review page and said “all are good.”
Then we rendered clip one. It came back fine. Twelve seconds, lip-sync clean, Moreen in the background. One note from me, one re-roll, approved.
Then clip two.
Getting the background actor right
Clip two is eight seconds. Marguerite holds fishnet tights to her cheek and explains that fishnet plus chrome eyeshadow is how you fake mermaid scales. Behind her, Moreen deals with the lampshade — and getting that right turned out to be the actual cost driver.
The problem was never Moreen’s face. It was subtlety. She needed to genuinely struggle with the lampshade, not perform struggling, then bump into a shelf on her way out without it reading as staged. More than once the model skipped the contact entirely: a pair of glasses would slide off a shelf with nothing in frame close enough to have touched them, so it just looked broken instead of funny. Getting a believable, unforced physical beat out of a background character took a lot more iteration than any one number captures. The five takes below are the ones that changed the direction; there were quite a few more in between, just testing whether the physics read as real.
Take 1. Moreen tugs at the lampshade, it flies off, hits a mannequin, Marguerite flinches and looks back. Seedance even produced the crash sound on its own. My note: Moreen looks like she is acting for the camera. She is grinning at the lens. She should be genuinely struggling and never look at us.
Take 2. We re-edited the start frame so she is side-on and straining, no smile, and rewrote the prompt as documentary direction. My note: she still turns to the camera. I do not want to see her face at all. After it comes off, move her out of frame.
Take 3. Back of head or profile only, and she exits after the lampshade flies off. My note: put her in another aisle, further back. I want her stuck, pulling at it, and then she falls out of view. And Marguerite should not react at all. Ignore her completely.
Take 4. New start frame with Moreen one aisle back, small, out of focus, seen from behind. She tugs, staggers, topples out of view. No crash, no reaction from Marguerite. My note: instead of falling, have her give up and walk off, bump into a rack, and go round the corner. Also, when Marguerite holds the fishnet to her face she needs to be applying makeup through it, so the viewer sees the stencil working.
Take 5. Brush edited into the start frame. She dabs shimmer through the net, peels it away to show the scale pattern on her cheek, and finishes the tip. Moreen, far back and out of focus, gives up on the lampshade and wanders off toward the racks with it still on her head. Approved.
Take five is genuinely better than take one. That is what makes this hard to write off. Every note I gave was a real creative improvement. The problem is that every one of them arrived after a paid render instead of before it.
The numbers
I want to be exact, because “it took a while” hides the shape of the problem. Claude itself cost nothing extra here — that runs on a flat-rate subscription. The real money was Higgsfield credits.
| Higgsfield generation jobs | 37 (28 images, 7 video, 2 audio) |
| Video credits spent | 448 |
| Real cost, at the Higgsfield plan rate | ~$19.63 USD / ~$27.76 CAD |
Seventy percent of that spend went into takes we threw away, for clips totalling eighteen seconds.
What actually went wrong
I have gone back through the whole transcript, and it is not one mistake. It is five, and they compound.
I approved the look, not the action. When I said “all are good” about the start frames, I was checking whether Marguerite’s face held, whether the cuff was on the correct wrist, whether the shower curtain was in the cart. I was not asking what each person in the frame was about to do. But a video model treats the start frame as direction. Moreen was beaming straight at the lens in that frame, so in the clip she performed for the lens. The frame was the instruction. Nobody, including me, read it that way.
We changed the character and did not notice the contradiction. Moreen started as a deadpan buzz-cut. In the space of five minutes I asked for long hair, then petite, then “always have a smile.” Reasonable requests. But the joke depends on her being oblivious in the background, and a character who is always smiling at the camera is the opposite of oblivious. The direction contradicted the comedy engine, and it went straight into the still, and from the still into three wasted takes.
We directed by watching renders. Each note I gave was something I could only articulate once I saw it moving. Fair enough. But at 56 credits a take, “let me see it and then tell you” is an expensive way to find out what you want. The storyboard existed as text. There was no version of it I could watch.
The script changed during the shoot. The makeup-through-the-fishnet demonstration is the best idea in the video, and it was not in any of the eight script versions. It came to me on take four. In a real production that is a rewrite mid-shoot, and everyone knows what that costs.
Everything got rewritten on every note. Each round trip re-edited the prompt, re-edited the storyboard page, re-rendered a start frame, refreshed the review page, tiled the frames, and described it all back to me. Helpful in the moment, but it’s exactly how a session balloons.
What we are changing
The short answer to “should we be storyboarding this first” is yes, but not the way we did. A text shot list is not a storyboard. Here is the production order from the Wednesday Addams teaser onward.
- Casting brief before any render. One line from me per character: age, build, hair, default expression, and the one thing they never do. Two takes maximum, then locked. Moreen would have cost two images instead of five.
- Blocking card per shot, approved for action. Each start frame comes with a plain-language card: who is where, who looks where, what happens in beats one, two and three, what the last frame looks like, and what is not allowed. I approve the card, not just the picture. “Looks good” is no longer an approval.
- An animatic before any video credit. The start frames, held for the length of each voice clip, stitched with the audio. Thirty-six seconds I can watch end to end. That is where “spiritually, a tail” would have been caught, where the missing makeup demonstration would have been caught, and where Moreen’s grin would have been caught. It costs nothing.
- End frames where the end state matters. Seedance takes a start image and an end image. Moreen gone round the corner with the lampshade on the floor is an end frame, not a paragraph of prompt hoping the model gets there.
- Test blocking at 480p. Seedance renders at 480p for less. Blocking questions (“does she face the lens”) get answered cheaply, and only an approved blocking gets a 720p render.
- Two-take rule. A clip gets at most two paid renders. If the second is wrong, the fix is upstream, in the frame or the card, not another roll. New ideas that arrive after a render go into the next video, unless I explicitly decide they are worth a re-shoot.
- Separate the sessions. Storyboard and animatic in one session, ending with an approved page. Production in a fresh session that starts from that page and nothing else. Review notes batched, not one per render.
The house rule that came out of this is now written down where the assistant reads it every session: background characters react to the gag, never to the lens. Only the person talking to camera knows the camera exists.
Where the mermaid teaser stands
Clips one and two are the eighteen seconds above. Clips three and four, the moment Marguerite finally turns and sees Moreen and the checkout with the pile of finds, got made afterward under the new rules — and the finished four-clip video is live now on Instagram and Facebook. Whether the new rules actually held that second half’s cost down is its own accounting; this post is just about what the first half cost, and why.
This is the part of AI Secret Sauce nobody demos
Anyone can show you a good render. Weeks two and three of the AI Secret Sauce class at Cowork Chilliwack show what happens between the renders: locking a character, directing a video model with frames instead of adjectives, and keeping the cost of a wrong guess small. Marguerite's outtakes are course material now.
See the course →Keep reading
- Meet Marguerite: We Built an AI Influencer to See If She Can Hit 21,000 Followers by Halloween: the first post, the target, and how she is built
- Meet Rae: The AI Strategist Who Rewrites Every Script Before It’s Shot: the persona whose notes are all over this script
- We Tried Building an AI Team on Grok. Here’s Why We Moved to Self-Hosted Rakazo Instead: the bot platform that posts for her
Comments
Loading comments…
Leave a comment