A build-in-public case study: one founder, one AI agent, 90+ creator videos, and a children's book about a small brave owl β wired into a system that tests everything, learns from every dollar, and traces every sale back to the exact video that caused it.
The idea
Most ad accounts let the algorithm pick favorites before creatives get a fair shot. We flipped it: every video gets the exact same budget, the data picks winners, and the reasons they won get extracted and fed into the next generation of content.


What actually happened
77 Instagram videos from 4 UGC creators, pulled via the creator platform's own API, integrity-checked (a whole batch once arrived silently missing audio β caught by verification, re-downloaded with the fix). Every video transcribed locally with Whisper. Total media cost: $0.
Every transcript scored against a brand rubric: does it mention the book? The tree-planting? The lessons? Is it book-focused or life-focused? Then a visual pass caught what transcripts can't see β 21 "silent" videos were actually text-overlay ads, some of the strongest brand content in the library.
All 77 videos went live simultaneously β one campaign, 77 ad sets, $2/day each, budget optimization off so the algorithm couldn't starve anyone. Three days, ~$460, and every creative got a genuinely fair read: ~660 impressions and 15β25 clicks each.
1,370 visitors, zero attributed leads β the conversion pixel was firing from a sandboxed context with no click ID. We rebuilt tracking so every ad click writes its ID to the site, every captured email gets tagged in the store with the exact ad and placement that earned it, and every conversion event carries the thread back to Meta. No attribution model required β the data is just there.
Top 12 creatives moved to lead generation, then the best into sales campaigns behind an advertorial. New creator videos now flow in on a conveyor: export, dedupe across platforms, rank by organic views, download, transcribe, build, launch β a fresh 15-video sales test went from CSV to live campaign in one afternoon.
The data
Winners averaged a 40% hook rate against the losers' 19%. The best creative delivered site visitors at $0.07 β six times cheaper than the account's historical $0.42 average. And the patterns were unmistakable:
Five of the top twelve ads were anti-iPad content. The audience β moms 25β45 β is actively fighting this fight and stops scrolling for it every time. The book's natural position: the thing that replaced the tablet.
Book-focused videos won when they opened with parent identity ("the quality I most want to nurture is leadershipβ¦") and arrived at the book as the tool. Same content as the losers, opposite packaging.
The #1 and #2 ads were literally part 1 and part 2 of the same ongoing story. Cliffhangers aren't just for TV.
We expected "bad organic = good ads." The data said no: the same story-style hooks win both feeds. The real divide is story-style vs. list-style.
Next frontier
The winning patterns became script templates. The scripts got a cast: named, reusable AI moms β designed against a strict spec (real kitchens, tired-but-warm, zero influencer glam) so sequels stay consistent. Meet Meg, generated from a casting prompt, headed for a head-to-head test against the human winners on cost-per-result.


How-to
Everything below was built in conversation with an AI agent (Claude). No dashboards clicked, no code written by hand. These are my real prompts, pasted verbatim β typos, rambles, and all β because the whole point is that you don't need perfect prompts, you need clear intent.
All quotes are copy-pasted from the real chat. Yes, including the spelling.
What happened: the agent found the 39 downloaded videos, moved them into an organized library, transcribed all of them locally for free, wrote the campaign architecture doc, and explained why my instinct about budget-starving was right (and how to do it with one campaign instead of 39).
π‘ State the goal, the fear, and the half-formed plan. Let the agent structure it.What happened: it verified all 77 budgets via the API, set a $500 hard spend cap on the campaign, and explained the worst-case math (Meta can flex 25% on a single day, never more than 7Γ weekly). Every campaign since gets built paused with a cap before a cent moves.
π‘ Ask for hard caps, paused-by-default builds, and verification β not promises.What happened: the agent had recommended a dedicated landing page. Instead of defending it, it loaded my homepage in a mobile browser, looked at it, and retracted the recommendation β the homepage already was a landing page. The real problems were elsewhere (attribution), and we fixed those instead.
π‘ When something smells off, challenge it. A good agent verifies instead of arguing.What happened: it joined ad performance to transcripts and frame-grabs, compared winners vs losers statistically, and wrote the five patterns into a learnings doc β which then became the script templates for the next generation of content. ~$460 of spend became a permanent playbook.
π‘ The test isn't the product. The extracted learnings are the product.What happened: dynamic URL macros on every ad, plus a tiny snippet pushed to the store theme via CLI β so every captured email is permanently tagged with the exact ad, placement, campaign, and even which form on the page converted. No attribution modeling, just facts.
π‘ If you can say the question ("which ad drove this email?"), the tracking can be built to answer it.What happened: the download job was already running. The agent updated the target list, messaged the running worker to drop the extra video, and added a safety exclusion downstream. Total ceremony: one sloppy sentence.
π‘ You can change your mind whenever. Say it plainly; the agent handles the cleanup.What happened: a scheduled task now runs every morning: pulls every ad's numbers, computes cost-per-result, ranks creatives, flags the kill list, and saves the data to disk so future analysis can join against it. I read it with coffee and make the calls.
π‘ Automate the reporting, keep the decisions. That division of labor is the whole trick.Live right now
| Campaign | What it tests | Budget |
|---|---|---|
| Proven winners β advertorial | The 5 battle-tested creatives selling the book | $36/day |
| Fresh batch β advertorial | 15 brand-new creator videos, same funnel | $50/day |
| Episode funnel test | One 3D animated episode β 3 funnel depths (advertorial vs product page vs straight to checkout) | $30/day |
Every morning at 8am, an automated readout ranks every ad by cost-per-purchase, joins it to the video's transcript and placement, and flags what to cut. The humans drink coffee and make the calls.