63 AI agents that take raw footage to a published video, then out to YouTube, Instagram and TikTok. I built it to run my own channel.
A video essay is a script, a shoot, a cut, subtitles, a grade, music, citations, a thumbnail, chapters, an article and a schedule. I do that across two channels, a website and a book.
The question was never how do I automate this. It was: which decisions actually need me?
The system measured what it did. My audience experiences what it produced. Nothing had ever compared the two.
Zero fillers in 495 words. I don't say "um", I restart. So the filler cutter every editor sells is worth nothing here, and the restart cutter found thirteen in one take.
Pull the frame and look at it. A step's return value is never evidence of what it produced.
Three fallbacks guessed where a card should go. All three were deleted, not improved.
"Found nothing" and "there is nothing" are the same silence. 0-of-N is now a fault.
A guard I wrote deleted a finished render over a bug in the guard.
Asked to approve choices video by video, I stopped answering. That is a queue, not consent.
I write for myself, not for a parser. The step accommodates the person.
Which font, what size — nobody can answer that in the abstract. So I made it a rendering job: seven directions, my sentence, my footage, full size. I decided in one pass.
That the same input gives the same output, and that a system knows when it failed. Neither holds for a model.
They worked, and they were the largest single source of wrong things on screen. Improving them would only have hidden that better.
The checks do the mechanical half so that watching one video is enough, not so that watching none is. I no longer think that should be removed.
"Verify what I would see or hear. Never that a mechanism reported success."