AI Agents · Social Video · YouTube, Instagram, TikTok

crew.

63 AI agents that take raw footage to a published video, then out to YouTube, Instagram and TikTok. I built it to run my own channel.

AI in UX Systems Design Information Architecture Agent Design Content Ops
Project
Crew — Agent System
Context
Self-directed · 55 videos live
Role
Designer, builder, user
Year
2026
Palette

"Every mistake it published passed every check it ran."

Why it exists The problem The system The edit What it makes The rules Reflection
Crew — an agent system for a one-person studio
01 — Why it exists

I cannot be in four places at once.

A video essay is a script, a shoot, a cut, subtitles, a grade, music, citations, a thumbnail, chapters, an article and a schedule. I do that across two channels, a website and a book.

The question was never how do I automate this. It was: which decisions actually need me?

02 — The problem

it was confident, fluent, and wrong.

Six real faults reached the published video. Not one had failed a check.
What the system reported
What I saw or heard
0 of 60 captions cross my face
The opening was unreadable
N cards placed
7 of 8 videos had none
5 of 5 study cards placed
One was a raw URL, one was my private notes
delivered: 2469 frames
The sound ran out 165ms before the picture
The counts were green every single time.

The system measured what it did. My audience experiences what it produced. Nothing had ever compared the two.

03 — The system

eleven stages, three of them mine.

The pipeline from raw take to scheduled upload, with three stages marked as needing me
Filming, approving a device, and watching one finished video. Everything else runs under a rule I approved once.
The 63 agents grouped into families by stage of work, with routers and a shared foundation
An agent I cannot find is worse than no agent, so the naming is the interface.
04 — The long-form edit

what one take goes through.

Seven frames from a finished 9:48 video, marked on a runtime bar, each labelled with the device it shows
The template stays constant. What goes on top is specific to each video — if I name a book, the cover appears.

Zero fillers in 495 words. I don't say "um", I restart. So the filler cutter every editor sells is worth nothing here, and the restart cutter found thirteen in one take.

A raw 5:51 take shown as a bar of kept, restart and dead-air segments, collapsing to a 2:02 delivered cut
Nothing I said once is removed. Every abandoned run has to come back straight after, or the cut is refused.
Eight steps of the long-form edit, six run by the system and two by me
Judgement lives in a decisions file, never in the thresholds. I loosened one once and it destroyed the video’s spine.
05 — What it makes

everything below came out of the system.

chapter cards

Chapter card reading 'stage one, the merge'
Placed on the word I say, or dropped.
Chapter card reading 'part 1, the fear, said out loud'
The eyebrow says where you are in the argument.
Chapter card reading 'conversation three, how you fight, and how you stop'
Fifty generated per batch.

research figures

Bar chart figure titled 'what actually predicts living longer'
When I cite a study, the system draws it. This nearly did not exist: the parser wanted one exact heading, so 30 of my 31 scripts were invisible to it.
06 — The rules

six rules, each paid for by something that shipped.

Verify the artifact, not the process

Pull the frame and look at it. A step's return value is never evidence of what it produced.

Refuse rather than approximate

Three fallbacks guessed where a card should go. All three were deleted, not improved.

Zero is louder than missing

"Found nothing" and "there is nothing" are the same silence. 0-of-N is now a fault.

Only destroy what is provably unusable

A guard I wrote deleted a finished render over a bug in the guard.

Approve the class, not the instance

Asked to approve choices video by video, I stopped answering. That is a queue, not consent.

Match the parser to the human

I write for myself, not for a parser. The step accommodates the person.

The pattern I am proudest of

choosing a look by looking at it

Which font, what size — nobody can answer that in the abstract. So I made it a rendering job: seven directions, my sentence, my footage, full size. I decided in one pass.

Seven typographic directions rendered over the same video frame, with Playfair Display marked as chosen
Shown side by side here, but I reviewed them one at a time, full size — a grid shrinks every option to a thumbnail.
Caption set in cream Playfair Display over the video frame
Chosen. Sentence case, ragged, one italic accent word.
Caption set in all-caps with a yellow accent word over the video frame
Rejected — and it was the existing default until I saw it next to the rest.
07 — Reflection

what it changed about how I design.

Lesson 01

Non-determinism breaks a quiet assumption

That the same input gives the same output, and that a system knows when it failed. Neither holds for a model.

Lesson 02

Deleting three working features was the best call

They worked, and they were the largest single source of wrong things on screen. Improving them would only have hidden that better.

Lesson 03

Still unsolved: the last look

The checks do the mechanical half so that watching one video is enough, not so that watching none is. I no longer think that should be removed.

"Verify what I would see or hear. Never that a mechanism reported success."
— Crew, 2026
← → navigate · Esc close