Operations

How I Actually Make Content: A Step-by-Step Deep Dive

The primer said decide what you're the authority on. This is the machine that turns one walk with a voice recorder into a blog post, a video, and a week of posts, and what the whole thing costs.

15 min read1

Every few weeks somebody asks me how much of my content I make myself, and the question always arrives a little skeptically, because they've looked at the blog and the videos and the posts and they've priced out what that would take. They're picturing a team. They're doing the math on a team, and the math doesn't work, so they assume there's something they're not seeing.

There isn't. It's one person with a phone, a walk, and about five hours in a normal week, and the reason that's enough is that almost none of the work is making things. It's capture, and then it's cutting one captured thing into the shapes each platform wants. I've been running this version since June, and it produces a blog post most weeks, a long-form video roughly monthly, and the short pieces cut out of both, without a single week where I sat down to write from nothing.

A content production system, plainly, is the repeatable path a single idea takes from your mouth to a published page. That's all it is. Most of what gets sold as one is a calendar, which is a schedule pretending to be a system, and a schedule doesn't help you on the Tuesday when the box is empty.

So here is the actual order of operations, the tools, the prices, and the four places it breaks.

What it costs, before anything else

Nobody puts this number at the top, so I will. Between about $20 and $122 a month, and the floor is real, not a trick.

MonthlyWhat you get
Floor~$20Phone recorder, local transcription, one AI subscription, native platform schedulers
Typical~$55The above plus Descript for video editing and captions
Ceiling~$122The above plus a scheduling tool that can report into your own dashboards

That's the production side only. The site you publish onto is a separate bill, roughly $21 to $76 a month, and I broke that down in the website deep dive. Nothing here has a setup fee and nothing here has a contract.

The reason the floor holds is that the two most expensive-sounding parts of content production, transcription and writing, both collapsed in price. Transcription runs on my laptop for nothing. And the draft doesn't get written from scratch, so what I'm paying for is editing, not authorship.

The shape of the thing

Before the steps, the shape, because the steps don't make sense without it.

One capture per week. One long piece built from that capture. Everything else cut out of the long piece.

That's it. I record one voice memo, and out of that comes the blog post, and out of the blog post comes the LinkedIn post and the image card, and separately I shoot a long-form video from a script built off the same transcript, and out of the video come the shorts. One idea, one recording, seven or eight published artifacts.

The direction matters and it's the part people get backwards. You go long first and cut down. Starting with short posts and hoping a body of work assembles itself out of them works for a person building a personality and fails for a business trying to build authority, and I made the longer version of that argument in the content primer.

Step 1: Lock one concept, and apply the only filter that matters

Monday, one concept for the week, tied to one of five standing topics I've decided I'm the authority on.

The filter I actually use is whether I care about it that week. Not whether it's strategically optimal, and not whether it fits a gap in the calendar. If I don't have something to say about it on Monday I will produce something hollow by Wednesday, and hollow is worse than late. So the concept moves and the week doesn't.

One row, one week, one topic. Not one row per platform. I ran per-platform planning for a while and it was noise, because the platform breakdown is mechanical once the idea exists and it doesn't need to be decided in advance.

Cost: $0.

Step 2: Write the prompts before you record

This is the step everybody skips and it's the one that makes the difference between a usable recording and forty minutes of circling.

Five questions inside the week's topic. Each one gets a couple of sub-questions, and each one gets a rescue line for when I stall, something like "if you're stuck, just describe the last time a client asked you this." Then it gets printed, because I'm going to be walking and I don't want a screen in my hand.

Writing prompts feels like procrastinating on the real work. It is the real work. An unprompted recording rambles, and a rambling transcript produces a draft with no spine, and then you're editing structure instead of editing language, which is ten times the effort.

Cost: $0.

Step 3: Record the voice memo on a walk

Ten to fifteen minutes, phone in pocket, walking, working through the printed questions in order.

Three rules I hold to. Don't script it, because a scripted voice memo sounds like a press release read aloud. Don't re-record either, and this one took me a while to accept, because the second take is always smoother and always worse. The third is that when you say something clumsy you keep going, since the clumsy phrasing is usually closer to what you actually mean than the clean one is, and fixing it in text later costs nothing.

Walking specifically, and I can't fully defend why. Something about moving makes me explain things the way I'd explain them to a person instead of the way I'd write them.

Cost: $0.

Step 4: Transcribe it locally

The audio goes to text on my own laptop, using an open-source speech model, and it costs nothing and nothing gets uploaded anywhere. For a fifteen minute memo it takes a couple of minutes.

Any transcription tool works here and several are free. The only thing I'd push on is doing it locally if the recording is ever about a client, because the cheap cloud transcribers are the least considered link in most people's chain.

Raw transcript gets filed by date and topic and never edited. It's the source material and I want to be able to go back to what I actually said.

Cost: $0.

Step 5: Draft the long piece from the transcript

Now the writing, except it isn't writing from nothing, it's editing a person who already said the thing.

The transcript goes in with the week's concept and the structure, and what comes back is a draft that has my arguments in roughly my order. Then I spend real time on it, and this is the part that can't be handed off. I cut the throat-clearing and I fix the register wherever the draft got polished and I didn't, and then I go back and put in the specific number or the specific client situation that my talking self gestured at and never actually said.

I run a voice check on every draft before I'll look at it as finished, and it's mostly a list of my own tells. Em dashes, which I overuse in speech and which pile up in drafts. Certain words that read as marketing rather than as me. And a rhythm problem, where three or four very short sentences in a row means the draft has drifted into somebody else's cadence, because my natural rhythm is medium sentences chained together with "and" and "but" and "I think."

That last one is the whole game on voice. Sounding like yourself isn't a matter of vocabulary, it's sentence length and how you join clauses.

Cost: ~$20/mo for the AI subscription, which is the single line item doing the most work in this entire system.

Step 6: Cut the written derivatives

The blog post exists, so the rest of the writing is extraction rather than composition.

  • The LinkedIn post. Text, not video. I tried vertical shorts as the native LinkedIn play and a considered text post outperformed it and read better besides. Shorts still get cross-posted there when the piece is a video promo, but I don't force vertical video into a place where writing fits better.
  • The image card. One pull quote, black on off-white, no accent color. For anything personal I use my own photograph instead of a generated card, because a real photo beats a designed quote every time on that register.
  • Short-form talking points. Bullets for the video cuts, drafted now while the argument is fresh, not later while staring at a timeline.

All of it is drafted in the same week, in one sitting, off the one artifact.

Cost: $0 beyond step 5.

Step 7: Shoot the long-form from a teleprompter script

The same transcript becomes a spoken script, and I read it off a teleprompter, and the reason it doesn't sound read is that it started as speech. It's my sentences, in my order, tightened.

I shoot after the blog is drafted, never before. If the argument has a hole in it, the writing finds the hole, and finding it in text costs twenty minutes while finding it on camera costs a reshoot.

Cost: $0 if you already own a phone and a light.

Step 8: Edit in Descript, and plan for two passes

Editing is text-based, which means cutting a sentence out of the video is cutting the sentence out of the transcript, and for talking-head content that is the whole reason to use it.

Plan on two passes and stop trying to land it in one. First pass proposes the shape, cut selection and captions and audio cleanup. Then I watch it, and then a second pass lands it. On one recent short, the two passes together ran about 68 AI credits, which is well inside a month's allowance and is worth knowing because the credits are the real constraint on that plan, not the seat price.

Cost: $35/mo month-to-month for the Creator plan, or $24/mo if you pay for the year. There's a free tier and it will not survive one long-form video, because it caps at an hour of media a month.

Step 9: Caption to a locked spec

Captions are where amateur video announces itself, so this is settled once and never re-decided:

  • One word visible at a time, synced to speech, so the natural pauses become silence and the lines land
  • DM Sans Medium, 70pt, white with a soft black drop shadow
  • Sentence case, never all caps
  • Positioned 82% down the frame for a standalone short, 68% when I'm seated at a desk and my hands are in the bottom third of the frame
  • No bounce, no pop, no color emphasis on the key beat

All caps reads as aggressive and the bouncing word animations read as somebody else's platform. Both undercut the register I'm going for, which is a person admitting something rather than a person selling something.

Step 10: Publish in order, blog first

The order is not cosmetic and I've broken it and paid for it.

  1. Blog goes live on the site and becomes the featured piece on the homepage
  2. Previous featured piece gets unfeatured
  3. Then, and only then, the social posts fire

Post the LinkedIn and the Instagram card before the blog is live and every click lands on a page-not-found error for however long the gap is. The social posts exist to point at the piece. The piece has to be there.

Step 11: Tag the links, or you'll be guessing

This is the step I was worst at, so it's the one I'll be most direct about.

For months my reach was healthy, and my Instagram was getting more non-followers than followers seeing the work, which is the number you want. And about 84% of my website traffic showed up in analytics as "direct," which is the bucket a visit lands in when nothing tells the site where the person came from. Which means I could not tell you which of those posts sent a single human to the site. I was optimizing the top of the funnel with no instrument on the middle of it.

The fix is unglamorous. Put tracking parameters on every link you post so your own analytics can tell you where the visit came from, keep a single naming convention, and use it every time. Note that most scheduling tools report which site referred the visit and stop there, which is not the same thing and won't tell you which post did it.

Fixing attribution beats chasing more reach. I had the reach.

Cost: $0 for the convention. $67/mo month-to-month, or $53 annually, if you want a scheduling platform that will feed those numbers into your own reporting automatically. Native schedulers on each platform are free and for low volume they're honestly fine.

One set, start to finish

Steps are easy to agree with and hard to picture, so here is a real one with the actual dates and the actual sizes. This is the set that became AI Doesn't Have Taste, and I picked it deliberately, because it worked and it also went wrong in a way worth showing.

WhenWhatSize
Jul 28Prompts written for the week's topic2,678 characters
Jul 29Voice memo recorded on a walk, transcribed10,983 characters
Jul 29Short-form talking points pulled from the transcript1,701 characters
Jul 30Essay drafted from the transcript8,024 characters
Jul 30Video spine built from the same transcript8,608 characters
Jul 30Teaser short published to LinkedIn and YouTube40 seconds
Aug 5Essay published on the site, social posts out8,164 characters, 6 min read
Aug 13Main video shot, three takes, one continuous roll29:12 raw
Aug 19First cut22:50

Five things in that table that I'd point at.

The prompts are the highest-leverage step and they're free. Under 2,700 characters of written questions produced almost 11,000 characters of me talking, which is four times the input, and the questions took twenty minutes on a Tuesday. Every time I've skipped this step the ratio has gone the other way.

The transcript barely shrank on its way to being an essay. Just under 11,000 characters of talking became a published piece of about 8,200. That is not a rewrite, it's a trim, and it's the whole argument for capture over composition. I was not short of content. I was short of a recording.

The cheap thing shipped first and the expensive thing shipped last. The teaser went out on July 30, the essay on August 5, and the video was not shot until August 13. That order is on purpose. The teaser cost forty seconds of edit and told me whether the hook landed with anybody before I spent a shoot day on it. Shooting first is how you find out you were wrong at the most expensive possible moment.

The video came in at nearly twice its target, and the plan was what was wrong. I was aiming at ten to twelve minutes. Raw was 29:12 across three takes. I made six cuts, all of them either false starts or takes that a later take had superseded, which removed 5:39 and left 22:50. Nothing that came out was padding. Getting to twelve would have meant cutting eleven minutes of material that was working, most of it out of the middle of the argument. So it's a sixteen-to-eighteen minute piece, and what was wrong was the estimate, not the take. Targets are guesses. The footage gets a vote.

The best moment in it wasn't in the plan. The script ended on the last line with no call to action, deliberately, and on camera I ignored that and asked the audience a real question instead. It was better than what I'd written. Which is the argument for a spine rather than a script, and against reading your own words back to yourself: the plan is there so you don't get lost, not so you obey it.

One more thing that only shows up in a real set. The six chapter titles on that video came out of phrases in my own transcript, not out of the outline. When you name sections with the words you actually used, the chapter list reads like the video instead of like a table of contents.

Two people worth listening to, who don't agree

I'd rather hand you a real disagreement than a consensus that isn't one, and on content production the disagreement is sharp.

Gary Vaynerchuk runs VaynerMedia and has been making this argument since 2017, when he told everyone to document rather than create. His model is a pyramid. Make one substantial pillar piece, then atomize it into dozens of smaller pieces aimed at every platform, and publish at a volume that most people find unreasonable. His position is that the constraint was never the quality of the idea, it's that almost nobody sees any given piece, so the only rational response is more surfaces and more repetitions.

Cal Newport is a computer science professor at Georgetown who has spent a decade arguing close to the opposite. His term for busy-looking output that isn't producing anything of value is pseudo-productivity, and in Slow Productivity he lands on doing fewer things, working at a natural pace, and obsessing over quality. He's also argued that the assumed obligation to maintain a constant social presence is worth examining rather than accepting. In his read, the atomization treadmill is a machine for producing volume, and volume is not the same asset as depth.

Where they differ is what the marginal hour buys. Vaynerchuk says another surface. Newport says another draft.

I've landed in a specific spot after running the Vaynerchuk shape for a year. The pyramid is correct and I use it, because cutting derivatives out of one long piece is close to free and skipping it wastes the expensive part. But the volume number attached to it assumes a team, and when one person tries to hit it the derivatives get thin, and thin derivatives don't just underperform, they teach the audience that your stuff is skimmable. So I cap it. One chopped reel per batch on Instagram at most, no matter how many good clips came out of the video, because I'd rather the one be good.

The pyramid is right. The number underneath it is wrong.

Where this breaks

Four failure modes, all of which I've hit.

No prompts. Recording without written questions produces a transcript with no structure, and you end up doing composition work you were trying to avoid.

Skipping the human pass. A draft off a transcript is 70% of a piece. Publishing at 70% is how you end up with a large archive nobody finishes reading.

Cutting a short that stops instead of finishes. A clip needs to resolve the tension it opened, not just end on a strong line. The test I run is whether the thought feels finished or feels like it just stopped, and if it's the second I extend to the resolution even if it breaks the target length.

Front-loading the tools. Every piece of this runs on a phone and one subscription. The video editor and the scheduler are conveniences that make it faster, and buying them first, before you've proven to yourself that you'll do the capture, is the most common way this fails. Nobody has ever been blocked from publishing by not having Descript.

Verified pricing

Read off each vendor's own pricing page, and worth checking, because these move.

ToolFree tierMonth-to-monthBilled annually
Claude ProNo$20/mo$17/mo
Descript Hobbyist1 hr media/mo, 100 one-time credits$24/mo$16/mo
Descript Creatorsame free tier$35/mo$24/mo
Descript Businesssame free tier$65/mo$50/mo
Metricool StarterYes$25/mo$20/mo
Metricool Advanced (automated reporting)Yes$67/mo$53/mo
Local transcriptionFree and unlimited$0$0
Native platform schedulingFree$0$0

Verified as of September 9, 2026.

One thing to watch, because it caught me. Both Descript and Metricool display the annual rate as the big number on the pricing page, so the price you remember seeing is not the price you'll pay if you pay monthly. Descript Creator reads as $24 and bills at $35. Always check which basis you're looking at.

An honest word for doing it the other way

If you have no intention of recording a voice memo every week, this system is worth nothing to you, and I'd rather say that plainly than sell you a workflow you won't run.

Hiring it out is a real option and a defensible one. An agency that ships something every week beats a beautiful system that ships nothing, and if the capture step is never going to happen, pay somebody. What you give up is voice, because the thing that makes this work is that it's actually you talking, and nobody can outsource that convincingly. The middle path most people land on is doing the capture yourself, ten minutes a week, and paying somebody for everything downstream of the transcript.

And if you're posting whenever you feel like it with no system at all, that's not nothing either. It's how most of this starts. The system is what you build when you notice you've said the same thing on eleven sales calls and never written it down once.

When to stop and call someone

Do the first cycle yourself. Not because it's cheap, though it is, but because until you've been through it once you don't know which step is your bottleneck, and everyone's is different.

Places where it's reasonable to get help:

  • You have an archive. Years of recorded calls, webinars, and workshops that nobody has converted to text. That's usually the largest content asset in the building and it's invisible. Processing it is a real project and it produces material for months.
  • You've done three cycles and something isn't landing. Not the tools, the pieces themselves. That's an outside-read problem, not a workflow problem.
  • The publishing surface is the constraint. If your site can't take a blog post without a developer, no amount of production discipline helps.

For that last one, I do a free teardown, on video, no call. There's a field on the form asking what you have so far, and if you've built something and want somebody to look at it, the option that says "something half-built" is the one. If the video is useful and you never speak to me again, that's a completely fine outcome.


Two doors, then. If you'd rather hand the whole thing to someone, that's what I do. If you'd rather run it yourself and have somebody look it over, the teardown is free.

Prices read off each vendor's own pricing page on September 9, 2026. Positions attributed to Gary Vaynerchuk (VaynerMedia, "Document, Don't Create") and Cal Newport (Georgetown, "Deep Work," "Slow Productivity") are my summary of their published work, not quotes.

David Kerns

David Kerns

Operator, builder, creative. Sharing thoughts on the intersection of operations, product, and making things that matter.

More about David

Enjoyed this post?

Subscribe to get new posts delivered to your inbox. No spam.