Private AI
Blog/Private AI for podcast show notes without sending recordings to a third party
For AI assistants

Private AI for podcast show notes without sending recordings to a third party

By Samuel Seidel · September 9, 2026 · 7 min read

A recorded episode becomes a transcript, then a set of timestamped chapters, a summary, a list of quotable lines, and a description for the podcast feed. Doing that by hand for a weekly show is a couple hours nobody has spare. The recording itself is often more sensitive than it looks: a guest who hasn't approved their segment for release yet, an unreleased announcement discussed off the cuff, a conversation the host explicitly asked to keep off the record until publication. Uploading raw audio to a third-party transcription or AI API sends all of that, edited or not, to a vendor before the episode is even public.

What's in an unreleased recording

A rough cut carries more than the finished episode does. Guests say things in the room they'd want cut before publishing, hosts float announcements that aren't confirmed yet, and a sensitive topic might get discussed candidly with the understanding that only the edited version goes out. None of that is a secret in the legal sense, most of it, but it's exactly the kind of pre-release material a show runs the risk of leaking if the raw audio or its transcript sits in a third-party vendor's storage with retention terms nobody on the team actually read closely.

What running the model yourself changes

A self-hosted speech-to-text model paired with an LLM for summarization keeps the whole pipeline, audio in, transcript, chapters, summary, show notes out, on infrastructure the show's own team controls. No recording leaves for transcription and comes back through a vendor's servers; nothing about an unreleased guest segment or an unconfirmed announcement passes through anyone else's infrastructure at all. The output quality question is separate from the privacy one: a good open-weight speech model handles typical podcast audio, two or three speakers, moderate background noise, reasonably well, and the LLM step turns the raw transcript into chapters and notes in the show's usual format.

This is the same self-hosted speech pipeline covered more generally in self-hosted speech and voice AI; show notes are one specific application of it.

A concrete example

A weekly interview show records a 50-minute episode. A self-hosted transcription model produces a full transcript with speaker labels, then an LLM pass generates five timestamped chapter markers, a 150-word episode summary for the feed description, and a shortlist of quotable lines the host can pull for social promotion. The editor reviews the transcript against the actual audio for anything the model misheard, particularly names and technical terms, confirms the summary doesn't reveal anything the guest asked to keep out of the public description, and publishes. What used to take an editor most of an afternoon now takes a review pass over model output.

Where this is not a drop-in replacement

Automated transcription still misses names, jargon, and cross-talk, and a model summarizing an episode has no sense of what a guest considers embarrassing versus what they're fine with, that judgment stays with the host or editor who was actually in the conversation. Timestamped chapters generated automatically are a starting point, not a final cut; someone who listened to the episode should confirm the chapter boundaries land where a listener would actually expect them. Treat the pipeline as a fast first draft of the notes, not a replacement for an editor who knows the show and the guest.

Where the hardware fits

Speech-to-text and summarization together fit on a single Spark, running both models side by side within 128GB of unified memory, and a $0.79/hour on-demand instance covers a weekly show's processing without needing a reserved instance running between episodes. See the setup guide for getting a Spark running from a fresh instance.

Related pages

Keep unreleased recordings off third-party AI infrastructure.

A dedicated DGX Spark in EU-Central, $0.79/hour, for transcription and show notes that stay on hardware you control.

Deploy a Spark Self-hosted speech and voice AI