Recording

What it takes to ship an AI notetaker inside your product

Aug 25, 2026 · 3 min read

An AI notetaker demo takes an afternoon: point a transcription model at an audio file, ask a language model for a summary, done. The product takes months, because the hard part is everything that happens before the audio file exists and after the summary does. Here is the actual pipeline, and where it breaks.

Six stages, five handoffs

A working notetaker is a chain: detect the meeting on the calendar, get a bot into the call, capture the media, transcribe after the call ends, summarize the transcript, and deliver the result back to your application over a webhook. Each handoff is a place where state can be lost. The demo skips the first two stages and the last one, which is why the demo is easy.

It starts in the calendar, not the call

Before anything joins anything, you need to know a meeting exists, that it is virtual, and that this user wants it recorded. That means watching calendar events for join links, handling reschedules and cancellations, and deciding policy: does the bot join everything, only meetings the user created, or only when explicitly asked? Auto-join sounds convenient until the bot shows up in a meeting the user forgot it would attend. An ask-first mode, where the product suggests recording and a human confirms, is the right default for most products.

Reschedules deserve their own attention. A meeting moved twenty minutes is a new join time, and a bot that keeps the original slot records an empty room while the real conversation happens without it. Cancellation is the same problem inverted: a bot that joins a cancelled meeting is a bug your customer's guests get to witness.

The join is where reliability dies

Getting a bot admitted to a call is the least glamorous stage and the most failure-prone. Waiting rooms need a human to admit the bot. Hosts start late. Join links expire or get regenerated. Platform differences matter too: some meeting platforms require the full join URL with its embedded context, while others let you construct entry from a native meeting ID. Treat every join as an operation that can fail loudly, with a status your application can observe, not a fire-and-forget.

  • Join failures: waiting rooms, expired links, hosts who never show. Surface these as observable states, not silence.
  • Duplicate bots: a retry after a timeout can start a second recorder. Check for an active bot first, and reconcile conflicts by adopting the existing one.
  • Platform drift: each meeting platform changes its join flow on its own schedule, and your bot layer has to absorb those changes so your product does not.

Post-call beats live, and it is not close

Live transcription is the feature everyone asks for and the one that hurts most to promise. Live streams drop words and mislabel speakers, and they degrade under exactly the conditions (bad audio, crosstalk) where a transcript matters most. Post-call transcription works from the complete recording, gets a second pass at speaker separation, and produces one authoritative document. Ship post-call as the reliable product. Offer live as a labeled best-effort beta if you offer it at all. Your users will forgive a transcript that arrives five minutes after the call. They will not forgive one that is confidently wrong in the moment.

The webhook is the product

A transcript is worthless sitting in the recorder. Your application needs to hear about it: a signed webhook when the recording completes, another when the transcript and summary are ready, with stable IDs so your handlers can dedupe. Snapshot the completed transcript somewhere your API can serve it long after the call, because the meeting platform will not do that for you.

Where the months go

If you build this yourself, the transcription is the easy fifth. The calendar watcher, the bot fleet, the retry logic, and the webhook plumbing are the other four fifths, and none of them differentiate your product. That slice is what Horato sells as an API: calendar detection through bot join, capture, post-call transcript, summary, and signed webhooks back to your app, so the part you build is the part your users actually see.