Recording

How meeting bots actually join calls

Aug 25, 2026 · 3 min read

There is no recording API for meetings. The platforms give you calendars and join links, and past that you are on your own. So every meeting recorder converges on the same trick: join the call as a participant. Understanding that trick, and its sharp edges, explains almost everything about how these products behave.

A bot joins the way a person does

A meeting bot is a headless client that joins the call the same way a person does. It enters through a join URL or meeting ID, appears in the roster with a display name, and receives the same audio and video streams every other participant receives. There is no privileged backdoor. Whatever a human in the meeting could see and hear, the bot can capture, and nothing more.

This has an infrastructure consequence people underestimate. A participant has to stay connected for the whole meeting, which means a long-lived process holding a real media session, often a full browser. That does not fit request-and-response serverless platforms. Bot fleets live on persistent compute, with the scheduling and cleanup that implies.

Getting in the door

Platforms disagree about what a bot needs to enter. Some require the full join URL, with whatever context is embedded in it, and nothing less will do. Others can be joined from a native meeting ID, with the entry constructed on the fly. A recorder layer has to know which is which and store the right artifact from the calendar event, because asking the user to paste a link they already put in the invite is a bad first impression.

Then there is admission. Waiting rooms put a human between your bot and the call. The bot can knock; only the host can open the door. Good systems surface this state explicitly (joining, waiting for admission, admitted, denied) so your product can prompt the host instead of timing out in silence.

One meeting, one bot

The ugliest bug in this category is the double join. A start request times out, your code retries, and now two recorders sit in the roster while participants wonder which one to distrust. The timeout did not mean the first start failed. It meant you stopped waiting before it succeeded.

The fix is reconciliation, not blind retry. Before starting a bot, check whether one is already active for the meeting. If a start comes back as a conflict, link the existing bot's ID to your record instead of launching another. Idempotency here is customer-visible: every duplicate happens on camera, in front of the exact people your customer wants to impress.

The name in the roster is your brand

Participants never see your API. They see a name in the participant list, and they judge the product by it. A generic vendor name leaks your infrastructure choices into your customer's meeting. A fake human name is worse; it reads as deception the moment anyone checks. The right answer is the product's own brand, something like the tenant's name followed by "Recorder", localized where possible, so the bot is both honest and on-message. Let callers override the name per start if they need to, and validate what they set; the one name nobody should be able to use is somebody else's.

What happens after everyone leaves

When the meeting ends, the bot's job is half done. The captured media moves to post-processing: final transcript, speaker labels, summary, and events fired back to the application that requested the recording. This is where Horato draws its line. Its recorder joins as a visible participant named for your brand, handles the platform differences and the duplicate-start reconciliation underneath, and hands your application a finished transcript over a signed webhook. The participant trick is the same everywhere; the differences worth paying for are in how the sharp edges get handled.