DEV Community

Lucie Savoir
Lucie Savoir

Posted on

How to transcribe Zoom, Meet, and Teams without a bot in the participant list

The audio you need to transcribe a Zoom meeting is already playing on your computer. You don't need a bot to fetch it. You don't need the Zoom API. You don't even need permission from the host.

On Mac and Windows, the operating system will hand you that audio directly — through ScreenCaptureKit and WASAPI loopback respectively. Nothing joins the call. Nothing appears in the participant list. The other side sees a normal participant, because that's all there is.

This article explains how to transcribe Zoom meetings without a bot — and do the same for Google Meet and Microsoft Teams — using system audio instead of the Zoom API. CatchMeet, which we built, shows up later as one implementation of the approach, not the headline.

Disclosure: we build CatchMeet. The capture method below is the approach. CatchMeet is one implementation of it.

Why a bot in the participant list is a problem

When a transcription bot joins a meeting, it becomes a participant. That sounds harmless until the practical consequences become clear:
`

Problem What happens in practice
Visibility The bot appears in the roster with its own name and tile
Notifications Host and participants may receive join alerts
Waiting room The bot may never be admitted if the host doesn't approve
Breakout rooms The bot can't follow you in unless the host explicitly allows it
Corporate policies Many organizations block unknown third-party apps
Social friction Participants ask what it is, some object, some leave

`
There's also a practical limitation: bots join as a separate identity. They aren't the participant. They don't have the participant's microphone, audio context, or local playback. They capture a stream the platform gives them — if the platform gives them one at all.

The alternative is to capture the audio on the participant's side. That's what system audio capture does.

The core idea: system audio, not Zoom API

There are three common ways to get meeting audio for transcription:
`

Approach Adds a participant? Works across Zoom, Meet, Teams? Depends on platform approval? Live audio? Cost
Bot Yes Mostly Yes Yes Usually subscription
Zoom API / cloud recording No No (Zoom only) Yes, often enterprise-tier Sometimes Often paid tier
System audio capture No Yes No Yes Free or one-time tool

`
The Zoom API is the official path. It's powerful for cloud recordings, calendar automation, and enterprise compliance. But it requires scopes, approvals, and often a paid plan. It doesn't give a universal live audio stream for every meeting type, and it certainly doesn't help with Google Meet or Microsoft Teams.

System audio is different. It captures whatever the computer is playing — the incoming voices from Zoom, Meet, or Teams — directly from the output device. From the platform's perspective, nothing unusual is happening. It's just a participant with speakers on.
`

Criterion Bot Zoom API System audio
Visible to host Yes No No
Works on Meet / Teams Partially No Yes
Needs platform approval Yes Yes No
Requires paid plan Usually Often No
Per-speaker audio Sometimes Sometimes No (needs diarization)
Works offline / locally No No Yes

`
That's the key insight: a bot isn't necessary, and neither is the Zoom API. What's needed is the audio that's already coming out of the machine.

How system audio capture works on Mac — ScreenCaptureKit

On macOS, the modern way to capture system audio is ScreenCaptureKit. Apple introduced it in macOS 12.3, but the functionality evolved significantly across versions.
`

macOS version System audio capture Per-app audio isolation Notes
12.3–13.x Yes No Captures all system audio, not just the meeting app
14.2+ Yes Yes First version with reliable per-app audio filtering
15.x Yes Yes Privacy alert when bypassing private window picker is normal
16+ (expected) Yes Yes Microphone capture expected in ScreenCaptureKit

`
The minimum practical version: for reliable per-application system audio capture, macOS 14.2 is the realistic floor. On 12.3–14.3, system audio can be captured, but it will be a mix of everything playing — not just the meeting.

Known issues: developers have reported crashes with EXC_BAD_ACCESS during system audio capture on certain configurations. There's also a privacy alert on macOS 15 that appears when an app requests to bypass the system private window picker — this is normal and expected.

What the other side sees: nothing new. The participant is still just a participant. Zoom, Meet, and Teams don't know system audio is being captured, because the capture happens at the OS level, below the meeting app.
`

Aspect What happens
Participant list Unchanged
Host notification None
Recording indicator None (unless built-in recording is on)
Screen share Visible only if the transcription window is shared
Microphone leakage Possible with speakers, avoided with headphones

`
Limitations: the right macOS version, the system audio recording permission, and a stable output device are required. If audio devices are switched mid-call, the capture may need to be restarted. And critically — ScreenCaptureKit does not separate audio by participant. It mixes all remote voices into a single system audio stream.

How system audio capture works on Windows — WASAPI loopback

On Windows, the equivalent is WASAPI loopback. WASAPI (Windows Audio Session API) lets an application capture the audio being sent to an output device without a virtual cable.

In practice:

  • The app enumerates audio output devices.
  • It opens the selected device in loopback mode.
  • It receives the same audio the user hears — including remote participants' voices in Zoom, Meet, or Teams.
  • The transcription engine processes that stream.

What the other side sees: again, nothing. No bot joins. No notification fires.

The hard parts — and how to handle them:
`

Issue Symptom Fix
Exclusive mode Loopback captures silence Disable exclusive mode in Sound Settings → Device Properties → Advanced
Sample rate mismatch Stream drops or distorts Resample upstream, or use a virtual audio cable as fallback
Wrong process targeted Zero bytes captured Hook the Media Foundation process tree, not the UI shell
Device switch mid-call Capture stops Restart capture, or lock the output device

`
Details on each:

  • Exclusive mode. If Teams or another app has exclusive control of the audio device, WASAPI loopback will fail to capture. The fix is to disable exclusive mode in Sound Settings → Device Properties → Advanced, or use per-process loopback to target the meeting app directly.
  • Sample rate mismatch. If the capture client is set to 44.1 kHz but Teams forces the hardware to 48 kHz, the engine may drop the connection. Modern libraries handle some of this automatically — NAudio 3, for example, now adapts bit depth and channels in exclusive mode, leaving only sample-rate mismatches for upstream resampling. In stubborn cases, a virtual audio cable remains the industry-standard fallback.
  • Process targeting. Teams spawns multiple processes. Hooking the UI shell process instead of the Media Foundation process yields zero bytes. Since Windows 10 build 19041, per-process loopback with PROCESS_LOOPBACK_MODE_INCLUDE_TARGET_PROCESS_TREE captures the entire process tree.

`

Windows version Basic loopback Per-process loopback Notes
Windows 10 (pre-2004) Yes No Exclusive-mode issues may appear
Windows 10 build 19041+ Yes Yes Recommended baseline
Windows 11 Yes Yes Full support

`
Limitations: the output device must be active. Exclusive-mode audio can block loopback. And like ScreenCaptureKit, WASAPI loopback gives a mixed stream — not per-speaker audio.

The diarization problem (and what to do about it)

This is the part most articles skip. System audio capture gives a single mixed stream of all remote participants. It does not indicate who said what.
`

Source of audio Per-speaker separation Metadata available
Platform API (Zoom, Teams) Sometimes Names, join times, speaker events
Bot Sometimes Participant list
System audio capture No None — audio waveform only

`
Why it's harder: platform APIs often provide participant metadata — names, join times, speaker events. System audio capture gives none of that. It provides the audio waveform and nothing else.

What can be done:
`

Approach How it works Trade-off
Built-in diarization AssemblyAI, Deepgram, Rev analyze voice characteristics Costs extra, accuracy varies
Meeting metadata mapping Map diarized labels to participant list manually Manual effort
Accept the limitation Use "Speaker 1", "Speaker 2" labels No names

`

What doesn't work: expecting the OS to magically separate speakers. Neither ScreenCaptureKit nor WASAPI loopback does this. A separate layer is required.

What the other side sees

This is the part most people care about, so precision matters.

When system audio capture is used for transcription:
`

What Visible to other side?
Bot in participant list No
Bot join notification No
Platform recording indicator Only if built-in recording is on
Transcription window Only if screen is shared
Audio leakage via microphone Only if speakers are used


But here are the scenarios people actually search for:

Scenario What happens
Host enables recording Platform indicator appears; local capture unaffected
Corporate DLP enabled Native DLP can't inspect live audio; endpoint monitoring may detect capture
Zoom audio watermarking on Watermark in played audio; can identify recorder if audio is shared externally
Participant complains Legal question, not technical — see consent section

`
What if the host enables recording? Then Zoom, Meet, or Teams will show a recording indicator to everyone. That indicator is about the platform's own recording, not local capture. Transcription continues unaffected.

What if the meeting has DLP enabled? Enterprise DLP tools in Teams, Slack, and Zoom are built for chat and file content, not live audio streams. Native DLP typically cannot inspect voice or screen-share content. That said, some organizations deploy endpoint monitoring that can detect unusual audio capture. Company policy should be checked.

What if Zoom has audio watermarking enabled? Zoom can embed an inaudible watermark in the audio played through each participant's speakers. If someone records the meeting with a separate microphone or third-party tool and shares the audio, Zoom can potentially identify which participant was responsible. Local and cloud recordings made directly through Zoom do not contain this watermark.

What if someone complains? That's a legal question, not a technical one. See below.

What can give it away: sharing the screen with the transcription window visible, or using speakers so the microphone picks up the transcription audio. Headphones avoid the latter.

Consent: the part that can't be skipped

System audio capture is invisible to the platform. But invisibility is not the same as legality.
`

Jurisdiction Consent rule Examples
US federal One-party consent —
All-party consent states All parties must consent California, Florida, Pennsylvania, Washington, Illinois
EU / GDPR Lawful basis required Consent or legitimate interest
Corporate policy Varies by organization Many prohibit third-party transcription

`
In the United States, federal law operates on a one-party consent basis — it's legal to record if at least one person in the conversation is aware and consents. But over a dozen states, including California, Florida, and Pennsylvania, require all-party consent. California's Invasion of Privacy Act is particularly broad: it prohibits not just recording, but also "reading, attempting to read, or learning" the contents of communications without all parties' consent.

Recent lawsuits against AI notetakers like Otter.ai have focused on exactly this gap. Plaintiffs allege that Otter's notetaker joins meetings, records conversations including those of non-users, and uses the data to train machine learning models — all without proper consent or disclosure.
`

Situation Recommended action
Single-party call, one-party state Consent from one participant is enough
Group call, all-party state Inform and get consent from all participants
Mixed jurisdictions Apply the strictest rule
Corporate meeting Check company policy first
Public webinar Follow platform and organizer rules

`
What this means in practice: if transcribing a meeting with participants in all-party consent states, or in multiple jurisdictions, the strictest rule applies. Participants should be told. Their permission should be obtained. The absence of a bot doesn't remove that obligation.

Corporate policies: some organizations explicitly prohibit third-party transcription tools. One policy states: "Personal AI transcribers or any other third-party transcription tools are strictly prohibited during group calls." Company rules should be checked before capturing anything.

Step-by-step: transcribe Zoom, Meet, and Teams without a bot

Before starting

`

Step Why it matters
Get consent Legal requirement in many jurisdictions
Choose language and model Affects accuracy
Decide where to save Local storage is best for privacy
Use headphones Prevents microphone picking up system audio

`

Mac — ScreenCaptureKit

`

Step Action
1 Verify macOS 14.2 or later
2 Open the transcription tool
3 Grant Screen Recording permission
4 Select display or application to capture
5 Ensure system audio is included
6 Start capture, then transcription
7 Join the meeting; confirm no bot appears

`

Windows — WASAPI loopback

`

Step Action
1 Disable exclusive mode in Sound Settings → Advanced
2 Open the transcription tool
3 Select the output device playing the meeting
4 Enable loopback capture
5 Target the meeting app's process tree (build 19041+)
6 Start capture, then transcription
7 Join the meeting; confirm participant list unchanged

`

Zoom, Google Meet, Microsoft Teams

`

Aspect Zoom Google Meet Microsoft Teams
System audio capture works Yes Yes Yes
Bot join notification Yes Yes Yes
Breakout rooms Captured if you're in one Captured if you're in one Captured if you're in one
Corporate restrictions Common Common Common
Built-in recording indicator Yes Yes Yes

`

  • What's common: all three play remote audio through the system output. System audio capture works identically.
  • Breakout rooms: system audio captures whatever is heard. In a breakout room, that room's audio is captured. In the main room, breakout rooms aren't heard — and neither is the capture.
  • Corporate accounts: some organizations restrict third-party apps. System audio capture is unaffected, since it runs locally — but company policy may still prohibit it.

CatchMeet as one implementation

CatchMeet is a practical example of how this approach comes together. It uses ScreenCaptureKit on Mac and WASAPI loopback on Windows to capture system audio. It transcribes locally, without joining the meeting as a participant, and without the Zoom API.
`

Feature CatchMeet
Joins as participant No
Uses Zoom API No
Mac capture method ScreenCaptureKit
Windows capture method WASAPI loopback
Works on Meet and Teams Yes
Requires host approval No

`
What CatchMeet doesn't do: it doesn't connect to Zoom, Meet, or Teams as a bot. It doesn't appear in the participant list. It doesn't require host approval. It's one way to implement the system-audio approach — not the only way.

For those who'd rather not assemble the pieces themselves, CatchMeet is a reasonable starting point.

Limitations and when a bot or API is still better

System audio capture is powerful, but it isn't universal.
`

Scenario Best approach
Local meeting with human present System audio capture
Server-side recording only Platform API or cloud recording
Fully automated pipeline, no human Bot
Per-speaker audio required API or diarization engine
Enterprise compliance workflow Zoom API
Cross-platform, invisible capture System audio capture

`
When it doesn't fit:

  • Meetings where audio isn't played locally (server-side recording, fully automated pipelines with no human present).
  • Scenarios where the computer is off or the meeting runs on a device that isn't controlled.
  • Situations requiring per-speaker audio without additional diarization work.

When the Zoom API is better:

  • Enterprise compliance and cloud recording retrieval.
  • Calendar-driven automation.
  • Workflows that need official platform data.

When a bot is better:

  • Large-scale automated meeting capture where a visible participant is acceptable.
  • Platforms or configurations where the API doesn't provide the audio needed.

Trade-offs of system audio: requires an active machine, depends on OS permissions, diarization is harder, and consent compliance is the user's responsibility.

FAQ

How do I transcribe Zoom meetings without a bot?
Capture system audio on the local machine using ScreenCaptureKit on Mac or WASAPI loopback on Windows. The transcription happens locally, and no bot joins the meeting.

Will the other side know I'm transcribing?
Not from the participant list or a bot notification. But if the law requires informing participants, do so. Using headphones also prevents audio leakage through the microphone.

Does system audio capture work with breakout rooms?
Yes — whatever audio plays through the speakers or headphones is captured. In a breakout room, that room's audio is captured. In the main room, breakout rooms aren't heard, and neither is the capture. A bot can't follow a participant into a breakout room unless the host explicitly allows it; system audio capture follows automatically.

What happens if the host enables recording?
The platform shows a recording indicator to everyone. That indicator is about the platform's built-in recording, not local capture. Transcription continues unaffected.

Does corporate DLP detect system audio capture?
Native DLP in Teams, Slack, and Zoom is built for chat and file content, not live audio streams. It typically cannot inspect voice or screen-share content. However, some organizations deploy endpoint monitoring that may detect unusual audio capture. Company policy should be checked.

Can I transcribe Google Meet and Microsoft Teams the same way?
Yes. System audio capture is platform-agnostic. If the audio plays through the output device, it can be captured.

Does CatchMeet join as a participant?
No. It captures system audio locally and does not appear in the participant list.

Is it legal to transcribe a meeting without telling participants?
It depends on the jurisdiction. Federal law in the US operates on one-party consent, but over a dozen states require all-party consent. California's CIPA is particularly broad. When in doubt, tell people and get permission.

What about diarization — who said what?
System audio gives a mixed stream. A transcription engine with built-in diarization (AssemblyAI, Deepgram, Rev) is needed to separate speakers. Even then, mapping speakers to names requires additional work.

Does this work on macOS 12 or 13?
System audio capture works from macOS 12.3, but reliable application-specific capture requires macOS 14.2 or later. On older versions, all system audio is captured, not just the meeting.

Does this work on Windows 10?
Per-process loopback requires Windows 10 build 19041 (version 2004) or later. Basic WASAPI loopback works on earlier versions, but exclusive-mode or sample-rate issues may appear.

Conclusion

A bot in the participant list isn't necessary to transcribe a meeting. The Zoom API isn't necessary either. What's needed is system audio — the audio the computer is already playing — captured at the OS level.
`

Platform Method Minimum version
macOS ScreenCaptureKit 14.2 for per-app audio
Windows WASAPI loopback Build 19041 for per-process
Zoom, Meet, Teams All supported —

`
On Mac, that's ScreenCaptureKit (macOS 14.2+ for reliable app-specific capture). On Windows, that's WASAPI loopback (with per-process targeting on build 19041+). The other side sees a normal participant, because that's all there is.

CatchMeet is one implementation of this approach. For anyone looking to transcribe Zoom meetings without a bot, it's a reasonable place to start.

Top comments (3)

Collapse
 
iberezh profile image
Ilya Berezhniak •

We've been using Otter for 6 months and the bot thing is killing our sales calls. Prospects literally ask "is that recording us?" and the tone changes immediately. Does the system audio approach actually fix this? Like, does the client see absolutely nothing different on their end?

Collapse
 
vedernikova profile image
Ria •

Curious about the WASAPI loopback approach - does per-process capture actually work with Teams? Last time I tried, Teams spawned like 5 processes and I couldn't figure out which one to hook.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.