The audio you need to transcribe a Zoom meeting is already playing on your computer. You don't need a bot to fetch it. You don't need the Zoom API. You don't even need permission from the host.
On Mac and Windows, the operating system will hand you that audio directly — through ScreenCaptureKit and WASAPI loopback respectively. Nothing joins the call. Nothing appears in the participant list. The other side sees a normal participant, because that's all there is.
This article explains how to transcribe Zoom meetings without a bot — and do the same for Google Meet and Microsoft Teams — using system audio instead of the Zoom API. CatchMeet, which we built, shows up later as one implementation of the approach, not the headline.
Disclosure: we build CatchMeet. The capture method below is the approach. CatchMeet is one implementation of it.
Why a bot in the participant list is a problem
When a transcription bot joins a meeting, it becomes a participant. That sounds harmless until the practical consequences become clear:
`
| Problem | What happens in practice |
|---|---|
| Visibility | The bot appears in the roster with its own name and tile |
| Notifications | Host and participants may receive join alerts |
| Waiting room | The bot may never be admitted if the host doesn't approve |
| Breakout rooms | The bot can't follow you in unless the host explicitly allows it |
| Corporate policies | Many organizations block unknown third-party apps |
| Social friction | Participants ask what it is, some object, some leave |
`
There's also a practical limitation: bots join as a separate identity. They aren't the participant. They don't have the participant's microphone, audio context, or local playback. They capture a stream the platform gives them — if the platform gives them one at all.
The alternative is to capture the audio on the participant's side. That's what system audio capture does.
The core idea: system audio, not Zoom API
There are three common ways to get meeting audio for transcription:
`
| Approach | Adds a participant? | Works across Zoom, Meet, Teams? | Depends on platform approval? | Live audio? | Cost |
|---|---|---|---|---|---|
| Bot | Yes | Mostly | Yes | Yes | Usually subscription |
| Zoom API / cloud recording | No | No (Zoom only) | Yes, often enterprise-tier | Sometimes | Often paid tier |
| System audio capture | No | Yes | No | Yes | Free or one-time tool |
`
The Zoom API is the official path. It's powerful for cloud recordings, calendar automation, and enterprise compliance. But it requires scopes, approvals, and often a paid plan. It doesn't give a universal live audio stream for every meeting type, and it certainly doesn't help with Google Meet or Microsoft Teams.
System audio is different. It captures whatever the computer is playing — the incoming voices from Zoom, Meet, or Teams — directly from the output device. From the platform's perspective, nothing unusual is happening. It's just a participant with speakers on.
`
| Criterion | Bot | Zoom API | System audio |
|---|---|---|---|
| Visible to host | Yes | No | No |
| Works on Meet / Teams | Partially | No | Yes |
| Needs platform approval | Yes | Yes | No |
| Requires paid plan | Usually | Often | No |
| Per-speaker audio | Sometimes | Sometimes | No (needs diarization) |
| Works offline / locally | No | No | Yes |
`
That's the key insight: a bot isn't necessary, and neither is the Zoom API. What's needed is the audio that's already coming out of the machine.
How system audio capture works on Mac — ScreenCaptureKit
On macOS, the modern way to capture system audio is ScreenCaptureKit. Apple introduced it in macOS 12.3, but the functionality evolved significantly across versions.
`
| macOS version | System audio capture | Per-app audio isolation | Notes |
|---|---|---|---|
| 12.3–13.x | Yes | No | Captures all system audio, not just the meeting app |
| 14.2+ | Yes | Yes | First version with reliable per-app audio filtering |
| 15.x | Yes | Yes | Privacy alert when bypassing private window picker is normal |
| 16+ (expected) | Yes | Yes | Microphone capture expected in ScreenCaptureKit |
`
The minimum practical version: for reliable per-application system audio capture, macOS 14.2 is the realistic floor. On 12.3–14.3, system audio can be captured, but it will be a mix of everything playing — not just the meeting.
Known issues: developers have reported crashes with EXC_BAD_ACCESS during system audio capture on certain configurations. There's also a privacy alert on macOS 15 that appears when an app requests to bypass the system private window picker — this is normal and expected.
What the other side sees: nothing new. The participant is still just a participant. Zoom, Meet, and Teams don't know system audio is being captured, because the capture happens at the OS level, below the meeting app.
`
| Aspect | What happens |
|---|---|
| Participant list | Unchanged |
| Host notification | None |
| Recording indicator | None (unless built-in recording is on) |
| Screen share | Visible only if the transcription window is shared |
| Microphone leakage | Possible with speakers, avoided with headphones |
`
Limitations: the right macOS version, the system audio recording permission, and a stable output device are required. If audio devices are switched mid-call, the capture may need to be restarted. And critically — ScreenCaptureKit does not separate audio by participant. It mixes all remote voices into a single system audio stream.
How system audio capture works on Windows — WASAPI loopback
On Windows, the equivalent is WASAPI loopback. WASAPI (Windows Audio Session API) lets an application capture the audio being sent to an output device without a virtual cable.
In practice:
- The app enumerates audio output devices.
- It opens the selected device in loopback mode.
- It receives the same audio the user hears — including remote participants' voices in Zoom, Meet, or Teams.
- The transcription engine processes that stream.
What the other side sees: again, nothing. No bot joins. No notification fires.
The hard parts — and how to handle them:
`
| Issue | Symptom | Fix |
|---|---|---|
| Exclusive mode | Loopback captures silence | Disable exclusive mode in Sound Settings → Device Properties → Advanced |
| Sample rate mismatch | Stream drops or distorts | Resample upstream, or use a virtual audio cable as fallback |
| Wrong process targeted | Zero bytes captured | Hook the Media Foundation process tree, not the UI shell |
| Device switch mid-call | Capture stops | Restart capture, or lock the output device |
`
Details on each:
- Exclusive mode. If Teams or another app has exclusive control of the audio device, WASAPI loopback will fail to capture. The fix is to disable exclusive mode in Sound Settings → Device Properties → Advanced, or use per-process loopback to target the meeting app directly.
- Sample rate mismatch. If the capture client is set to 44.1 kHz but Teams forces the hardware to 48 kHz, the engine may drop the connection. Modern libraries handle some of this automatically — NAudio 3, for example, now adapts bit depth and channels in exclusive mode, leaving only sample-rate mismatches for upstream resampling. In stubborn cases, a virtual audio cable remains the industry-standard fallback.
-
Process targeting. Teams spawns multiple processes. Hooking the UI shell process instead of the Media Foundation process yields zero bytes. Since Windows 10 build 19041, per-process loopback with
PROCESS_LOOPBACK_MODE_INCLUDE_TARGET_PROCESS_TREEcaptures the entire process tree.
`
| Windows version | Basic loopback | Per-process loopback | Notes |
|---|---|---|---|
| Windows 10 (pre-2004) | Yes | No | Exclusive-mode issues may appear |
| Windows 10 build 19041+ | Yes | Yes | Recommended baseline |
| Windows 11 | Yes | Yes | Full support |
`
Limitations: the output device must be active. Exclusive-mode audio can block loopback. And like ScreenCaptureKit, WASAPI loopback gives a mixed stream — not per-speaker audio.
The diarization problem (and what to do about it)
This is the part most articles skip. System audio capture gives a single mixed stream of all remote participants. It does not indicate who said what.
`
| Source of audio | Per-speaker separation | Metadata available |
|---|---|---|
| Platform API (Zoom, Teams) | Sometimes | Names, join times, speaker events |
| Bot | Sometimes | Participant list |
| System audio capture | No | None — audio waveform only |
`
Why it's harder: platform APIs often provide participant metadata — names, join times, speaker events. System audio capture gives none of that. It provides the audio waveform and nothing else.
What can be done:
`
| Approach | How it works | Trade-off |
|---|---|---|
| Built-in diarization | AssemblyAI, Deepgram, Rev analyze voice characteristics | Costs extra, accuracy varies |
| Meeting metadata mapping | Map diarized labels to participant list manually | Manual effort |
| Accept the limitation | Use "Speaker 1", "Speaker 2" labels | No names |
`
What doesn't work: expecting the OS to magically separate speakers. Neither ScreenCaptureKit nor WASAPI loopback does this. A separate layer is required.
What the other side sees
This is the part most people care about, so precision matters.
When system audio capture is used for transcription:
`
| What | Visible to other side? |
|---|---|
| Bot in participant list | No |
| Bot join notification | No |
| Platform recording indicator | Only if built-in recording is on |
| Transcription window | Only if screen is shared |
| Audio leakage via microphone | Only if speakers are used |
But here are the scenarios people actually search for:
| Scenario | What happens |
|---|---|
| Host enables recording | Platform indicator appears; local capture unaffected |
| Corporate DLP enabled | Native DLP can't inspect live audio; endpoint monitoring may detect capture |
| Zoom audio watermarking on | Watermark in played audio; can identify recorder if audio is shared externally |
| Participant complains | Legal question, not technical — see consent section |
`
What if the host enables recording? Then Zoom, Meet, or Teams will show a recording indicator to everyone. That indicator is about the platform's own recording, not local capture. Transcription continues unaffected.
What if the meeting has DLP enabled? Enterprise DLP tools in Teams, Slack, and Zoom are built for chat and file content, not live audio streams. Native DLP typically cannot inspect voice or screen-share content. That said, some organizations deploy endpoint monitoring that can detect unusual audio capture. Company policy should be checked.
What if Zoom has audio watermarking enabled? Zoom can embed an inaudible watermark in the audio played through each participant's speakers. If someone records the meeting with a separate microphone or third-party tool and shares the audio, Zoom can potentially identify which participant was responsible. Local and cloud recordings made directly through Zoom do not contain this watermark.
What if someone complains? That's a legal question, not a technical one. See below.
What can give it away: sharing the screen with the transcription window visible, or using speakers so the microphone picks up the transcription audio. Headphones avoid the latter.
Consent: the part that can't be skipped
System audio capture is invisible to the platform. But invisibility is not the same as legality.
`
| Jurisdiction | Consent rule | Examples |
|---|---|---|
| US federal | One-party consent | — |
| All-party consent states | All parties must consent | California, Florida, Pennsylvania, Washington, Illinois |
| EU / GDPR | Lawful basis required | Consent or legitimate interest |
| Corporate policy | Varies by organization | Many prohibit third-party transcription |
`
In the United States, federal law operates on a one-party consent basis — it's legal to record if at least one person in the conversation is aware and consents. But over a dozen states, including California, Florida, and Pennsylvania, require all-party consent. California's Invasion of Privacy Act is particularly broad: it prohibits not just recording, but also "reading, attempting to read, or learning" the contents of communications without all parties' consent.
Recent lawsuits against AI notetakers like Otter.ai have focused on exactly this gap. Plaintiffs allege that Otter's notetaker joins meetings, records conversations including those of non-users, and uses the data to train machine learning models — all without proper consent or disclosure.
`
| Situation | Recommended action |
|---|---|
| Single-party call, one-party state | Consent from one participant is enough |
| Group call, all-party state | Inform and get consent from all participants |
| Mixed jurisdictions | Apply the strictest rule |
| Corporate meeting | Check company policy first |
| Public webinar | Follow platform and organizer rules |
`
What this means in practice: if transcribing a meeting with participants in all-party consent states, or in multiple jurisdictions, the strictest rule applies. Participants should be told. Their permission should be obtained. The absence of a bot doesn't remove that obligation.
Corporate policies: some organizations explicitly prohibit third-party transcription tools. One policy states: "Personal AI transcribers or any other third-party transcription tools are strictly prohibited during group calls." Company rules should be checked before capturing anything.
Step-by-step: transcribe Zoom, Meet, and Teams without a bot
Before starting
`
| Step | Why it matters |
|---|---|
| Get consent | Legal requirement in many jurisdictions |
| Choose language and model | Affects accuracy |
| Decide where to save | Local storage is best for privacy |
| Use headphones | Prevents microphone picking up system audio |
`
Mac — ScreenCaptureKit
`
| Step | Action |
|---|---|
| 1 | Verify macOS 14.2 or later |
| 2 | Open the transcription tool |
| 3 | Grant Screen Recording permission |
| 4 | Select display or application to capture |
| 5 | Ensure system audio is included |
| 6 | Start capture, then transcription |
| 7 | Join the meeting; confirm no bot appears |
`
Windows — WASAPI loopback
`
| Step | Action |
|---|---|
| 1 | Disable exclusive mode in Sound Settings → Advanced |
| 2 | Open the transcription tool |
| 3 | Select the output device playing the meeting |
| 4 | Enable loopback capture |
| 5 | Target the meeting app's process tree (build 19041+) |
| 6 | Start capture, then transcription |
| 7 | Join the meeting; confirm participant list unchanged |
`
Zoom, Google Meet, Microsoft Teams
`
| Aspect | Zoom | Google Meet | Microsoft Teams |
|---|---|---|---|
| System audio capture works | Yes | Yes | Yes |
| Bot join notification | Yes | Yes | Yes |
| Breakout rooms | Captured if you're in one | Captured if you're in one | Captured if you're in one |
| Corporate restrictions | Common | Common | Common |
| Built-in recording indicator | Yes | Yes | Yes |
`
- What's common: all three play remote audio through the system output. System audio capture works identically.
- Breakout rooms: system audio captures whatever is heard. In a breakout room, that room's audio is captured. In the main room, breakout rooms aren't heard — and neither is the capture.
- Corporate accounts: some organizations restrict third-party apps. System audio capture is unaffected, since it runs locally — but company policy may still prohibit it.
CatchMeet as one implementation
CatchMeet is a practical example of how this approach comes together. It uses ScreenCaptureKit on Mac and WASAPI loopback on Windows to capture system audio. It transcribes locally, without joining the meeting as a participant, and without the Zoom API.
`
| Feature | CatchMeet |
|---|---|
| Joins as participant | No |
| Uses Zoom API | No |
| Mac capture method | ScreenCaptureKit |
| Windows capture method | WASAPI loopback |
| Works on Meet and Teams | Yes |
| Requires host approval | No |
`
What CatchMeet doesn't do: it doesn't connect to Zoom, Meet, or Teams as a bot. It doesn't appear in the participant list. It doesn't require host approval. It's one way to implement the system-audio approach — not the only way.
For those who'd rather not assemble the pieces themselves, CatchMeet is a reasonable starting point.
Limitations and when a bot or API is still better
System audio capture is powerful, but it isn't universal.
`
| Scenario | Best approach |
|---|---|
| Local meeting with human present | System audio capture |
| Server-side recording only | Platform API or cloud recording |
| Fully automated pipeline, no human | Bot |
| Per-speaker audio required | API or diarization engine |
| Enterprise compliance workflow | Zoom API |
| Cross-platform, invisible capture | System audio capture |
`
When it doesn't fit:
- Meetings where audio isn't played locally (server-side recording, fully automated pipelines with no human present).
- Scenarios where the computer is off or the meeting runs on a device that isn't controlled.
- Situations requiring per-speaker audio without additional diarization work.
When the Zoom API is better:
- Enterprise compliance and cloud recording retrieval.
- Calendar-driven automation.
- Workflows that need official platform data.
When a bot is better:
- Large-scale automated meeting capture where a visible participant is acceptable.
- Platforms or configurations where the API doesn't provide the audio needed.
Trade-offs of system audio: requires an active machine, depends on OS permissions, diarization is harder, and consent compliance is the user's responsibility.
FAQ
How do I transcribe Zoom meetings without a bot?
Capture system audio on the local machine using ScreenCaptureKit on Mac or WASAPI loopback on Windows. The transcription happens locally, and no bot joins the meeting.
Will the other side know I'm transcribing?
Not from the participant list or a bot notification. But if the law requires informing participants, do so. Using headphones also prevents audio leakage through the microphone.
Does system audio capture work with breakout rooms?
Yes — whatever audio plays through the speakers or headphones is captured. In a breakout room, that room's audio is captured. In the main room, breakout rooms aren't heard, and neither is the capture. A bot can't follow a participant into a breakout room unless the host explicitly allows it; system audio capture follows automatically.
What happens if the host enables recording?
The platform shows a recording indicator to everyone. That indicator is about the platform's built-in recording, not local capture. Transcription continues unaffected.
Does corporate DLP detect system audio capture?
Native DLP in Teams, Slack, and Zoom is built for chat and file content, not live audio streams. It typically cannot inspect voice or screen-share content. However, some organizations deploy endpoint monitoring that may detect unusual audio capture. Company policy should be checked.
Can I transcribe Google Meet and Microsoft Teams the same way?
Yes. System audio capture is platform-agnostic. If the audio plays through the output device, it can be captured.
Does CatchMeet join as a participant?
No. It captures system audio locally and does not appear in the participant list.
Is it legal to transcribe a meeting without telling participants?
It depends on the jurisdiction. Federal law in the US operates on one-party consent, but over a dozen states require all-party consent. California's CIPA is particularly broad. When in doubt, tell people and get permission.
What about diarization — who said what?
System audio gives a mixed stream. A transcription engine with built-in diarization (AssemblyAI, Deepgram, Rev) is needed to separate speakers. Even then, mapping speakers to names requires additional work.
Does this work on macOS 12 or 13?
System audio capture works from macOS 12.3, but reliable application-specific capture requires macOS 14.2 or later. On older versions, all system audio is captured, not just the meeting.
Does this work on Windows 10?
Per-process loopback requires Windows 10 build 19041 (version 2004) or later. Basic WASAPI loopback works on earlier versions, but exclusive-mode or sample-rate issues may appear.
Conclusion
A bot in the participant list isn't necessary to transcribe a meeting. The Zoom API isn't necessary either. What's needed is system audio — the audio the computer is already playing — captured at the OS level.
`
| Platform | Method | Minimum version |
|---|---|---|
| macOS | ScreenCaptureKit | 14.2 for per-app audio |
| Windows | WASAPI loopback | Build 19041 for per-process |
| Zoom, Meet, Teams | All supported | — |
`
On Mac, that's ScreenCaptureKit (macOS 14.2+ for reliable app-specific capture). On Windows, that's WASAPI loopback (with per-process targeting on build 19041+). The other side sees a normal participant, because that's all there is.
CatchMeet is one implementation of this approach. For anyone looking to transcribe Zoom meetings without a bot, it's a reasonable place to start.
Top comments (3)
We've been using Otter for 6 months and the bot thing is killing our sales calls. Prospects literally ask "is that recording us?" and the tone changes immediately. Does the system audio approach actually fix this? Like, does the client see absolutely nothing different on their end?
Curious about the WASAPI loopback approach - does per-process capture actually work with Teams? Last time I tried, Teams spawned like 5 processes and I couldn't figure out which one to hook.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.