Speaker labels, computed on your device
A wall of text is hard to read and easy to misquote. JesRecap breaks a transcript into turns and marks which ones are yours — using a small local model, with no audio sent anywhere.
What speaker diarization means, in plain words
Transcription answers "what was said". Diarization answers the other half of the question: who was talking, and when. It is the difference between a paragraph of run-together sentences and a readable exchange of turns.
The model does not understand words for this. It listens to voice characteristics — pitch, timbre, the physical texture of a voice — and turns each stretch of speech into a numeric fingerprint. Fingerprints that look alike are grouped together, and each group becomes a speaker. The transcript is then split at the points where the voice changes hands.
Two things follow from that. It works on languages it was never told about, because voices are voices. And it does not learn identities — grouping voices is not the same as knowing whose they are.
How it works in JesRecap
Optional, local, and applied after the recording is finished.
A one-time 54 MB model
Turn speaker labels on in settings and JesRecap downloads a small diarization model once. After that it runs offline, on your machine, forever — like your transcription model.
Applied after recording
Labels are computed when you press stop, not during the call. Nothing extra runs in the audio path while you are in a meeting, so recording stays light.
"You" versus the others
Your turns are marked You. Everyone else is grouped as Speaker 1, Speaker 2, and so on, in the order they first speak.
No audio leaves the device
The only network request is the initial model download. Your recording, the voice fingerprints derived from it, and the finished transcript never go anywhere.
Carried into exports
Speaker turns appear in the transcript view and in what you export — Markdown notes, plain text, the JSON archive, and SRT/VTT captions.
Off by default, yours to switch
If you do not want labels, leave the setting off and no model is downloaded. You can enable it later and retranscribe an old meeting, because both audio tracks are still on disk.
The dual-track advantage: your words are never misattributed
Most tools face a single mixed recording and have to work out, purely from voice, which parts are the user's. It usually works. When it does not — a cold, a bad headset, two similar voices — your sentences get filed under someone else's name, which is the single most annoying failure a meeting transcript can have.
JesRecap does not have to solve that problem, because it never merges the evidence in the first place. Your microphone is recorded to its own file, and the meeting audio to another. Anything on the microphone track is you, by construction.
- "You" is exact, not inferred. It comes from which device the audio arrived on, so no amount of voice similarity can move your words to another speaker.
- Interruptions survive. When you and someone else talk over each other, the two voices are on two different tracks instead of one muddled waveform.
- Diarization gets an easier job. The model only has to separate the other participants, on a track that does not contain you.
- You can check the source. Export microphone-only or system-audio-only WAV to hear a single side of the conversation on its own.
This is what "keep the tracks separate, mix only for playback" buys you: a cleaner recording, and one label that is always right.
Honest limits
We would rather set expectations than oversell this. Speaker labels are a readability feature, not a forensic identification system.
- Other speakers are not named. They are Speaker 1, Speaker 2, and so on. JesRecap has no participant list, no calendar access, and no voice database, so it cannot know who they are — and it will not guess a name.
- Rooms are harder than calls. On a remote call each remote voice arrives cleanly. In a shared room, several people reach one microphone at different distances, and separation gets less reliable.
- Heavy crosstalk is difficult. When three people on the far side talk at once, the boundaries between their turns are genuinely ambiguous — for any system, local or cloud.
- Similar voices may merge. Two speakers with close vocal characteristics can occasionally be grouped as one, or one speaker split into two after a long gap.
- It costs a little time. Diarization is extra work after the recording stops, on top of transcription.
- Not evidence. Do not treat labels as proof of who said what in a dispute. Keep the audio — it is on your disk — and listen.
In practice, for the common case of a one-to-one or small call with headphones on, labels make a transcript dramatically easier to skim, and your own lines are always your own.
Speaker labels, answered
Is the diarization model really running locally?
Yes. It is a 54 MB file downloaded once to your machine, and it runs there. After that download you can disconnect from the network entirely and speaker labels still work, along with transcription.
Why can't it use people's real names?
Because knowing a name requires knowing a voice in advance, or reading your calendar and participant lists. JesRecap does neither — no accounts, no calendar sync, no voice profiles stored anywhere. Speakers are numbered instead of guessed.
How is "You" identified so reliably?
It is not identified at all, it is recorded separately. Your microphone track is a distinct file from the meeting audio, so your turns are known by their source rather than by voice matching.
Does it work in languages other than English?
Diarization works on voice characteristics rather than words, so it is not tied to a language. Your choice of transcription model is what determines which languages the text itself comes out well in.
Does it slow down my meeting?
No. Labels are computed after you press stop, so nothing extra competes for resources while you are recording. It adds a little processing time at the end.
Can I add labels to a meeting I already recorded?
Yes. Enable speaker labels in settings and retranscribe the meeting. Both audio tracks are kept on disk, so past recordings can be reprocessed at any time.
How many speakers can it handle?
There is no fixed limit, but accuracy is best on smaller conversations. Two-person calls are the easiest case; large meetings with lots of crosstalk are the hardest.
Do the labels appear in exports?
Yes — speaker turns are included in Markdown notes, plain text, the JSON archive, and SRT/VTT caption files.
Know who said what, privately
Local diarization, a separate track for your own voice, and no audio leaving your machine. From €34.95 once.