Benchmark
Testing AI notetakers on real Arabic meetings
We are running a blind test of meeting transcription on recordings in Egyptian, Levantine and Gulf Arabic, and on meetings that switch between Arabic and English.
Results pending
The test is still running, so this page shows no numbers yet. When it finishes, the full table will appear here with the method and the recording mix.
What we measure
- Word error rate, tolerant of dialect spelling differences
- Errors on names, companies and terms, with a right word in the wrong script counted separately
- Word error rate on code-switched Arabic-English segments
- Speaker errors: who said what
Systems in the test
- Qarar (Kalemio pipeline)
- Cohere Transcribe Arabic, the open-source model, unmodified
- Whisper large-v3
- ElevenLabs
- Fireflies.ai
- Granola (captured by hand)
The recordings
Real business meetings chosen by the Qarar team, each with a human-checked transcript, speaker turns, and a list of names and terms written in both Arabic and Latin script. The recordings are not published.