Research · Evaluation report · September 2026

M3: Arabic meeting transcription,
measured on real meetings

M3, the Arabic meeting transcription system inside MeetriX, recorded a 1.62% character error rate on real multi-dialect Arabic business meetings with human-verified references, the lowest of the seven systems compared. Google chirp_3 followed at 2.26%.

Research

  • Published
  • VersionEvaluation Edition, September 2026
  • AuthorshipLisan Research
1.62%Character error rateMeasured on the same human-verified meetings as every compared system
28% to 83%Lower CER than each compared systemMeasured, approximate, against six speech APIs
about 45 sTo transcribe one hour of audioTeam-reported, about 80x real time on an L4 GPU, batch not live
7Systems scored on the same audio, reference and rulesM3 and six commercial speech APIs

Why Arabic meetings are different

A business meeting in the region is rarely held in one register. The chair may open in Modern Standard Arabic, a colleague answers in a Gulf or Levantine dialect, and the technical terms, product names and numbers arrive in English inside the Arabic sentence. The evaluated meetings contain all of this: multiple Arabic dialects, English technical terminology inside Arabic speech, several participants in one recording, overlapping speakers, sub-second utterances, silence, and sentences split across recorder boundaries. Qualitative.

Each of those conditions defeats a transcription system in a different way. A sub-second acknowledgement can be classified as silence and dropped. An English abbreviation can be replaced by a similar-sounding Arabic word. A pause can become an invented filler, and a hard segment boundary can cost the first words of a turn. A meeting record is only useful if the numbers, the names and the decisions survive these conditions, which is why the evaluation was run on real meetings rather than read speech.

M3 records the lowest character error rate of seven systems on the same human-verified meetings

Lower is better. Character error rate (CER) after defined lexical normalisation.

Same source audio, same human-verified reference, same scoring rules for every system. Every system received the same participant-separated audio, so speaker attribution is not part of the comparison. Evaluation snapshot September 2026. See "How we measured" for the scoring rules.

Measured. The chart shows the character error rate of the seven systems on the same human-verified evaluation set, after the lexical normalisation described under "How we measured". M3 recorded 1.62%, followed by Google chirp_3 at 2.26%. Cohere (4.93%), ElevenLabs Scribe (5.01%) and Deepgram Nova-3 (6.09%) form a middle group, and Google latest_long at 9.01% and Sonix at 9.8% close the field. In plain terms, CER is the share of characters that had to be added, removed or changed to match what a human verified was said, so 1.62% means roughly 16 wrong characters in every 1,000.

The gap to chirp_3 is 0.64 percentage points, about 28% lower CER. The gap to Sonix is 8.18 points, about 83% lower. All of these figures come from one evaluation set, and the "Limitations" section says what that does and does not support.

The leaderboard

The table below carries the same seven results with their character accuracy (100 minus CER) and M3's approximate relative reduction against each system. Read the reduction column as a description of the gap on this evaluation set, not as a ranking of vendors by quality: a longer gap to a system says as much about that system's configuration for Arabic as about M3.

Character error rate and character accuracy for the seven systems, with M3's approximate relative reduction versus each. Lower CER is better. Measured on the same human-verified meetings; scoring rules under How we measured. best in column
RankSystemCER (%), lower is betterCharacter accuracy (%)M3 vs this system
1M31.62 (best in column)98.38 (best in column)baseline
2Google chirp_32.2697.74about 28% lower (0.64 points)
3Cohere4.9395.07about 67% lower (3.31 points)
4ElevenLabs Scribe5.0194.99about 68% lower (3.39 points)
5Deepgram Nova-36.0993.91about 73% lower (4.47 points)
6Google latest_long9.0190.99about 82% lower (7.39 points)
7Sonix9.890.2about 83% lower (8.18 points) (best in column)

What the transcripts look like

Aggregate rates hide what actually goes wrong in a transcript. The cards below show the human-verified reference and every system's output for individual segments, copied verbatim from the evaluation. Speaker labels are fictional placeholders. Each card carries the source clip: an excerpt of an internal Lisan meeting, the same audio every system transcribed. We chose segments that a reader without Arabic can still inspect: an empty cell, a Latin term, a digit, or a switch into another script is visible in any language. Where M3 itself makes a mistake, the caption says so. Under the scoring rules a phonetic spelling of an English term is not an error, so several of the differences below affect readability and search rather than the CER.

Example 1English terminology
  • Speaker Kareem Saleh
  • Segment 38faf0f4_6
  • Clip length 42.66 s

SDK, rendering, editor, and Chrome extension

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

صباح الخير بالنسبة لي امبارح كان معظم الشغل مع تيم العربي الجديد بخصوص الـ SDK (preserved term) اتسكر تقريبا مشكلتين و إلهن علاقة بالـ rendering (preserved term) الـ highlights (preserved term) بالإديتور اللي معمل له الـ Configuration (preserved term) من عندن بصفى مشكلة وحدة طلبنا منهم اجتماع مشان نكمل معهم اليوم ان شاء الله، كمان امبارح صرت اشتغل على موضوع الـ Chrome extension (preserved term) بين كوركتو و Lisan (preserved term) ما كثير يعني خلصت هذا الموضوع، اليوم ان شاء الله بدي كمل مثل ما قلت بخصوص العربي الجديد وكمان الـ Chrome extension مع كوركتو، يعطيكم العافية.

Reference

صباح الخير، بالنسبة لي امبارح كان معظم الشغل مع تيم العربي الجديد بخصوص الـ SDK اتسكر تقريبا مشكلتين ولهم علاقة بالـ rendering للـ highlights بالـ editor اللي معمل له configuration من عندن بصفى مشكلة وحدة طلبنا منن اجتماع مشان نكمل معن اليوم ان شاء الله كمان امبارح صرت اشتغل على موضوع الـ Chrome extension بين كوركتا و Lisan ما كتير خلصت هذا الموضوع اليوم إن شاء الله بدي كمل متل ما قلت بخصوص العربي الجديد وكمان الـ (Chrome extension) مع العربي الجديد

Why this example matters A daily stand-up full of developer vocabulary. M3 keeps SDK, rendering, highlights, configuration, Chrome extension and Lisan in Latin script. ElevenLabs Scribe keeps them too, with heavy filler. chirp_3 is complete but phonetic, which the scoring rules accept; the difference is what a search box will find later.

  • Preserved: SDK, rendering, highlights, Configuration, Chrome extension, Lisan
  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: content dropped
  • Google latest_long: content dropped
  • Sonix: content dropped
  • Cohere: content dropped
Example 2Filler and punctuation
  • Speaker Rami Haddad
  • Segment a3ef0ea8_735
  • Clip length 20.81 s

Sprint Planning and VAT

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

صباح الخير يعطيكم العافية. آآآه (filler-like insertion) امبارح كان في عندنا متابعة باجتماعات الـ Sprint Planning (preserved term) آآه (filler-like insertion) كان في كمان متابعة مع آآه (filler-like insertion) الشي المطلوب من هيئة الزكاة والضرائب بالسعودية آآه (filler-like insertion) موضوع الـ آآه (filler-like insertion) الـ VAT (preserved term) ضريبة القيمة المضافة. كنت عم بتابع بهدول المواضيع يعطيكم العافية.

Reference

صباح الخير يعطيكم العافية امبارح كان في عنا متابعة باجتماعات السبرنت بلاننج كان في كمان متابعة مع الشي المطلوب من هيئة الزكاة والضرائب بالسعودية موضوع ال الزاد ضريبة القيمة المضافة كان كنت عم بتابع بهدول المواضيع يعطيكم العافية

Why this example matters A hesitant status update. M3 keeps Sprint Planning and VAT in Latin with the Arabic expansion of VAT, while Cohere prints five hesitation tags and several systems garble or drop the VAT term. Honest note: M3 transcribes the speaker's hesitation sounds here, where the reference has none, so this clip does not show M3 removing fillers.

  • Preserved: Sprint Planning, VAT
  • Cohere: unsupported text
  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: unsupported text
  • Google latest_long: unsupported text
  • Sonix: content dropped
  • Google latest_long: content dropped
Example 3Short speech
  • Speaker Omar Khalil
  • Segment ce9f3a3a_10
  • Clip length 0.71 s

A sub-second spoken greeting

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

كيفكم؟

Reference

كيفكم؟

Why this example matters A sub-second greeting. M3, Cohere, Sonix and chirp_3 match the reference exactly. Deepgram Nova-3 and Google latest_long returned nothing, which scores as deletions, and ElevenLabs Scribe produced an English word that was not spoken.

  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: content dropped
  • Google latest_long: content dropped
Example 4Numbers and terms
  • Speaker Tareq Faris
  • Segment 38faf0f4_8
  • Clip length 39.23 s

Files, quality figures, and model work

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

مساء الخير (not supported by the audio) امبارح خلصت الملفات كلهم طلعوا جوا الـ تبع الـ translated (preserved term) اكسل هلأ هنن حوالي 200 (preserved term) file (preserved term) الدقة طلعت يعني بالـ style (preserved term) حوالي الـ 98 (preserved term) يعني بالـ translation (preserved term) كجودة عم بـ 94 (preserved term) هلأ هدا بعد طبعا عدة تعديلات هلأ اليوم بدي ارفع هلأ هي الـ version (preserved term) وكان في كمان شغل على الـ NBOE (preserved term) يعني حاولنا نعمل صاين تيونينغ هلأ اليوم كمان حتابع

Reference

صباح الخير امبارح خلصت الملفات كلهم تبع لينا تبع TranslateX حوالي 200 file الدقه طلعت يعني بالـ style حوالي 98 يعني بالـ translation كجوده بال 94 هلا هذا بعد طبعا عده تعديلات هلا اليوم بدي ارفع هلا هي الـ version وكان في كمان شغل على MOE حاولنا نعمل fine-tunning، وسأتابع اليوم كمان

Why this example matters Three spoken figures: 200 files, a style score of 98 and a translation quality score of 94. M3, Cohere and chirp_3 keep all three. Elsewhere 94 became 40, 98 became eight, 200 was dropped, or 98 became 80 90. M3 is not clean here: it opens with the wrong greeting and misspells two English terms. One spoken first name is replaced by a placeholder in the transcripts.

  • Preserved: translated, file, style, translation, version, NBOE, 200, 98, 94
  • Cohere: unsupported text
  • ElevenLabs Scribe: unsupported text
  • Deepgram Nova-3: unsupported text
  • Sonix: content dropped
  • Google latest_long: content dropped
  • Google chirp_3: content dropped
Example 5English terminology
  • Speaker Kareem Saleh
  • Segment d974dd6c_4
  • Clip length 23.99 s

SDK, dark theme, web app, and monitoring

Source clipAn excerpt of an internal Lisan meeting; the same audio every system transcribed.
  • Preserved term
  • Unsupported text
  • Filler-like insertion
  • No transcript
M3

صباح الخير بالنسبة ليوم الخميس الشغل كان مع فضاءات كان في عندهم مشاكل العلاقة بالـ SDK (preserved term) والسججشنز اللي عم تطلع انحل اغلب الاخطاء صفيان وحدة بدنا نعمل كول عليها اليوم ان شاء الله كمان بدي اشتغل بنسخة الـ dark theme (preserved term) بالـ web app (preserved term) وكمان بدي اعمل مونتورنج للنسخة الجديدة تبع الاي اي ايجنتس يعطيكم العافية

Reference

صباح الخير، بالنسبة ليوم الخميس الشغل كان مع فضاءات، كان في عندهم مشاكل العلاقة بالـ SDK والـ suggestions اللي عم تطلع، انحل اغلب الاخطاء، صفيان وحدة بدنا نعمل كول عليها اليوم إن شاء الله، كمان بدي اشتغل بالـ بنسخة الـ dark theme بالـ web app، وكمان بدي أعمل monitoring للنسخة الجديدة تبع الـ AI agents، يعطيكم العافية.

Why this example matters Consecutive product terms. M3 keeps SDK, dark theme and web app in Latin and writes suggestions, monitoring and AI agents phonetically, which the scoring rules count as correct. ElevenLabs Scribe keeps every term in Latin. Sonix drops most of them, and Google latest_long garbles most of the rest.

  • Preserved: SDK, dark theme, web app
  • ElevenLabs Scribe: unsupported text
  • Sonix: content dropped
  • Google latest_long: content dropped
  • Deepgram Nova-3: content dropped
  • Google chirp_3: content dropped

Output behaviour across systems

Qualitative. Alongside the CER, the reviewers recorded how each system behaved on the evaluation set: whether English terms survived inside Arabic speech, whether content was lost at the start of segments, whether fillers, event tags or repeated punctuation crept in, and whether outputs looped. No separate aggregate rates were calculated for these observations. The matrix shows the reviewers' phrase in every cell so the mapping to strong, mixed, weak and not rated can be checked against the wording.

The honest reading is that Google chirp_3 did well. It handled English technical terminology strongly within Arabic speech, showed low content loss, and produced substantially fewer empty outputs than the V1 configuration. On several of the example segments above it is as complete as M3, with the difference limited to phonetic rather than Latin spellings. ElevenLabs Scribe preserved English terminology particularly well, but produced frequent spurious filler-like vocalisations and one switch into an unexpected language or script. Cohere was highly accurate on Classical Arabic and supported the dialects present, while the review identified content omissions, inconsistent terminology, repeated punctuation and repetition loops. Deepgram Nova-3 showed participant-name errors and segment-start omissions. Google latest_long correctly returned no transcript for many non-speech segments but missed a small number of speech-bearing ones and frequently omitted English terms. Sonix showed errors in names and English terms, segment-start omissions and excessive sentence-final punctuation. All seven systems supported the Arabic dialects present in the evaluated meetings.

M3 recorded the lowest CER and, in the reviewers' notes, preserved business-critical content, numbers and terminology more reliably and did not show the same concentration of filler, tag or repeated-punctuation artifacts in the review. Speaker identification was not evaluated for the external systems: each received the same participant-separated audio prepared by the M3 capture workflow and was scored only on transcription.

Nine criteria, seven systems, one phrase per cell. The M3 column is marked with the brand rule only; status colours are never used for emphasis. The final row repeats the measured CER for reference.

StrongMixedWeakNot rated
Observed output behaviour by criterion and system, from the reviewers' notes on the evaluation set. Levels follow one rule from the report's wording: strong where the reviewers state a positive result, weak where they state a defect, mixed where the note is qualified or says not separately scored or no recurring issue was isolated, and not rated where the criterion was not evaluated. The phrase is shown in every cell. Separate aggregate rates were not calculated.
CohereElevenLabs ScribeDeepgram Nova-3Google latest_longGoogle chirp_3SonixM3
Classical ArabicStrongHighly accurate.StrongAccurate.StrongAccurate.WeakMany errors.StrongAccurate.StrongAccurate.StrongHighly accurate.
English inside ArabicWeakEnglish technical and business terms were sometimes omitted, mistranscribed, or rendered inconsistently within Arabic speech.StrongEnglish technical terminology was preserved particularly well in the reviewed output.MixedEnglish-term performance was not scored separately; reviewed examples showed inconsistent rendering alongside segment-start omissions.WeakEnglish technical terms were frequently omitted.StrongEnglish technical terminology was preserved strongly within Arabic speech.WeakNames and English technical terminology were frequently transcribed incorrectly or omitted.StrongPreserved English technical and business terminology more consistently within Arabic speech.
Text omissionWeakMay omit or distort business-critical content, including numbers, technical terms, and organisation or product names.MixedNo separate omission rate was calculated; the principal qualitative issue was inserted filler-like content rather than a recurring segment-start omission pattern.WeakSubstantial content loss was observed, particularly at the beginning of source segments.MixedCorrectly rejected many non-speech segments, but missed a small number of speech-bearing segments and sometimes omitted English terms or numbers.StrongLow content loss was observed, with substantially fewer empty outputs than the V1 configuration.WeakContent loss was observed particularly at the beginning of source segments.StrongPreserves business-critical content more reliably and reduces the loss of important numbers, terminology, and organisation or product names.
PunctuationWeakRepeated punctuation patterns that reduce readability were observed in the reviewed output.MixedNot separately scored; readability was affected more by filler-like insertions than by a recurring punctuation defect.MixedNot separately scored; no recurring punctuation defect was isolated in the qualitative review.MixedNot separately scored; no recurring punctuation defect was isolated in the qualitative review.MixedNot separately scored; no recurring punctuation defect was highlighted in the qualitative review.WeakPunctuation quality was weak, with excessive sentence-final marks observed in the reviewed output.StrongProduced clearer sentence boundaries without the same concentration of repeated-punctuation artifacts.
Spelling and terminologyWeakTechnical and business terms may appear inconsistently or in phonetic form.MixedEnglish technical terms were handled well, although filler-like insertions reduced overall transcript cleanliness.WeakParticipant names were frequently transcribed incorrectly.WeakEnglish technical terms were frequently missed, and some numbers were omitted.StrongEnglish technical terms were rendered reliably within Arabic speech.WeakPerformance was weak on participant names and English technical terminology.StrongProduces more consistent business terminology for reading and search.
Word repetitionWeakToken and phrase repetition loops were observed in the reviewed output.MixedNo separate repetition rate was calculated; frequent filler-like vocalisations were the more prominent insertion issue.MixedNo recurring repetition issue was isolated in the qualitative review.MixedNo recurring repetition issue was isolated in the qualitative review.MixedNo recurring repetition issue was isolated in the qualitative review.MixedNo recurring word-repetition issue was isolated; punctuation was the more prominent readability problem.StrongShowed substantially fewer repetition problems in the reviewed output.
DialectsStrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.StrongAll systems supported the Arabic dialects present in the evaluated meetings.
Speaker identificationNot ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.Not ratedNot evaluated for the external systems. Each system received the same participant-separated audio segments prepared by the M3 capture workflow and was scored only on transcription.StrongUses participant-linked audio channels and platform participant information. Platform-provided speaker attribution remained linked throughout the evaluated meetings.
Filler words and event tagsWeakFiller-like insertions, event-style tags, and repeated-punctuation artifacts were observed in the reviewed output.WeakFrequent spurious filler-like vocalisations such as “آآآ” were observed, and one segment switched to an unexpected language or script.MixedNo recurring filler-word or event-tag pattern was isolated in the qualitative review.MixedHandled many non-speech segments correctly, although a small number of speech-bearing segments were also rejected.MixedNo recurring filler-word or event-tag pattern was highlighted in the qualitative review.MixedNo recurring event-tag pattern was isolated; excessive punctuation was the more prominent artifact.StrongDid not show the same concentration of filler-like, event-tag, or repeated-punctuation artifacts in the review.
Character error rate (CER), lower is better4.93%5.01%6.09%9.01%2.26%9.8%1.62%

Speed

Team-reported. The M3 meeting transcription workflow processes a one-hour meeting in approximately 45 seconds on an L4 GPU, about 80x real time. The figure describes batch throughput for audio that is already available, measured on a meeting-duration basis. It is not end-to-end live latency. M3 also supports live processing while a meeting is in progress, but live delivery is measured separately and no live latency figure is given in this report.

One hour of meeting audio in about 45 seconds

About 80x real time, team-reported, L4 GPU.

Team-reported batch throughput for the M3 meeting transcription workflow on an L4 GPU. Describes processing of recorded audio, not live speech latency.

From transcript to meeting record

Product capability. The benchmark above scores transcription only. In MeetriX, transcription is one stage of a longer workflow. It is described here as a product capability so that the two are not confused.

  • Meeting capture. Audio is captured with the available participant identity attached from the moment of capture, producing a clearer record of who said what.
  • Participant and time context. Each contribution is kept with its speaker identity and meeting time throughout processing. Platforms that provide separate participant channels and platforms that provide mixed audio with speaker-activity information are both supported.
  • Speech preparation. Silence, short turns, pauses and speech boundaries are handled before transcription, reducing cut-off words, repetition and non-speech artifacts.
  • Arabic transcription and terminology. Multi-dialect Arabic is transcribed, and English names and technical terms spoken inside Arabic sentences are written in their recognised business form so transcripts are easier to read and search.
  • Structured meeting record. Every transcript segment stays connected to its original speaker, meeting time and source audio segment, giving a speaker-linked, time-aligned record that can enter knowledge bases, analytics, CRM or archives.
  • Summary. A post-meeting stage turns the record into a concise summary that can include the meeting title, date and language, attendees, a general summary, key discussion points, decisions and recommendations, and next steps. The template is configurable.
  • Email delivery. The finished summary is sent to the registered email addresses of the participants, following the organisation's access, consent and retention policies.

In the evaluated deployment, the transcription path called no external transcription service. M3 can operate within customer-controlled infrastructure, keeping meeting audio, transcripts and summaries inside the organisation's data-governance boundary, and it adapts its workload to the available compute while preserving speaker identity, timestamps and transcript order. Platform-provided speaker attribution remained linked throughout every evaluated meeting in the M3 participant-channel workflow; this is a coverage statement, and no speaker-attributed error rate is reported.

Note

Product capability, not benchmark. The summary and delivery stages were not scored in this evaluation. The CER figures on this page cover transcription only.

How we measured

Every system was given the same material and scored by the same rules. The steps below condense the comparison protocol and the scoring rules from the report. They are published so that a reader can reproduce the scoring on their own meetings.

  1. Source material

    Real multi-dialect Arabic business meetings, with human-verified reference transcripts prepared for every scored segment.

  2. Same audio for every system

    All seven systems received the same source audio, pre-segmented by the M3 capture workflow into participant-separated segments. External systems were scored on transcription only; their speaker diarization, the step that splits audio by who is speaking, was not used or evaluated. Every competitor figure on this page was produced by running the vendor's API on the evaluation audio; none is taken from a vendor publication.

  3. Configurations

    Google Cloud Speech-to-Text V2 was run as chirp_3 with ar-XA; Google Cloud Speech-to-Text V1 as latest_long with ar-SA; ElevenLabs as scribe_v1 with language_code=ara; Cohere, Deepgram Nova-3 and Sonix.ai as named, with no further model string given in the report. M3 was run as the complete meeting system with participant-linked audio channels from its capture workflow.

  4. Metric

    CER = (substitutions + deletions + insertions) divided by the characters in the human-verified reference, computed after normalisation. Character accuracy is 100 minus CER.

  5. Dialect and standard forms

    An Arabic word is correct whether written in its dialect form or its Modern Standard Arabic form, as long as the meaning is the same.

  6. English words

    An English word is correct whether written in Latin letters, spelled phonetically in Arabic script, or translated into an equivalent Arabic word.

  7. Numbers

    Compared by value: 15, ١٥ and خمسة عشر count as the same thing. Numbers and English business terms stay in the scored content; they are not stripped out.

  8. Punctuation, case, letter variants, whitespace

    Punctuation is ignored entirely. Latin text is case-insensitive. Arabic letter variants such as the forms of alef are normalised. Whitespace is ignored, except that a missing space that fuses two words counts as one error.

  9. Remaining differences

    After these equivalences, every remaining missing, added or substituted character is an error. Only segments with a verified reference enter the CER. An empty output against a speech-bearing reference counts as deletions; extra unsupported output counts as insertions.

  10. Qualitative review

    Reviewers recorded content omission, names and English terminology, filler-like insertions, unexpected language changes, repeated punctuation and repetition loops per system. No separate aggregate rates were calculated.

Evidence map

ClaimScopeStatus
M3 1.62% CER versus 2.26% for Google chirp_3, 4.93% for Cohere, 5.01% for ElevenLabs Scribe, 6.09% for Deepgram Nova-3, 9.01% for Google latest_long and 9.8% for Sonix; relative reductions of approximately 28%, 67%, 68%, 73%, 82% and 83%.The same human-verified evaluation set of real meetings, same reference, same scoring rules.Measured.
Deepgram Nova-3 showed participant-name errors and segment-start omissions; ElevenLabs Scribe preserved English technical terms well but produced frequent spurious filler-like vocalisations and one unexpected language or script switch; Google latest_long handled many non-speech segments correctly but missed a small number of speech-bearing segments and frequently omitted English terms, with occasional number omissions; Google chirp_3 handled English technical terms strongly and showed low content loss; Sonix showed errors in names and English technical terms, segment-start omissions and excessive sentence-final punctuation.Qualitative review of the same evaluation set.Observed qualitatively; separate aggregate rates not calculated.
Platform-provided speaker attribution remained linked throughout every evaluated meeting.M3 participant-channel workflow; external systems were evaluated for transcription only.Measured for the M3 participant-channel workflow.
Approximately 80x batch throughput; about 45 seconds per hour of meeting audio.M3 meeting transcription workflow for a one-hour meeting on an L4 GPU.Team-reported; batch throughput, not live latency.
No external transcription call.Evaluated transcription path.Verified in the evaluated deployment.

Limitations

  • One evaluation set. Every CER figure comes from a single human-verified evaluation set of real meetings. It shows how the seven systems compared on that material and should not be read as a universal rate for any of them.
  • A snapshot in time. The compared APIs were called in September 2026 with the configurations listed under "How we measured". Vendors update their models, so the figures describe those versions on that date.
  • A small margin at the top. The gap between M3 and Google chirp_3 is 0.64 percentage points, and on several example segments chirp_3 is as complete as M3. The ordering is measured; its weight should be read with the single-set caveat.
  • Throughput is team-reported. The 45 seconds per hour and 80x figures were reported by the team for one hardware configuration and are not part of the scored benchmark. Live latency was not measured for this report.
  • Qualitative observations carry no aggregate rates. Statements about fillers, omissions, repetition and punctuation come from reviewer notes. The matrix levels follow one rule applied to the wording of those notes, and the phrases are shown so the mapping can be challenged.
  • External speaker diarization was not evaluated. Every external system received participant-separated audio prepared by M3, so the comparison isolates transcription and says nothing about how those systems attribute speakers on mixed audio. The M3 speaker-attribution result is a coverage statement; no speaker-attributed error rate is reported.
  • M3 is not perfect. The example cards show M3 transcribing hesitation sounds where the reference has none, and adding words that the reference does not contain.
  • Next steps. The next evaluation phase expands testing across additional organisations, meeting platforms, dialects, acoustic conditions and customer-specific terminology. Organisations can evaluate M3 on their own meetings now.
Full results
Full results, all seven systems. Measured on the same human-verified meetings, same reference, same scoring rules; see How we measured. best in column
RankSystemProtocol detailCER (%), lower is betterCharacter accuracy (%)M3 relative reduction (%), approximateAbsolute reduction, points
1Complete M3 meeting systemComplete M3 meeting system; participant-linked audio channels supplied by the M3 capture workflow1.62 (best in column)98.38 (best in column)baseline
2Google Cloud Speech-to-Text V2 (chirp_3)chirp_3, ar-XA2.2697.74280.64
3CohereCohere (no model string given in the white paper)4.9395.07673.31
4ElevenLabs Scribe v1scribe_v1, language_code=ara5.0194.99683.39
5Deepgram Nova-3Deepgram Nova-3 (no further model string given)6.0993.91734.47
6Google Cloud Speech-to-Text V1 (latest_long)latest_long, ar-SA9.0190.99827.39
7Sonix.aiSonix.ai (no model string given)9.890.283 (best in column)8.18 (best in column)

To cite this report, use the reference below or the BibTeX record.

Cite this page

Lisan Research (2026). M3: Arabic meeting transcription, measured on real meetings. Evaluation Edition, September 2026. Lisan, 15 September 2026. https://lisan.com/company/research/m3/
BibTeX
@techreport{lisan2026m3,
  author = {Lisan Research},
  title = {M3: Arabic meeting transcription, measured on real meetings},
  institution = {Lisan},
  type = {Evaluation report},
  year = {2026},
  month = {September},
  note = {Evaluation Edition, September 2026. Published 15 September 2026},
  url = {https://lisan.com/company/research/m3/},
}

Run M3 on your own meetings

The benchmark that matters is your audio. A pilot runs M3 on a sample of your meetings with a human-verified reference, and we publish the scoring rules to you before the run, so the result is measured the same way this report was.

Frequently asked questions

What is CER, and does this report use WER?

CER, the character error rate, is the share of characters in a transcript that had to be added, removed or changed to match what a human verified was said. This report scores CER after the normalisation rules listed under "How we measured", so dialect and standard spellings, Latin and phonetic English, and numeric forms count as equivalent. Word error rate is not reported here, and earlier MeetriX figures published as WER are superseded by this evaluation.

Was the evaluation independently verified?

No. Lisan Research ran the evaluation. The protocol, scoring rules, examples and evidence map are published on this page so the scoring can be reproduced, and a pilot on your own meetings with a human-verified reference is the way to check the result on your material.

How close is Google chirp_3?

Close. It recorded 2.26% against M3's 1.62%, a gap of 0.64 percentage points, and the review credits it with strong handling of English terms and low content loss. The differences that remain are Latin versus phonetic spellings of terms, the concentration of artifacts, and the speaker-linked workflow around M3, which was outside the transcription comparison.

Does M3 send meeting audio to a cloud speech API?

In the evaluated deployment the transcription path called no external transcription service. M3 can run within customer-controlled infrastructure, keeping meeting audio, transcripts and summaries inside the organisation's data-governance boundary.

Can we evaluate M3 on our meetings?

Yes. A pilot runs M3 on a sample of your meetings against a human-verified reference, with the scoring rules shared before the run. Write to support@lisan.com or book a pilot through MeetriX.

References

  • Lisan Research (2026). M3: Arabic meeting transcription, measured on real meetings. Evaluation Edition, September 2026 (the white paper this page condenses).
  • Google Cloud Speech-to-Text V2 documentation, chirp_3 model, language code ar-XA (configuration used in step 3 of "How we measured").
  • Google Cloud Speech-to-Text V1 documentation, latest_long model, language code ar-SA (configuration used in step 3).
  • ElevenLabs Scribe v1 documentation, scribe_v1 with language_code=ara (configuration used in step 3).
  • Cohere, Deepgram Nova-3 and Sonix.ai: vendor documentation for the products as named; the report gives no further model string.