Newsroom

Introducing M3

The Arabic meeting transcription system inside MeetriX, with an evaluation you can read: 1.62% character error rate on real multi-dialect Arabic meetings, the lowest of the seven systems compared.

M3, the Arabic meeting transcription system inside MeetriX, recorded a 1.62% character error rate on real multi-dialect Arabic business meetings with human-verified reference transcripts, the lowest of the seven systems compared. Today we are publishing the evaluation report behind that number: the scoring rules, the transcript examples and an evidence map, measured against six commercial speech APIs on the same audio and the same references. M3 is the part of MeetriX that turns what people said in a meeting into a speaker-linked, time-aligned Arabic transcript.

The full field, measured on the same audio and scored by the same rules: Google chirp_3 at 2.26%, then Cohere at 4.93%, ElevenLabs Scribe at 5.01%, Deepgram Nova-3 at 6.09%, Google latest_long at 9.01% and Sonix at 9.8%. The headline figure is measured after defined lexical normalisation, and the report says what one evaluation set does and does not support.

1.62%Character error rateMeasured. Lowest of the seven systems compared, on the same human-verified meetings.
2.26%Google chirp_3, the closest systemMeasured. A gap of 0.64 percentage points, about 28% lower CER for M3.
about 45 sTo transcribe one hour of meeting audioTeam-reported, approximate. About 80x real time on an L4 GPU, batch processing, not live latency.
No external transcription callEvaluated transcription pathVerified in the evaluated deployment.

M3 records the lowest character error rate of seven systems on the same human-verified meetings

Lower is better. Character error rate (CER) after defined lexical normalisation.

Same source audio, same human-verified reference, same scoring rules for every system. Evaluation snapshot September 2026. See "How we measured" in the report (/company/research/m3/#method) for the scoring rules.

What goes wrong, and what does not

A meeting record fails in specific ways, and the report shows them segment by segment. In a daily stand-up full of developer vocabulary, the human-verified reference keeps SDK, rendering, highlights, editor, configuration and Chrome extension in Latin script, and so does M3. Google chirp_3 transcribed the passage completely but phonetically, in Arabic letters; Cohere did the same and garbled the ending; Sonix dropped the partner name twice. A phonetic spelling is not an error under the scoring rules, so the difference is what a search across the archive will find later.

In another, a greeting shorter than a second is the whole segment. M3, Cohere, Sonix and chirp_3 matched the reference كيفكم؟ exactly. Deepgram Nova-3 and Google latest_long returned nothing, which the scoring rules count as deletions, and ElevenLabs Scribe produced an English word that was not said. Short turns are where meetings carry their yes, no and agreed, and the report treats them as content.

The comparison also credits what competitors did well. Google chirp_3 handled English technical terminology strongly within Arabic speech and showed low content loss, and on several example segments it is as complete as M3, with the difference limited to phonetic rather than Latin spellings. ElevenLabs Scribe preserved English terminology particularly well. The report's qualitative observations are labelled as such: no separate aggregate rates were calculated for them.

How to read the numbers

CER counts every character that is wrong, missing or extra, divided by the characters in the human-verified reference, after a set of published equivalences: a dialect spelling and its standard form count the same, an English term counts whether written in Latin letters or phonetically in Arabic script, numbers are compared by value, and punctuation is ignored. So 1.62% means roughly 16 wrong characters in every 1,000. The throughput figure, about 45 seconds for a one-hour meeting on an L4 GPU, is team-reported and describes batch processing of recorded audio, not live latency. The evaluation set is a single set of real meetings; the report says what that supports and what it does not, and the next phase expands to additional organisations, platforms, dialects and acoustic conditions.

M3 is not perfect, and the examples show that too: hesitation sounds transcribed where the reference has none, and words added that the reference does not contain. We publish those alongside the cases where M3 keeps a term or a name that other systems drop.

Where to go next

The full report, with the leaderboard, the qualitative matrix, five transcript examples, the methodology and the evidence map, is at lisan.com/company/research/m3. The product view, including what the result means for a government or enterprise meeting record, is on the MeetriX site. M3 is live in MeetriX today, and organisations can evaluate it on their own meetings against a human-verified reference through a pilot. Write to support@lisan.com.

Start with one product.
Keep the whole platform.

Open a free account and solve today's problem in the next ten minutes. When you are ready for more, six flagships and a 20+ app workspace are already on your account. Or talk to us and we will scope it with you.

Free to start, no card · Your data exportable, always · Trusted by more than 24 government entities