M3: Arabic meeting transcription
Evaluation report ·1.62% character error rate on real multi-dialect Arabic meetings with human-verified transcripts, the lowest of seven systems scored under the same rules. Google chirp_3 was next at 2.26%.
Type to search · ↑↓ to navigate · Enter to open · Esc to close
Company
Papers written from production systems, inventions nobody else has built, and benchmarks we measured instead of promised.
Lisan Research
Evaluations we designed and ran ourselves, published with the method, the competitors' versions and the limits.
1.62% character error rate on real multi-dialect Arabic meetings with human-verified transcripts, the lowest of seven systems scored under the same rules. Google chirp_3 was next at 2.26%.
80.29% edit F0.5 on a linguist-reviewed benchmark of real Arabic sentences, against 44.68% for Claude Fable 5.1, the strongest of 26 large language model configurations.
Why we publish
Behind the products sits over a decade of language engineering: 19.5 million lines of internally developed software, 32 AI components, and a data operation that never stops. The papers below document how the hardest parts actually work, and the innovations list shows the parts nobody else has built.
The core inventions
Most language AI throws compute at text one opaque layer at a time. Lisan's core stack does the opposite: compress the language using its own morphology, then process it through explicit, explainable spaces. The result is energy-efficient, sovereign AI that runs where transformers cannot, from your data center down to edge devices.
Arabic words decompose into roots and morphemes, so the representation shrinks far below traditional tokenization, losslessly. vs traditional tokenization in transformer models
The compressed stream crosses explicit, explainable spaces instead of one black box.
Papers
Written from production systems, not lab prototypes. Each one documents a problem the platform had to solve for real customers, prepared for peer-reviewed publication.
How the Lisan Engine pairs a ~10M-line rule-based Java engine with a neural model server over gRPC, routing each request to the engine that does it best: deterministic sub-second grammar checking on one side, transformer-scale paraphrasing and summarization on the other.
Traditional diarization guesses who spoke when and still gets 5-15% of it wrong. By capturing a separate audio stream per meeting participant, Lisan makes speaker identity deterministic: the error rate of guessing drops out of the equation entirely.
A production morphological analyzer that decomposes any Arabic word into root, pattern, and stem, handling the clitics and missing diacritics that give a single token dozens of valid readings, at speeds that keep grammar checking interactive.
Arabic's cursive script and contextual letter shapes break conventional OCR. OCR PLUS combines multiple recognition engines and corrects their output with an n-gram Viterbi decoder, so the language model fixes what the vision models miss.
Translating a document is easy; giving it back with the layout intact is not. This work covers TranslateX's format-aware pipeline for DOCX, XLSX, PPTX, PDF, and images, including the bidirectional-text handling Arabic demands.
Raw ASR output is fragmented and full of platform artifacts. Eleven specialized merging heuristics, tuned per meeting platform, turn it into transcripts that read the way the meeting actually sounded.
The Lisan data layer runs 162 specialized pipelines: scrapers, corpus processors, n-gram calculators, entity extractors, and dictionary builders in a continuous feedback loop that makes the production engine better every week.
Spell-check models train best on realistic mistakes. By modeling which keys sit next to each other on a real keyboard, synthetic typos match the errors humans actually make, and the corrector learns the right lessons.
LLM-based scoring gives a different answer every run. This system scores text with deterministic linguistic metrics instead, so the same paragraph always gets the same score, and the riskiest paragraphs always surface first.
One bot controller, four meeting platforms: Zoom, Microsoft Teams, Google Meet, and Webex, each behind an interchangeable adapter. The architecture that lets MeetriX join any meeting and leave with the minutes.
Measured, not promised
Lower is better. Character error rate (CER) after defined lexical normalisation. Lisan evaluation, September 2026.
Innovations
A selection from the platform's invention portfolio: from the language core that reads Arabic the way it is actually built, to files that carry their own identity.
The language core
6Documents & files
6Speech & meetings
3Inside the products
5The data operation
A curated lexicon of modern standard Arabic built by intersecting expert dictionaries with live corpora: 681,771 expert-verified words, plus 14,300 modern words validated by hand.
Our in-house annotation studio combines active learning and reinforcement learning with a team of 32 annotators, so every correction a human makes teaches the models.
Speech data spanning modern standard Arabic and 32 dialects across 17 countries. It is why the transcription holds up when the meeting switches from Cairo to Riyadh mid-sentence.
For research collaborations, benchmark methodology, or a technical deep dive under NDA, talk to the team that built it.
They are written from production systems and prepared for submission to peer-reviewed venues. The engineering they describe is running in products today; the publication process is underway.
Yes, it is published. The M3 evaluation report states the protocol (same audio, same human-verified reference, same rules for all seven systems), the full scoring rules for CER after normalisation, the configurations of every compared system, five transcript examples, an evidence map with the status of every claim, and the limitations. Read it at the M3 report. For a pilot on your own meetings, talk to us.
Open a free account and solve today's problem in the next ten minutes. When you are ready for more, six flagships and a 20+ app workspace are already on your account. Or talk to us and we will scope it with you.
Free to start, no card · Your data exportable, always · Trusted by more than 24 government entities