Introducing L3
- L3, the engine that checks Arabic inside WriteX, ranked first on every primary metric against 26 large language model configurations. Measured, September 2026.
- 80.29% edit F0.5 against 44.68% for Claude Fable 5.1, the strongest configuration; the only system with precision above 80% and recall above 60% together.
- Not independently verified. Second on throughput, behind Qwen 3.8 27B on Groq. Prompt, scoring and limits published in full.
For a long time the honest answer to "how good is the Arabic checker in WriteX?" was a demonstration: paste a sentence, count the flags. Today we can give a better answer. We have published the evaluation report for L3, the Lisan language engine that checks Arabic inside WriteX, and it compares L3 with 26 large language model configurations from Anthropic, OpenAI, Google, DeepSeek, Meta and Alibaba Qwen on a linguist-reviewed benchmark of real Arabic sentences.
The result
L3 ranked first on every primary metric. On edit F0.5, the headline measure, it scored 80.29%; the strongest large language model configuration, Claude Fable 5.1, scored 44.68%. It was the only system that combined precision above 80% with recall above 60%. Once it flagged a span correctly, its replacement matched the linguist's reference 96.38% of the time after Arabic normalization, a figure that is conditional on the flag and is labelled that way everywhere it appears. On the strict whole-sentence test, L3 matched the reference on 55.96% of sentences, double the rate of Claude Fable 5.1. All of this is Measured, on one snapshot dated 9 September 2026.
L3 leads on edit F0.5 at 80.29%, against 44.68% for Claude Fable 5.1
Edit F0.5, exact-span detection, top ten configurations. Higher is better. F0.5 weights precision above recall, because unnecessary changes cost more than missed ones.
Why we lead with F0.5
A language checker that flags everything is no more useful than one that flags nothing, and in an institution every wrong flag costs a reviewer's time. F0.5 combines precision and recall but weights precision more heavily, which is the trade-off editors actually make. The models split along exactly that line. Claude Fable 5.1 was precise, at 64.21%, but found only 20.15% of the reference edits. GPT-5.6 Sol found 41.08%, the most among the models, at 25.03% precision. Neither profile is what an institutional reviewer needs.
One sentence
Here is a row from the benchmark. The source: نعم يمكن لكتاب واحد يقرأه الإنسان في مرحلة من عمره فيجد فيه ما يضيئ له طريقه إلى النجاح. Two hamzas are seated on the wrong carrier. In "reads it" the hamza carries a damma, which outranks the fatha before it, so it belongs on a waw; in "lights up" it follows a long yeh, so it stands alone on the line. The reference also closes the sentence with a period: نعم يمكن لكتاب واحد يقرؤه الإنسان في مرحلة من عمره فيجد فيه ما يضيء له طريقه إلى النجاح. L3 returned that exactly. The models shown in the report corrected the second hamza and left the first, and some inserted a word, which the prompt's rule against rephrasing did not allow. The card in the report shows every output, and the rule.
What the report does not claim
Lisan Research ran the evaluation, so it is not independently verified; the method is published in full so that anyone with a linguist-reviewed Arabic set can repeat it. The prompt forbade rephrasing, so rows where the reference applies a house editorial standard favour the engine that implements that standard, and the report labels those rows. On speed, L3 was second, behind Qwen 3.8 27B on Groq. On energy, L3's figure is a direct measurement and the comparison figures are cited estimates, with the sensitivity range shown. And L3 missed edits and made changes the linguist did not; the precision and recall numbers describe both sides.
Read on
The results page shows what the numbers mean inside the product, with the charts, four benchmark rows and the method. Language checking in WriteX covers the categories and the rule behind every card, and WriteX against LLM chatbots now carries the measured comparison. The full report, with every table and the citation, is on lisan.com. Or skip the reading and paste your own text into the editor.
Every word your institution writes, checked.
Official letters to social posts, Arabic and English alike. One assistant that proofreads, enforces your style, and explains every correction, wherever your teams write.
12 error types · Institutional style guides · Works in Office, browsers, and the web editor