Type to search · ↑↓ to navigate · Enter to open · Esc to close
Research · September 2026
Measured against 26 LLM configurations. First on every primary metric.
L3, the Lisan language engine inside WriteX, reached 80.29% edit F0.5 on a linguist-reviewed benchmark of real Arabic sentences. Claude Fable 5.1, the strongest large language model configuration we tested, reached 44.68%.
84.54%Precision of its flagsMeasured · next best 64.21%
96.38%Correct replacement once flaggedMeasured · after Arabic normalization, conditional on the flag
0.000014kWh per 1,000 tokens, accelerator scopeMeasured on an NVIDIA A10 · 0.000042 with the three-times full-stack factor, an estimate
F0.5 combines precision and recall and weights precision more heavily. Precision is the share of proposed changes that landed where the linguist also changed the text; recall is the share of the linguist's edits the system found. Full definitions in the methodology.
What this means for an institution's writing
A language checker in a ministry, a newsroom or a bank is judged on four things. Does it flag what is actually wrong? Does it leave correct text alone? Is the suggested fix right? And can the people accountable for the text see why? The L3 evaluation measures the first three and the product answers the fourth.
Precise flags. Of the changes L3 proposed, 84.54% landed on a stretch of text the linguist also changed. That is the number a reviewer feels: a queue of suggestions where most are worth accepting. Measured.
Few false alarms, and few misses. L3 was the only system in the field with precision above 80% and recall above 60% together. The large language models were either cautious and incomplete or thorough and noisy; none managed both on this benchmark. Measured.
Suggestions you can accept. Once a span was flagged correctly, L3's replacement matched the reference 89.79% of the time exactly and 96.38% of the time after Arabic normalization. Those rates are conditional on the flag, and we say so wherever they appear. Measured.
Explanations and ownership. Every WriteX correction carries the rule behind it, and the engine deploys on-premises, so nothing has to leave your perimeter to be checked. The evaluation does not score those; they are what the product does. Product capability.
Energy. Measured on an NVIDIA A10 24 GB, L3 used approximately 0.000014 kWh per 1,000 tokens at accelerator scope. Measured. The measurement itself says nothing about hardware requirements; on-premises deployment is a product capability, described above. Product capability.
L3 leads on edit F0.5 at 80.29%, against 44.68% for Claude Fable 5.1
Edit F0.5, exact-span detection, top ten configurations. Higher is better. F0.5 weights precision above recall, because unnecessary changes cost more than missed ones.
Exact-span detection, micro-averaged, that is, pooled across every edit span before the rate is computed. Single output per system, L3 first suggestion only, the same verbatim four-line prompt. Every competitor score comes from our own run of that provider's service, not from the provider's publications. See the methodology.
The headline number
F0.5 is the headline because it matches how institutions experience a language checker. It combines precision and recall but weights precision more heavily, since an unnecessary change has to be reviewed, rejected and sometimes reverted by hand. L3 scored 80.29%; Claude Fable 5.1 was next at 44.68%, a 35.61-point gap. The strongest comparator changes with the metric: Claude Opus 5 was next on F1 at 34.58% against L3's 74.66%, and GPT-5.6 Sol was next on recall at 41.08% against 66.85%. L3 led on all four. Measured.
L3 is the only system with precision above 80% and recall above 60%
Exact-span precision against recall, every configuration with a results row. Higher is better on both axes. Recall on the horizontal axis, precision on the vertical; the shaded quadrant marks precision above 80% together with recall above 60%.
Exact-span detection. Guides at precision 80% and recall 60%. GPT-6 Astra was run in two reasoning settings; the report's appendices carry a single GPT-6 Astra row, so it appears as one point. Labelled: L3, Claude Fable 5.1, GPT-5.6 Sol, Claude Opus 5.
Few false alarms
The scatter is the chart to remember. Claude Fable 5.1 was the most precise large language model at 64.21%, and its suggestions were usually right, but it found only 20.15% of the reference edits; a document checked with it keeps most of its errors. GPT-5.6 Sol found 41.08% of the edits, the most among the models, at 25.03% precision; a document checked with it comes back with more changes to reject than to accept. Most other configurations sit close together in the lower left. L3 is the one point in the upper right. Measured.
For a writing team this is the difference between a tool people keep on and a tool they switch off. Reviewers tolerate a missed error more readily than a stream of wrong flags, which is exactly the asymmetry F0.5 encodes.
After a correct flag, L3's replacement matches the reference 96.38% of the time, 1.84 points ahead of Claude Fable 5.1
Correction accuracy, exact and after Arabic normalization, conditional on a correct flag. Higher is better. The axis starts at 70%.
Conditional on a correct flag, so read it beside recall: Claude Fable 5.1's rates apply to far fewer flagged spans than L3's. Ten configurations, sorted by the normalized rate. Legend: exact replacement, after Arabic normalization.
Suggestions you can accept
Flagging the right span is half the job. L3's replacement matched the linguist's reference byte for byte on 89.79% of correctly flagged spans, and on 96.38% after normalization, which removes diacritics and tatweel, the elongation character, and unifies alef, hamza and teh marbuta variants so a correct suggestion is not marked wrong for one vowel mark. Claude Fable 5.1 was close on this conditional measure, at 89.07% and 94.54%, but on a much smaller set of flags. High conditional accuracy with low recall means a polite tool that misses a lot; the two numbers have to be read together. Measured.
L3 matches the whole reference sentence on 55.96% of sentences, double Claude Fable 5.1's 27.98%
Sentence exact match, top ten configurations. Higher is better. The share of sentences whose complete output equals the complete reference, character for character.
Deliberately strict: one missed edit, one extra edit or one differing character fails the sentence. GPT-5.6 Terra and Gemini 2.5 Flash tie at 18.49% and keep the source order. Rates only.
Whole sentences, not just spans
An editor receives a sentence, not a list of spans. On the strict end-to-end test, where the whole output must equal the whole reference, L3 matched on 55.96% of sentences, double the 27.98% of Claude Fable 5.1. The test is unforgiving by design: one extra diacritic or one missing period fails the sentence. It is a summary of the whole pipeline, not an error rate. L3 also kept the sentence intact while correcting it, scoring 98.84% on ROUGE-1 F1 against the reference, ahead of Claude Fable 5.1 by 2.28 points. Measured.
What the corrections look like
Four benchmark rows, verbatim: the source, the linguist's reference, L3's output and the outputs of several large language model configurations. Three are outright errors; the fourth is an editorial standard, where the source is grammatical and the reference applies the house convention. The models were told not to rephrase, so that row shows what the engine's editorial layer does, not a grammar failure, and the caption says so. Where a model matched the reference, the card says so too.
Example 1
Outright error
Two hamza seats and a closing period; L3 matched the reference exactly, and the models shown corrected the second hamza but left the first.
Source
نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيئ (erroneous span) له طريقه إلى النجاح (erroneous span)
Reference
نعم يمكن لكتاب واحد يقرؤه (correction) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح. (correction)
L3
Match
نعم يمكن لكتاب واحد يقرؤه (correction) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح. (correction)
Claude Fable 5.1
Partial
نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)
Claude Opus 5
Partial
نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)
GPT-6 Astra
Partial
نعم يمكن لكتاب واحد أن يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح. (correction)
GPT-5.6 Sol
Partial
نعم يمكن لكتاب واحد أن يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)
Gemini 3.1 Pro
Partial
نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)
DeepSeek V4 Pro
Partial
نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)
Why the reference is right
The glottal-stop letter hamza sits on a different carrier depending on the surrounding vowels: in "reads it" it must sit on a waw because the hamza itself carries the u vowel (damma), which outranks the a vowel of the letter before it, and in "lights up" it must stand alone on the line after a long vowel; both source words used the wrong carrier, and the reference also closes the sentence with a period.
hamza seat
punctuation
Example 2
Outright error
A classical hamza-seat rule inside a very common technical term; no large language model configuration produced the reference spelling.
Source
ما المقصود بالتحليلات التنبؤية (erroneous span)؟
Reference
ما المقصود بالتحليلات التنبئية (correction)؟
L3
Match
ما المقصود بالتحليلات التنبئية (correction)؟
Claude Fable 5.1
Unchanged
ما المقصود بالتحليلات التنبؤية (erroneous span)؟
Claude Opus 5
Unchanged
ما المقصود بالتحليلات التنبؤية (erroneous span)؟
GPT-6 Astra
Unchanged
ما المقصود بالتحليلات التنبؤية (erroneous span)؟
GPT-5.6 Sol
Unchanged
ما المقصود بالتحليلات التنبؤية (erroneous span)؟
Gemini 3.1 Pro
Unchanged
ما المقصود بالتحليلات التنبؤية (erroneous span)؟
DeepSeek V4 Pro
Unchanged
ما المقصود بالتحليلات التنبؤية (erroneous span)؟
Why the reference is right
When the noun "prediction" becomes the adjective "predictive", the hamza moves from a waw carrier to a yeh carrier; the source kept the noun spelling inside the adjective, which is a spelling error every LLM left untouched.
hamza seat
Example 3
Outright error
A feminine noun that looks masculine; several configurations, Claude Fable 5.1 among them, matched the reference here.
Source
كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟
Reference
كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟
L3
Match
كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟
Claude Fable 5.1
Match
كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟
Claude Opus 5
Unchanged
كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟
GPT-6 Astra
Match
كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟
GPT-5.6 Sol
Unchanged
كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟
Gemini 3.1 Pro
Match
كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟
DeepSeek V4 Pro
Unchanged
كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟
Why the reference is right
The noun "pride" (al-kibriya') is grammatically feminine in Arabic, its final -aa' being the extended feminine ending, so the demonstrative "this" must take its feminine form.
gender agreement
Example 4
Editorial standard
Editorial standard, not an error: the reference prefers the plain passive verb, and the models, told not to rephrase, kept the construction as written.
Source
تم بناء (erroneous span) هذا النموذج عام 2024.
Reference
بني (correction) هذا النموذج عام 2024.
L3
Match after normalization
بُنِي (differs from the reference only in diacritics or spelling variant) هذا النموذج عام 2024.
Claude Fable 5.1
Unchanged
تم بناء (erroneous span) هذا النموذج عام 2024.
Claude Opus 5
Unchanged
تم بناء (erroneous span) هذا النموذج عام 2024.
GPT-6 Astra
Unchanged
تم بناء (erroneous span) هذا النموذج عام 2024.
GPT-5.6 Sol
Unchanged
تم بناء (erroneous span) هذا النموذج عام 2024.
Gemini 3.1 Pro
Unchanged
تم بناء (erroneous span) هذا النموذج عام 2024.
DeepSeek V4 Pro
Unchanged
تم بناء (erroneous span) هذا النموذج في عام 2024.
Why the reference is right
"The building of this model was completed" is grammatical but house style prefers the direct passive "this model was built"; the LLM prompt forbade rephrasing, so this is a style convention the reference applies, not an outright error.
editorial standard
L3 is second on throughput, behind Qwen 3.8 27B on Groq
Standardized corrected-output tokens per second. Higher is better. End-to-end wall-clock request time, the same tokenizer for every system.
Service throughput including request and prompt-processing overhead, not raw decoder throughput. Qwen 3.8 27B on Groq is faster than L3 and holds first place on the chart. Ten configurations shown.
Speed
L3 delivered 52.28 standardized output tokens per second at 1.041 seconds per timed sentence, second to Qwen 3.8 27B on Groq at 66.98 tokens per second. Claude Fable 5.1, the strongest model on quality, ran at 15.07 tokens per second and 3.585 seconds per sentence. The figure is service throughput, measured end to end with request overhead included, and it is what a writer waiting for the sidebar experiences. It is not the batched, saturated rate used for the energy measurement, and the two are not comparable. Measured.
L3's full-stack energy estimate is approximately 29 to 3,357 times below the cited large language model anchors
kWh per 1,000 processed tokens, log scale: each gridline is ten times the previous. Lower is better. L3 measured; the anchors are cited estimates. The same values are tabulated below.
L3's accelerator figure is a direct measurement on an NVIDIA A10; the full-stack figure applies a three-times overhead factor. The anchors are published or derived full-serving-stack values, cited estimates rather than measurements of the tested API models. The shaded band on the L3 row is the two-times to four-times sensitivity range.
Energy, and why on-premises is realistic
We measured L3 on an NVIDIA A10 24 GB: at least 95% utilization, a 30-second warmup, a 10-minute steady state, power sampled at 1 Hz, approximately 100 W mean GPU power at approximately 2,000 tokens per second. That is 0.05 J per token, or approximately 0.000014 kWh per 1,000 tokens at accelerator scope. Multiplying by a three-times factor for host, cooling, networking and facility overhead gives a full-stack estimate of 0.000042 kWh. Measured.
Energy per 1,000 processed tokens, the values plotted above. The Wh column is the kWh value multiplied by 1,000, so the figures can be read without leading zeros. L3 measured; the anchors are cited estimates, not measurements of the tested API models.
System
kWh per 1,000 tokens
Wh per 1,000 tokens
Relative to the baseline
Evidence and scope
L3
0.000014
0.014
before the 3x factor
Direct measurement on an NVIDIA A10 24 GB, accelerator scope
L3
0.000042
0.042
1x, the baseline
Three-times full-stack normalization, the comparison baseline
Gemini Apps
0.0012
1.2
28.6x
Google disclosure of the full serving stack, cited
GPT-4o-mini
0.0061
6.1
145.2x
IFP School estimate, linearly scaled and normalized, cited
GPT-5 medium
0.074
74
1,761.9x
IFP School estimate, linearly scaled and normalized, cited
GPT-5 high
0.141
141
3,357.1x
IFP School estimate, linearly scaled and normalized, cited
The comparison figures are cited anchors, because API providers publish no power telemetry: Google's Gemini Apps disclosure gives 0.0012 kWh per 1,000 tokens, and the IFP School per-query estimates, scaled the same way, give 0.0061 for GPT-4o-mini, 0.074 for GPT-5 medium and 0.141 for GPT-5 high. On that basis L3's full-stack estimate is approximately 29 to 3,357 times lower. The result depends on the overhead factor, so the report shows the two-times to four-times range; across it, L3 stays at least 21 times more efficient than the lowest anchor. The practical point for institutions is narrower than it sounds: the measurement itself ran on an NVIDIA A10 24 GB, and WriteX offers on-premises deployment. Measured for L3; cited estimates for the anchors; on-premises deployment is a product capability.
Sensitivity of the L3 full-stack estimate to the overhead factor. Gemini Apps is the lowest cited anchor and GPT-5 high the highest; whatever factor is chosen between two and four, L3 stays at least 21 times below the lowest.
Overhead factor
L3 kWh per 1,000 tokens
Advantage over Gemini Apps (0.0012)
Advantage over GPT-5 high (0.141)
2x
0.000028
42.9x
5,036x
3x, used in the report
0.000042
28.6x
3,357x
4x
0.000056
21.4x
2,518x
How the L3 accelerator-scope figure was measured, from the energy methodology in the report.
Measurement
Value
Reference accelerator
NVIDIA A10 24 GB
Steady-state utilization
at least 95%
Warmup and measurement window
30-second warmup, then a 10-minute steady state
Power sampling
1 Hz during the steady-state window
Mean GPU power
approximately 100 W
Sustained throughput
approximately 2,000 tokens per second, batched and saturated
3x, which rounds 1.5 (power-usage effectiveness) × 1.3 (redundancy and reliability) × 1.5 (remaining serving and grid overheads)
Excluded from accelerator scope
host power, cooling, networking, storage and facility overhead
Method
1
The benchmark
A linguist-reviewed set of real Arabic sentences. A linguistic expert wrote a corrected reference for each sentence, or kept it unchanged when nothing needed fixing, so systems are also scored on leaving correct text alone.
2
The prompt
Every model received the same four-line Arabic instruction, verbatim:
صحح الأخطاء في كل جملة من الجمل التالية إن وجدت.
لا تعد صياغة الجملة أو تحذف من كلماتها فقط صحح الأخطاء الإملائية أو القواعدية بها.
لا تضف التشكيل إلى الكلمات.
إذا لم تجد خطأ، أعد الجملة نفسها تماماً.
It forbade rephrasing, word deletion and added diacritics, and required the source back unchanged when no error was found.
3
Single output
One scored output per sentence per system; retries only to obtain a valid response, never to pick the best of several. L3 was scored on its first suggestion only.
4
Scoring
Exact source spans, matched by identical boundaries; precision, recall, F1 and F0.5 micro-averaged across spans. Correction accuracy is scored only on correctly flagged spans, exactly and after Arabic normalization. Sentence exact match is per sentence. ROUGE is macro-averaged per sentence after removing diacritics and tatweel.
5
Throughput and energy
Throughput: standardized tokens of every valid corrected output over end-to-end wall-clock time, request overhead included. Energy: direct A10 measurement, three-times full-stack factor, cited anchors, two-times to four-times sensitivity. The full definitions are in the report.
6
Competitor scores
Every large language model figure was produced by Lisan Research's own run against the provider's service, with hosting as noted in the report's roster. None is taken from a provider's published results.
Evidence
Claim
Scope
Status
80.29% edit F0.5, first among all configurations
exact-span detection, single output
Measured
84.54% precision with 66.85% recall; the only system above 80% and 60% together
52.28 tokens per second, second to Qwen 3.8 27B on Groq
service throughput with overhead
Measured
0.000014 kWh per 1,000 tokens accelerator scope; 0.000042 full stack; approximately 29 to 3,357 times lower than cited anchors
A10 measurement; anchors are cited estimates
Measured for L3, cited estimates for the anchors
Explained corrections, on-premises and sovereign deployment
WriteX
Product capability
Limits
The prompt forbade rephrasing, so rows where the reference applies an editorial convention favour the house standard the linguist worked to; they show the engine applying that standard, not the models failing at grammar.
The energy anchors are cited estimates, not measurements of the tested API models; the three-times factor is the report's choice and models without a published figure were tier-mapped as a scenario estimate.
Throughput includes request overhead and reflects each provider's serving stack on the day of the run.
One benchmark snapshot, 9 September 2026, version 1.0; not independently verified; designed and run by Lisan Research. The definitions, the verbatim prompt and every rate are published so that the protocol can be repeated on other material; the benchmark sentences themselves are available to researchers on request.
L3 missed reference edits and proposed changes the linguist did not; 84.54% precision and 66.85% recall describe both sides.
Inside WriteX
The engine in this evaluation is the one behind WriteX's Arabic language checking in the editor, the Office add-ins and the browser extensions. Product capability. To see how it surfaces in the product: language checking covers the error categories and the rule behind each card, the technology explains the specialised components that make the engine explainable and fast, deployment covers cloud, sovereign cloud and on-premises options, and WriteX against LLM chatbots puts these numbers beside the workflow differences. The full report on lisan.com carries every table, the citation and the version history.
Check your own text
The benchmark is ours; your documents are the test that matters. The editor is free to start, institutions can run L3 on their own infrastructure under a pilot, and methodology questions go to support@lisan.com.
Because an unnecessary change in institutional writing costs more than a missed one: it has to be reviewed, rejected and sometimes reverted. F0.5 weights precision above recall to reflect that. We publish precision, recall and F1 as well, and L3 leads on all four, so the headline choice does not change who is first.
Is this independently verified?
No. Lisan Research designed and ran the evaluation. The full method is published: the verbatim prompt, the single-output rule, the exact-span definitions, the normalization steps and every rate in the appendices. Anyone with a linguist-reviewed Arabic set can repeat it, and the editor is free for that purpose.
Why do the LLMs score low on recall?
The scoring is strict: a model earns credit only when it changes exactly the span the linguist changed, and a punctuation edit can be its own span. The prompt forbade rephrasing and diacritics, so editorial-standard rows and classical rules such as hamza seats in common technical terms were often left untouched. And the models differ: Claude Fable 5.1 changed little and was usually right, at 20.15% recall; GPT-5.6 Sol changed a great deal and was usually not, at 41.08%.
Does 96.38% mean WriteX is 96% accurate?
No. That rate applies to the spans L3 flagged correctly: of those, 96.38% carried the right replacement after Arabic normalization. Its recall was 66.85% and its strict whole-sentence match 55.96%. Read the three together.
Is the engine in the evaluation the one in the product?
Yes. L3 is the Lisan language engine behind WriteX's Arabic language checking in the editor, the Office add-ins, the browser extensions and the API. The evaluation scored its first suggestion only, to match the single-output rule applied to the models. Product capability.