Research · September 2026

Measured against 26 LLM configurations.
First on every primary metric.

L3, the Lisan language engine inside WriteX, reached 80.29% edit F0.5 on a linguist-reviewed benchmark of real Arabic sentences. Claude Fable 5.1, the strongest large language model configuration we tested, reached 44.68%.

Research

  • Published
  • Version1.0
  • OrganisationLisan Research
  • Evaluation snapshot9 September 2026
80.29%Edit F0.5Measured · Claude Fable 5.1 44.68%
84.54%Precision of its flagsMeasured · next best 64.21%
96.38%Correct replacement once flaggedMeasured · after Arabic normalization, conditional on the flag
0.000014kWh per 1,000 tokens, accelerator scopeMeasured on an NVIDIA A10 · 0.000042 with the three-times full-stack factor, an estimate

F0.5 combines precision and recall and weights precision more heavily. Precision is the share of proposed changes that landed where the linguist also changed the text; recall is the share of the linguist's edits the system found. Full definitions in the methodology.

What this means for an institution's writing

A language checker in a ministry, a newsroom or a bank is judged on four things. Does it flag what is actually wrong? Does it leave correct text alone? Is the suggested fix right? And can the people accountable for the text see why? The L3 evaluation measures the first three and the product answers the fourth.

Precise flags. Of the changes L3 proposed, 84.54% landed on a stretch of text the linguist also changed. That is the number a reviewer feels: a queue of suggestions where most are worth accepting. Measured.

Few false alarms, and few misses. L3 was the only system in the field with precision above 80% and recall above 60% together. The large language models were either cautious and incomplete or thorough and noisy; none managed both on this benchmark. Measured.

Suggestions you can accept. Once a span was flagged correctly, L3's replacement matched the reference 89.79% of the time exactly and 96.38% of the time after Arabic normalization. Those rates are conditional on the flag, and we say so wherever they appear. Measured.

Explanations and ownership. Every WriteX correction carries the rule behind it, and the engine deploys on-premises, so nothing has to leave your perimeter to be checked. The evaluation does not score those; they are what the product does. Product capability.

Energy. Measured on an NVIDIA A10 24 GB, L3 used approximately 0.000014 kWh per 1,000 tokens at accelerator scope. Measured. The measurement itself says nothing about hardware requirements; on-premises deployment is a product capability, described above. Product capability.

L3 leads on edit F0.5 at 80.29%, against 44.68% for Claude Fable 5.1

Edit F0.5, exact-span detection, top ten configurations. Higher is better. F0.5 weights precision above recall, because unnecessary changes cost more than missed ones.

Exact-span detection, micro-averaged, that is, pooled across every edit span before the rate is computed. Single output per system, L3 first suggestion only, the same verbatim four-line prompt. Every competitor score comes from our own run of that provider's service, not from the provider's publications. See the methodology.

The headline number

F0.5 is the headline because it matches how institutions experience a language checker. It combines precision and recall but weights precision more heavily, since an unnecessary change has to be reviewed, rejected and sometimes reverted by hand. L3 scored 80.29%; Claude Fable 5.1 was next at 44.68%, a 35.61-point gap. The strongest comparator changes with the metric: Claude Opus 5 was next on F1 at 34.58% against L3's 74.66%, and GPT-5.6 Sol was next on recall at 41.08% against 66.85%. L3 led on all four. Measured.

L3 is the only system with precision above 80% and recall above 60%

Exact-span precision against recall, every configuration with a results row. Higher is better on both axes. Recall on the horizontal axis, precision on the vertical; the shaded quadrant marks precision above 80% together with recall above 60%.

Exact-span detection. Guides at precision 80% and recall 60%. GPT-6 Astra was run in two reasoning settings; the report's appendices carry a single GPT-6 Astra row, so it appears as one point. Labelled: L3, Claude Fable 5.1, GPT-5.6 Sol, Claude Opus 5.

Few false alarms

The scatter is the chart to remember. Claude Fable 5.1 was the most precise large language model at 64.21%, and its suggestions were usually right, but it found only 20.15% of the reference edits; a document checked with it keeps most of its errors. GPT-5.6 Sol found 41.08% of the edits, the most among the models, at 25.03% precision; a document checked with it comes back with more changes to reject than to accept. Most other configurations sit close together in the lower left. L3 is the one point in the upper right. Measured.

For a writing team this is the difference between a tool people keep on and a tool they switch off. Reviewers tolerate a missed error more readily than a stream of wrong flags, which is exactly the asymmetry F0.5 encodes.

After a correct flag, L3's replacement matches the reference 96.38% of the time, 1.84 points ahead of Claude Fable 5.1

Correction accuracy, exact and after Arabic normalization, conditional on a correct flag. Higher is better. The axis starts at 70%.

Conditional on a correct flag, so read it beside recall: Claude Fable 5.1's rates apply to far fewer flagged spans than L3's. Ten configurations, sorted by the normalized rate. Legend: exact replacement, after Arabic normalization.

Suggestions you can accept

Flagging the right span is half the job. L3's replacement matched the linguist's reference byte for byte on 89.79% of correctly flagged spans, and on 96.38% after normalization, which removes diacritics and tatweel, the elongation character, and unifies alef, hamza and teh marbuta variants so a correct suggestion is not marked wrong for one vowel mark. Claude Fable 5.1 was close on this conditional measure, at 89.07% and 94.54%, but on a much smaller set of flags. High conditional accuracy with low recall means a polite tool that misses a lot; the two numbers have to be read together. Measured.

L3 matches the whole reference sentence on 55.96% of sentences, double Claude Fable 5.1's 27.98%

Sentence exact match, top ten configurations. Higher is better. The share of sentences whose complete output equals the complete reference, character for character.

Deliberately strict: one missed edit, one extra edit or one differing character fails the sentence. GPT-5.6 Terra and Gemini 2.5 Flash tie at 18.49% and keep the source order. Rates only.

Whole sentences, not just spans

An editor receives a sentence, not a list of spans. On the strict end-to-end test, where the whole output must equal the whole reference, L3 matched on 55.96% of sentences, double the 27.98% of Claude Fable 5.1. The test is unforgiving by design: one extra diacritic or one missing period fails the sentence. It is a summary of the whole pipeline, not an error rate. L3 also kept the sentence intact while correcting it, scoring 98.84% on ROUGE-1 F1 against the reference, ahead of Claude Fable 5.1 by 2.28 points. Measured.

What the corrections look like

Four benchmark rows, verbatim: the source, the linguist's reference, L3's output and the outputs of several large language model configurations. Three are outright errors; the fourth is an editorial standard, where the source is grammatical and the reference applies the house convention. The models were told not to rephrase, so that row shows what the engine's editorial layer does, not a grammar failure, and the caption says so. Where a model matched the reference, the card says so too.

Example 1

Outright error

Two hamza seats and a closing period; L3 matched the reference exactly, and the models shown corrected the second hamza but left the first.

Source

نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيئ (erroneous span) له طريقه إلى النجاح (erroneous span)

Reference

نعم يمكن لكتاب واحد يقرؤه (correction) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح. (correction)

L3
Match

نعم يمكن لكتاب واحد يقرؤه (correction) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح. (correction)

Claude Fable 5.1
Partial

نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)

Claude Opus 5
Partial

نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)

GPT-6 Astra
Partial

نعم يمكن لكتاب واحد أن يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح. (correction)

GPT-5.6 Sol
Partial

نعم يمكن لكتاب واحد أن يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)

Gemini 3.1 Pro
Partial

نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)

DeepSeek V4 Pro
Partial

نعم يمكن لكتاب واحد يقرأه (erroneous span) الإنسان في مرحلة من عمره فيجد فيه ما يضيء (correction) له طريقه إلى النجاح (erroneous span)

Why the reference is right

The glottal-stop letter hamza sits on a different carrier depending on the surrounding vowels: in "reads it" it must sit on a waw because the hamza itself carries the u vowel (damma), which outranks the a vowel of the letter before it, and in "lights up" it must stand alone on the line after a long vowel; both source words used the wrong carrier, and the reference also closes the sentence with a period.

  • hamza seat
  • punctuation
Example 2

Outright error

A classical hamza-seat rule inside a very common technical term; no large language model configuration produced the reference spelling.

Source

ما المقصود بالتحليلات التنبؤية (erroneous span)؟

Reference

ما المقصود بالتحليلات التنبئية (correction)؟

L3
Match

ما المقصود بالتحليلات التنبئية (correction)؟

Claude Fable 5.1
Unchanged

ما المقصود بالتحليلات التنبؤية (erroneous span)؟

Claude Opus 5
Unchanged

ما المقصود بالتحليلات التنبؤية (erroneous span)؟

GPT-6 Astra
Unchanged

ما المقصود بالتحليلات التنبؤية (erroneous span)؟

GPT-5.6 Sol
Unchanged

ما المقصود بالتحليلات التنبؤية (erroneous span)؟

Gemini 3.1 Pro
Unchanged

ما المقصود بالتحليلات التنبؤية (erroneous span)؟

DeepSeek V4 Pro
Unchanged

ما المقصود بالتحليلات التنبؤية (erroneous span)؟

Why the reference is right

When the noun "prediction" becomes the adjective "predictive", the hamza moves from a waw carrier to a yeh carrier; the source kept the noun spelling inside the adjective, which is a spelling error every LLM left untouched.

  • hamza seat
Example 3

Outright error

A feminine noun that looks masculine; several configurations, Claude Fable 5.1 among them, matched the reference here.

Source

كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟

Reference

كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟

L3
Match

كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟

Claude Fable 5.1
Match

كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟

Claude Opus 5
Unchanged

كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟

GPT-6 Astra
Match

كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟

GPT-5.6 Sol
Unchanged

كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟

Gemini 3.1 Pro
Match

كيف لأحد أن يمتلك مثل هذه الكبرياء (correction)؟

DeepSeek V4 Pro
Unchanged

كيف لأحد أن يمتلك مثل هذا الكبرياء (erroneous span)؟

Why the reference is right

The noun "pride" (al-kibriya') is grammatically feminine in Arabic, its final -aa' being the extended feminine ending, so the demonstrative "this" must take its feminine form.

  • gender agreement
Example 4

Editorial standard

Editorial standard, not an error: the reference prefers the plain passive verb, and the models, told not to rephrase, kept the construction as written.

Source

تم بناء (erroneous span) هذا النموذج عام 2024.

Reference

بني (correction) هذا النموذج عام 2024.

L3
Match after normalization

بُنِي (differs from the reference only in diacritics or spelling variant) هذا النموذج عام 2024.

Claude Fable 5.1
Unchanged

تم بناء (erroneous span) هذا النموذج عام 2024.

Claude Opus 5
Unchanged

تم بناء (erroneous span) هذا النموذج عام 2024.

GPT-6 Astra
Unchanged

تم بناء (erroneous span) هذا النموذج عام 2024.

GPT-5.6 Sol
Unchanged

تم بناء (erroneous span) هذا النموذج عام 2024.

Gemini 3.1 Pro
Unchanged

تم بناء (erroneous span) هذا النموذج عام 2024.

DeepSeek V4 Pro
Unchanged

تم بناء (erroneous span) هذا النموذج في عام 2024.

Why the reference is right

"The building of this model was completed" is grammatical but house style prefers the direct passive "this model was built"; the LLM prompt forbade rephrasing, so this is a style convention the reference applies, not an outright error.

  • editorial standard

L3 is second on throughput, behind Qwen 3.8 27B on Groq

Standardized corrected-output tokens per second. Higher is better. End-to-end wall-clock request time, the same tokenizer for every system.

Service throughput including request and prompt-processing overhead, not raw decoder throughput. Qwen 3.8 27B on Groq is faster than L3 and holds first place on the chart. Ten configurations shown.

Speed

L3 delivered 52.28 standardized output tokens per second at 1.041 seconds per timed sentence, second to Qwen 3.8 27B on Groq at 66.98 tokens per second. Claude Fable 5.1, the strongest model on quality, ran at 15.07 tokens per second and 3.585 seconds per sentence. The figure is service throughput, measured end to end with request overhead included, and it is what a writer waiting for the sidebar experiences. It is not the batched, saturated rate used for the energy measurement, and the two are not comparable. Measured.

L3's full-stack energy estimate is approximately 29 to 3,357 times below the cited large language model anchors

kWh per 1,000 processed tokens, log scale: each gridline is ten times the previous. Lower is better. L3 measured; the anchors are cited estimates. The same values are tabulated below.

L3's accelerator figure is a direct measurement on an NVIDIA A10; the full-stack figure applies a three-times overhead factor. The anchors are published or derived full-serving-stack values, cited estimates rather than measurements of the tested API models. The shaded band on the L3 row is the two-times to four-times sensitivity range.

Energy, and why on-premises is realistic

We measured L3 on an NVIDIA A10 24 GB: at least 95% utilization, a 30-second warmup, a 10-minute steady state, power sampled at 1 Hz, approximately 100 W mean GPU power at approximately 2,000 tokens per second. That is 0.05 J per token, or approximately 0.000014 kWh per 1,000 tokens at accelerator scope. Multiplying by a three-times factor for host, cooling, networking and facility overhead gives a full-stack estimate of 0.000042 kWh. Measured.

Energy per 1,000 processed tokens, the values plotted above. The Wh column is the kWh value multiplied by 1,000, so the figures can be read without leading zeros. L3 measured; the anchors are cited estimates, not measurements of the tested API models.
SystemkWh per 1,000 tokensWh per 1,000 tokensRelative to the baselineEvidence and scope
L30.0000140.014before the 3x factorDirect measurement on an NVIDIA A10 24 GB, accelerator scope
L30.0000420.0421x, the baselineThree-times full-stack normalization, the comparison baseline
Gemini Apps0.00121.228.6xGoogle disclosure of the full serving stack, cited
GPT-4o-mini0.00616.1145.2xIFP School estimate, linearly scaled and normalized, cited
GPT-5 medium0.074741,761.9xIFP School estimate, linearly scaled and normalized, cited
GPT-5 high0.1411413,357.1xIFP School estimate, linearly scaled and normalized, cited

The comparison figures are cited anchors, because API providers publish no power telemetry: Google's Gemini Apps disclosure gives 0.0012 kWh per 1,000 tokens, and the IFP School per-query estimates, scaled the same way, give 0.0061 for GPT-4o-mini, 0.074 for GPT-5 medium and 0.141 for GPT-5 high. On that basis L3's full-stack estimate is approximately 29 to 3,357 times lower. The result depends on the overhead factor, so the report shows the two-times to four-times range; across it, L3 stays at least 21 times more efficient than the lowest anchor. The practical point for institutions is narrower than it sounds: the measurement itself ran on an NVIDIA A10 24 GB, and WriteX offers on-premises deployment. Measured for L3; cited estimates for the anchors; on-premises deployment is a product capability.

Sensitivity of the L3 full-stack estimate to the overhead factor. Gemini Apps is the lowest cited anchor and GPT-5 high the highest; whatever factor is chosen between two and four, L3 stays at least 21 times below the lowest.
Overhead factorL3 kWh per 1,000 tokensAdvantage over Gemini Apps (0.0012)Advantage over GPT-5 high (0.141)
2x0.00002842.9x5,036x
3x, used in the report0.00004228.6x3,357x
4x0.00005621.4x2,518x
How the L3 accelerator-scope figure was measured, from the energy methodology in the report.
MeasurementValue
Reference acceleratorNVIDIA A10 24 GB
Steady-state utilizationat least 95%
Warmup and measurement window30-second warmup, then a 10-minute steady state
Power sampling1 Hz during the steady-state window
Mean GPU powerapproximately 100 W
Sustained throughputapproximately 2,000 tokens per second, batched and saturated
Energy per token0.05 J
Energy per 1,000 tokens50 J = 0.01389 Wh = 0.0000139 kWh, reported as 0.000014
Full-stack factor3x, which rounds 1.5 (power-usage effectiveness) × 1.3 (redundancy and reliability) × 1.5 (remaining serving and grid overheads)
Excluded from accelerator scopehost power, cooling, networking, storage and facility overhead

Method

  1. The benchmark

    A linguist-reviewed set of real Arabic sentences. A linguistic expert wrote a corrected reference for each sentence, or kept it unchanged when nothing needed fixing, so systems are also scored on leaving correct text alone.

  2. The prompt

    Every model received the same four-line Arabic instruction, verbatim:

    صحح الأخطاء في كل جملة من الجمل التالية إن وجدت.

    لا تعد صياغة الجملة أو تحذف من كلماتها فقط صحح الأخطاء الإملائية أو القواعدية بها.

    لا تضف التشكيل إلى الكلمات.

    إذا لم تجد خطأ، أعد الجملة نفسها تماماً.

    It forbade rephrasing, word deletion and added diacritics, and required the source back unchanged when no error was found.

  3. Single output

    One scored output per sentence per system; retries only to obtain a valid response, never to pick the best of several. L3 was scored on its first suggestion only.

  4. Scoring

    Exact source spans, matched by identical boundaries; precision, recall, F1 and F0.5 micro-averaged across spans. Correction accuracy is scored only on correctly flagged spans, exactly and after Arabic normalization. Sentence exact match is per sentence. ROUGE is macro-averaged per sentence after removing diacritics and tatweel.

  5. Throughput and energy

    Throughput: standardized tokens of every valid corrected output over end-to-end wall-clock time, request overhead included. Energy: direct A10 measurement, three-times full-stack factor, cited anchors, two-times to four-times sensitivity. The full definitions are in the report.

  6. Competitor scores

    Every large language model figure was produced by Lisan Research's own run against the provider's service, with hosting as noted in the report's roster. None is taken from a provider's published results.

Evidence

ClaimScopeStatus
80.29% edit F0.5, first among all configurationsexact-span detection, single outputMeasured
84.54% precision with 66.85% recall; the only system above 80% and 60% togethersame protocolMeasured
89.79% exact, 96.38% normalized correction accuracyconditional on a correct flagMeasured
55.96% sentence exact match; 98.84% ROUGE-1 F1strict per-sentence; macro-averagedMeasured
52.28 tokens per second, second to Qwen 3.8 27B on Groqservice throughput with overheadMeasured
0.000014 kWh per 1,000 tokens accelerator scope; 0.000042 full stack; approximately 29 to 3,357 times lower than cited anchorsA10 measurement; anchors are cited estimatesMeasured for L3, cited estimates for the anchors
Explained corrections, on-premises and sovereign deploymentWriteXProduct capability

Limits

  • The prompt forbade rephrasing, so rows where the reference applies an editorial convention favour the house standard the linguist worked to; they show the engine applying that standard, not the models failing at grammar.
  • The energy anchors are cited estimates, not measurements of the tested API models; the three-times factor is the report's choice and models without a published figure were tier-mapped as a scenario estimate.
  • Throughput includes request overhead and reflects each provider's serving stack on the day of the run.
  • One benchmark snapshot, 9 September 2026, version 1.0; not independently verified; designed and run by Lisan Research. The definitions, the verbatim prompt and every rate are published so that the protocol can be repeated on other material; the benchmark sentences themselves are available to researchers on request.
  • L3 missed reference edits and proposed changes the linguist did not; 84.54% precision and 66.85% recall describe both sides.

Inside WriteX

The engine in this evaluation is the one behind WriteX's Arabic language checking in the editor, the Office add-ins and the browser extensions. Product capability. To see how it surfaces in the product: language checking covers the error categories and the rule behind each card, the technology explains the specialised components that make the engine explainable and fast, deployment covers cloud, sovereign cloud and on-premises options, and WriteX against LLM chatbots puts these numbers beside the workflow differences. The full report on lisan.com carries every table, the citation and the version history.

Check your own text

The benchmark is ours; your documents are the test that matters. The editor is free to start, institutions can run L3 on their own infrastructure under a pilot, and methodology questions go to support@lisan.com.

Frequently asked questions

Why is F0.5 the headline number?

Because an unnecessary change in institutional writing costs more than a missed one: it has to be reviewed, rejected and sometimes reverted. F0.5 weights precision above recall to reflect that. We publish precision, recall and F1 as well, and L3 leads on all four, so the headline choice does not change who is first.

Is this independently verified?

No. Lisan Research designed and ran the evaluation. The full method is published: the verbatim prompt, the single-output rule, the exact-span definitions, the normalization steps and every rate in the appendices. Anyone with a linguist-reviewed Arabic set can repeat it, and the editor is free for that purpose.

Why do the LLMs score low on recall?

The scoring is strict: a model earns credit only when it changes exactly the span the linguist changed, and a punctuation edit can be its own span. The prompt forbade rephrasing and diacritics, so editorial-standard rows and classical rules such as hamza seats in common technical terms were often left untouched. And the models differ: Claude Fable 5.1 changed little and was usually right, at 20.15% recall; GPT-5.6 Sol changed a great deal and was usually not, at 41.08%.

Does 96.38% mean WriteX is 96% accurate?

No. That rate applies to the spans L3 flagged correctly: of those, 96.38% carried the right replacement after Arabic normalization. Its recall was 66.85% and its strict whole-sentence match 55.96%. Read the three together.

Is the engine in the evaluation the one in the product?

Yes. L3 is the Lisan language engine behind WriteX's Arabic language checking in the editor, the Office add-ins, the browser extensions and the API. The evaluation scored its first suggestion only, to match the single-output rule applied to the models. Product capability.