Comparison
A chatbot rewrites.
A proofreader accounts.
Pasting a ministry letter into a chatbot trades control for convenience: the text leaves your perimeter and comes back different, with no record of what changed or why.
| WriteX (Relay) | General LLM chatbot | |
|---|---|---|
| Accuracy on language tasksSpecialised models versus a generalist | High | Average |
| Real-time while typingFeedback as the sentence is written | Yes | Paste and wait |
| Explains each correctionCategory plus the rule behind the fix | Yes | If asked, unverifiably |
| Deterministic and auditableSame input, same result, tomorrow too | Yes | No |
| Classifies errors by typeCounted per category, reportable | Yes | No |
| Leaves the rest of the text aloneOnly flagged spans change | Yes | Rewrites freely |
| Data stays in your perimeter | On-premises | Usually cloud |
| Enforces your institution's style guide | Yes | Only via prompting |
| Serving cost at institutional volume | Far lower | GPU-bound |
| Open-ended drafting and reasoningWriting something from nothing | Not its job | Yes |
This compares a proofreading engine with a general assistant. They are different tools, and the point is knowing which job you are doing.
The numbers
The same sentences, 25 model configurations, one language checker.
The row "accuracy on language tasks" above is now a measured result. In version 2.0 of the L3 evaluation, published September 2026, the Lisan language engine behind WriteX and 25 large language model configurations from Anthropic, OpenAI, Google, DeepSeek, Meta and Alibaba Qwen received the same 1,000 linguist-reviewed Arabic sentences, the same strict prompt and one output each. L3 scored 75.67% edit F0.5; GPT-6 Astra, the strongest configuration, scored 65.60%.
The models deserve a fair reading. Claude Fable 5.1 matched L3 on precision, at 78.18%, and its suggestions were usually right, but it found only 30.63% of the reference edits. GPT-6 Astra found the most among the models, 45.92%, at 73.47% precision, and once it found a span its replacement was as reliable as L3's. What no model managed was both at once: L3 found 66.88% of the edits at 78.24% precision, alone in the upper right of the precision-and-recall scatter. The models also led on punctuation and on pronouns that agree with a distant referent, and the report shows those sentences too. The full method, the intervals, the limits and every table are on the results page.
L3 leads on edit F0.5 at 75.67%, against 65.60% for GPT-6 Astra
Edit F0.5, diacritic-insensitive detection on exact source spans, top ten of 26 systems. Higher is better. F0.5 weights precision above recall, because unnecessary changes cost more than missed ones.
The distinction
Institutions need an audit trail, not an oracle.
A chatbot cannot tell you which rule it applied, whether it will apply the same rule tomorrow, or why the rewrite quietly changed a legal nuance. WriteX shows the error, the category, the rule, and the fix, and your linguists can tune all of it.
- A rewrite you cannot diff is a rewrite you cannot approve.
- Pasting a ministry letter into a public chatbot is a data decision, not a writing one.
- "Improve this text" cannot be reported on; error counts per category can.
Morphology rule, stated and auditable: the sound plural is preferred for this pattern in formal writing.
Choose WriteX if
The text has to be correct, explainable, and yours.
Every change is classified, justified by a rule, and confined to the span that was wrong. Nothing leaves your network if you do not want it to, and the numbers roll up for whoever is accountable.
See the technologyReach for a chatbot when
You are drafting from scratch, brainstorming, or summarizing, and no one needs to audit what changed. Many teams draft with an LLM and then run WriteX as the gate.
Settle it yourself
Do not take our word for it.
Run a paragraph through a chatbot and through the checker on this page. One hands back new text; the other hands back a list of errors with the rule behind each one.
Frequently asked questions
Can WriteX and an LLM work together?
Yes, and they often do: teams draft with an LLM, then run WriteX as the compliance and correctness gate before anything is published or sent.
Is WriteX itself an LLM?
No. It is Lisan's own Relay Architecture: 32 specialised language components chained together, which is why it responds in real time, explains each correction, and can run on your own servers.
Chatbots keep improving. Will this gap close?
Their language quality improves, but the structural differences do not: a generative model still rewrites rather than classifies, still cannot promise the same output twice, and still runs on infrastructure you do not control.
Every word your institution writes, checked.
Official letters to social posts, Arabic and English alike. One assistant that proofreads, enforces your style, and explains every correction, wherever your teams write.
12 error types · Institutional style guides · Works in Office, browsers, and the web editor