Home / Services / Practice 01
Arabic language processing
Systems that read Modern Standard Arabic accurately and Gulf dialect natively — deployable inside your own network, with no text leaving the country.
The problem
Your text is not the Arabic these tools were trained on
Commercial NLP is trained overwhelmingly on newswire Modern Standard Arabic. Real institutional text is not that. Citizen feedback, social posts, support tickets and internal correspondence are written in dialect, with inconsistent spelling, mixed script, and grammatical particles that do not exist in MSA at all.
The failures are not subtle. Below is a single Kuwaiti sentence run through a standard MSA pipeline and through ours.
| Token | Standard MSA pipeline | Our analyser |
|---|---|---|
| الخدمات | noun, plural ✓ | noun, plural ✓ |
| صارت | verb, past ✓ | verb, past ✓ |
| وايد | unknown token — dropped | intensifier adverb, Kuwaiti — carries the sentiment |
| هالسنة | unknown token — dropped | demonstrative + noun, Kuwaiti — resolves the time reference |
Two dropped tokens out of six. Both of them the ones that told you how the citizen actually felt and when they meant. Multiply that across a hundred thousand comments and the report you brief your minister on is measuring something other than public opinion.
Capabilities
What we build
Spelling & grammar APIمدقق إملائي
A REST endpoint that checks Arabic orthography and morphology, with a browser plug-in for staff and a bulk mode for correcting archives. Already deployed in publishing and academic environments.
Morphological analysisالتحليل الصرفي
Root, pattern, part of speech and diacritisation for every token, including dialectal forms. This is the layer everything else sits on, and the reason our downstream accuracy holds up on real text.
Sentiment & dialect identificationتحليل المشاعر واللهجة
Polarity and intensity scoring calibrated on Kuwaiti and wider Gulf usage, plus automatic dialect labelling so you can segment an audience by how they write, not only by what they say.
Entity & topic extractionاستخراج الكيانات والموضوعات
Names of people, entities, laws and places pulled out of unstructured Arabic and normalised against your own reference lists, so the same ministry is not counted under four spellings.
Retrieval for Arabic archivesالبحث الدلالي
Semantic search and question answering over your own documents, wired to an LLM of your choosing — including models that run entirely inside your data centre.
How we work
Three stages, and you can stop after the first
Sample assessment
You send a representative extract of your text. We run it through our pipeline and return a written assessment: what is extractable, what accuracy to expect, what would need building, and whether an off-the-shelf tool would in fact serve you better.
Pilot on live data
A working system on a bounded slice of your operation — one department, one campaign, one archive — with measured accuracy against a human-annotated benchmark you can audit.
Deployment & handover
Production install inside your environment, integration with your existing dashboards, documentation in Arabic and English, and training for the team that will own it after we leave.
Proof
Where this has already run
Extending a global text-mining platform
We worked with SAS R&D on the Arabic processing inside SAS Text Miner. When a vendor of that size brings in an outside team for a language, it is a reasonable signal.
Read more →National AssemblyDialect-aware public opinion analysis
Our Social Intelligence Analyzer ran as an add-on to Pulsar for parliamentary and broadcast use, and was written up as a published Pulsar case study.
Read more →Kuwait TVAnalytics behind the news bulletin
Outputs from our systems were used to generate trend segments broadcast on national television news.
Read more →Start here
Tell us what your text is doing wrong.
Send us a sample — a set of citizen comments, a document archive, a support inbox — and we will come back with a written read on what is achievable, what it would take, and whether you need us at all.