Information Age Consulting Request a briefing

Home / Services / Practice 01

Arabic language processing

Systems that read Modern Standard Arabic accurately and Gulf dialect natively — deployable inside your own network, with no text leaving the country.

Typical clientMinistries, regulators, broadcasters, banks
Engagement6–16 weeks to production
DeploymentOn-premise or Kuwait-region cloud
LanguagesMSA, Kuwaiti, Gulf dialects

The problem

Your text is not the Arabic these tools were trained on

Commercial NLP is trained overwhelmingly on newswire Modern Standard Arabic. Real institutional text is not that. Citizen feedback, social posts, support tickets and internal correspondence are written in dialect, with inconsistent spelling, mixed script, and grammatical particles that do not exist in MSA at all.

The failures are not subtle. Below is a single Kuwaiti sentence run through a standard MSA pipeline and through ours.

TokenStandard MSA pipelineOur analyser
الخدماتnoun, plural ✓noun, plural ✓
صارتverb, past ✓verb, past ✓
وايدunknown token — droppedintensifier adverb, Kuwaiti — carries the sentiment
هالسنةunknown token — droppeddemonstrative + noun, Kuwaiti — resolves the time reference

Two dropped tokens out of six. Both of them the ones that told you how the citizen actually felt and when they meant. Multiply that across a hundred thousand comments and the report you brief your minister on is measuring something other than public opinion.

Capabilities

What we build

Component

Morphological analysisالتحليل الصرفي

Root, pattern, part of speech and diacritisation for every token, including dialectal forms. This is the layer everything else sits on, and the reason our downstream accuracy holds up on real text.

Component

Sentiment analysisتحليل المزاج العام

Polarity and intensity scored against Kuwaiti and wider Gulf usage rather than newswire Arabic, so the posts carrying the strongest opinion are not returned as neutral. Automatic dialect labelling lets you segment an audience by how they write, not only by what they say. This is the layer behind our social media analytics, where it scores public conversation at national scale.

Component

Topic & entity extractionاستخراج الموضوعات والكيانات

The dominant subjects across a body of text over a given period, clustered by theme rather than by hashtag or keyword — and the names of people, entities, laws and places, normalised against your own reference lists so the same ministry is not counted under four spellings.

Component

Retrieval for Arabic archivesالبحث الدلالي

Semantic search and question answering over your own documents, wired to an LLM of your choosing — including models that run entirely inside your data centre.

How we work

Three stages, and you can stop after the first

Stage 01

Sample assessment

You send a representative extract of your text. We run it through our pipeline and return a written assessment: what is extractable, what accuracy to expect, what would need building, and whether an off-the-shelf tool would in fact serve you better.

2 weeks · fixed fee
Stage 02

Pilot on live data

A working system on a bounded slice of your operation — one department, one campaign, one archive — with measured accuracy against a human-annotated benchmark you can audit.

4–6 weeks
Stage 03

Deployment & handover

Production install inside your environment, integration with your existing dashboards, documentation in Arabic and English, and training for the team that will own it after we leave.

6–10 weeks

Start here

Tell us what your text is doing wrong.

Send us a sample — a set of citizen comments, a document archive, a support inbox — and we will come back with a written read on what is achievable, what it would take, and whether you need us at all.

Contact Book a 30-minute briefing →

Send a sample of your text with the form. We reply in writing.