Skip to content
Algeria live · Maghreb expanding

Human Data for the Languages AI Doesn't Understand Yet.

Speech, evaluation and expert human feedback for Algerian Darija, Maghrebi Arabic, Tamazight and Francophone North Africa.

  • Managed delivery — we recruit, verify, train and pay the workforce
  • Dialect-level coverage, not just “Arabic”
  • Consent, NDA and anonymisation enforced per project

Live collection sample

DZ · Oran

Prompt · Algerian Darija

راني رايح نخلص الفاتورة غدوة

“I'm going to pay the bill tomorrow.”

Duration
3.4s
SNR
27 dB
QA
Passed
Languages & dialect families live
5

Languages & dialect families live

Algerian wilayas addressable
48

Algerian wilayas addressable

First-pass QA approval target
94%

First-pass QA approval target

Trusted human data infrastructure

An operating company, not a sourcing directory.

LAHJA AI is built to take full operational responsibility for a dataset: sourcing the right people, proving they are who they say they are, running the work and standing behind the output.

A managed workforce, not a marketplace

We recruit, verify, qualify, train, schedule and pay every contributor. You receive data, reports and deliverables — never a hiring problem.

Specification-driven delivery

Audio format, sample rate, regional quotas, age distribution, rubric design and redundancy are all configured per project before a single task goes live.

Quality gates at every step

Automatic signal checks, human review, gold-standard items and contributor scoring run continuously — not as a final inspection.

Consent and confidentiality by design

Recording consent, data-processing consent and per-project NDAs are captured and versioned before work begins.

LAHJA AI Language Lab

How well does AI understand Algerian Darija?

Benchmarks, datasets and evaluation methodology for the languages of the Maghreb — with the answer keys kept private, the method published, and no number shown that we did not measure.

Benchmarks
0
Language profiles
0
Worked examples
0

Languages we cover

Dialect-level coverage, configured as data.

Every language, dialect and regional variant on the platform is a database record — which is how we add a new country without rewriting the product.

Full language catalogue

Algerian Darija

Live

الدارجة الجزائرية

Algiers · Oran · Tlemcen · Constantine · Annaba · Sétif · Sahara

Modern Standard Arabic

Live

العربية الفصحى

Formal register for documents, media and institutional content

French

Live

Français

Algerian-accented and native/near-native varieties

Arabic / French code-switching

Live

التبديل اللغوي

Mid-sentence switching as spoken in daily Algerian life

Tamazight

Live

ⵜⴰⵎⴰⵣⵉⵖⵜ

Kabyle · Chaoui · Mozabite · Tuareg · Chenoui

Country coverage

  • AlgeriaLive
  • MoroccoQ3 2026
  • TunisiaQ4 2026
  • Libya2027
  • Mauritania2027

Use cases

What teams build with Maghrebi human data.

Speech recognition for Maghrebi markets

ASR teams shipping into North Africa use our regional speech corpora to cut word error rate on dialect and code-switched input.

Voice agents that survive a real phone call

Banking, telecom, delivery and public-service bots tested by real callers on real networks, in real noise.

LLMs that actually understand Darija

Evaluation sets and preference data that expose where a model silently switches to MSA or invents local facts.

Regulated and high-stakes domains

Credential-verified doctors, lawyers, engineers and accountants reviewing domain output before it reaches end users.

Enterprise chatbots for local operations

Customer-support assistants judged on register, politeness and cultural fit, not just semantic similarity.

Benchmarks for under-served languages

Research groups building the first credible Maghrebi benchmarks with documented methodology.

How it works

From requirement to delivered dataset.

A single managed pipeline. You approve the specification and review the output; we run everything in between.

  1. 01

    Scope

    We turn your requirement into a delivery specification: volumes, dialect and regional quotas, technical format, rubric design, timeline and acceptance criteria.

  2. 02

    Recruit & qualify

    We source contributors from our Algerian network, verify identity and dialect, and run project-specific qualification tests before anyone is assigned work.

  3. 03

    Collect

    Tasks are distributed through our platform with built-in guidance, consent capture and device checks. Progress is visible to you in real time.

  4. 04

    Quality-assure

    Automatic checks run on every submission, humans review a defined sample or the full set, and contributor scores update continuously.

  5. 05

    Deliver

    You receive a structured dataset — audio, metadata, transcripts, judgements — with a quality report and anonymised contributor identifiers.

Quality assurance

Four layers between a contributor and your dataset.

Quality is the product. Our QA pipeline is deliberately lightweight and inspectable — no black-box scoring you cannot audit.

1

Automatic signal checks

Runs on 100% of submissions

Duration bounds, silence ratio, clipping, sample rate, channel layout, file integrity and estimated noise floor are computed on every audio submission.

2

Human review

Full or sampled, per contract

Trained reviewers approve, reject or request resubmission against a project rubric, with keyboard-driven review to keep throughput high.

3

Gold standards & agreement

Continuous calibration

Known-answer items are seeded into evaluation batches, and inter-annotator agreement is tracked per contributor and per batch.

4

Contributor scoring

Updated after every review

Audio quality, annotation quality, reliability, completion and approval rates roll into an overall score that gates access to sensitive work.

Contributor network

The people behind the data.

Teachers, engineers, students, doctors, lawyers, accountants and linguists across Algeria — verified, scored and paid for skilled language work.

  • Identity, dialect and skill verification before first assignment
  • Equipment and environment checks for every audio project
  • Qualification tests per language, dialect and domain
  • Transparent earnings ledger and performance scoring

Anonymised contributor record

Contributor ID
DZ-000428
Country
Algeria
Region
Tlemcen
Native dialect
Western Algerian Darija
French
C1
MSA
C2
English
B1
Profession
Engineer
Audio quality92
Annotation quality89
Reliability96
Overall score92

Illustrative record. Clients receive aggregate distributions and pseudonymous IDs — never contributor contact details.

Enterprise security

Built for teams with a security review.

Security controls are part of the architecture, not a policy document written afterwards.

Least-privilege access

Role-based access control enforced on the server for every request. Contributors never see other contributors' data; clients only ever see their own projects.

Protected file access

No public buckets. Every audio file and deliverable is served through short-lived signed URLs after an authorization check.

Audit trail

Every privileged action — approvals, exports, payments, score overrides — is written to an append-only audit log.

Anonymised delivery

Datasets ship with pseudonymous contributor IDs. Personal details are never exported unless contractually required and consented to.

Working toward SOC 2 and ISO 27001 alignment as the platform scales. Read the security overview.

Tell us the language your model gets wrong.

Send us the requirement — volumes, dialects, timeline — and you will get a scoped proposal with pricing, quotas and a delivery plan.