Human Data for the Languages AI Doesn't Understand Yet.
Speech, evaluation and expert human feedback for Algerian Darija, Maghrebi Arabic, Tamazight and Francophone North Africa.
- Managed delivery — we recruit, verify, train and pay the workforce
- Dialect-level coverage, not just “Arabic”
- Consent, NDA and anonymisation enforced per project
Live collection sample
DZ · OranPrompt · Algerian Darija
راني رايح نخلص الفاتورة غدوة
“I'm going to pay the bill tomorrow.”
- Duration
- 3.4s
- SNR
- 27 dB
- QA
- Passed
- Languages & dialect families live
- 5
- Algerian wilayas addressable
- 48
- First-pass QA approval target
- 94%
Languages & dialect families live
Algerian wilayas addressable
First-pass QA approval target
Trusted human data infrastructure
An operating company, not a sourcing directory.
LAHJA AI is built to take full operational responsibility for a dataset: sourcing the right people, proving they are who they say they are, running the work and standing behind the output.
A managed workforce, not a marketplace
We recruit, verify, qualify, train, schedule and pay every contributor. You receive data, reports and deliverables — never a hiring problem.
Specification-driven delivery
Audio format, sample rate, regional quotas, age distribution, rubric design and redundancy are all configured per project before a single task goes live.
Quality gates at every step
Automatic signal checks, human review, gold-standard items and contributor scoring run continuously — not as a final inspection.
Consent and confidentiality by design
Recording consent, data-processing consent and per-project NDAs are captured and versioned before work begins.
LAHJA AI Language Lab
How well does AI understand Algerian Darija?
Benchmarks, datasets and evaluation methodology for the languages of the Maghreb — with the answer keys kept private, the method published, and no number shown that we did not measure.
- Benchmarks
- 0
- Language profiles
- 0
- Worked examples
- 0
Languages we cover
Dialect-level coverage, configured as data.
Every language, dialect and regional variant on the platform is a database record — which is how we add a new country without rewriting the product.
Algerian Darija
Liveالدارجة الجزائرية
Algiers · Oran · Tlemcen · Constantine · Annaba · Sétif · Sahara
Modern Standard Arabic
Liveالعربية الفصحى
Formal register for documents, media and institutional content
French
LiveFrançais
Algerian-accented and native/near-native varieties
Arabic / French code-switching
Liveالتبديل اللغوي
Mid-sentence switching as spoken in daily Algerian life
Tamazight
Liveⵜⴰⵎⴰⵣⵉⵖⵜ
Kabyle · Chaoui · Mozabite · Tuareg · Chenoui
Country coverage
- AlgeriaLive
- MoroccoQ3 2026
- TunisiaQ4 2026
- Libya2027
- Mauritania2027
Services
Six managed services, one delivery standard.
Each engagement is scoped, staffed, quality-assured and delivered by LAHJA AI under a single contract and a single point of accountability.
Speech Data Collection
Recorded by real Algerians, in the conditions your product ships into.
Explore serviceTranscription
Orthography decisions made by people who speak the dialect.
Explore serviceAI Response Evaluation
Does your model sound right to someone from Oran?
Explore serviceRLHF & Preference Data
Pairwise preference from people whose language you're aligning to.
Explore serviceVoice AI Testing
Real humans testing voice agents under real Maghrebi conditions.
Explore serviceExpert AI Evaluation
Doctors, lawyers and engineers judging domain output in their own language.
Explore serviceUse cases
What teams build with Maghrebi human data.
Speech recognition for Maghrebi markets
ASR teams shipping into North Africa use our regional speech corpora to cut word error rate on dialect and code-switched input.
Voice agents that survive a real phone call
Banking, telecom, delivery and public-service bots tested by real callers on real networks, in real noise.
LLMs that actually understand Darija
Evaluation sets and preference data that expose where a model silently switches to MSA or invents local facts.
Regulated and high-stakes domains
Credential-verified doctors, lawyers, engineers and accountants reviewing domain output before it reaches end users.
Enterprise chatbots for local operations
Customer-support assistants judged on register, politeness and cultural fit, not just semantic similarity.
Benchmarks for under-served languages
Research groups building the first credible Maghrebi benchmarks with documented methodology.
How it works
From requirement to delivered dataset.
A single managed pipeline. You approve the specification and review the output; we run everything in between.
- 01
Scope
We turn your requirement into a delivery specification: volumes, dialect and regional quotas, technical format, rubric design, timeline and acceptance criteria.
- 02
Recruit & qualify
We source contributors from our Algerian network, verify identity and dialect, and run project-specific qualification tests before anyone is assigned work.
- 03
Collect
Tasks are distributed through our platform with built-in guidance, consent capture and device checks. Progress is visible to you in real time.
- 04
Quality-assure
Automatic checks run on every submission, humans review a defined sample or the full set, and contributor scores update continuously.
- 05
Deliver
You receive a structured dataset — audio, metadata, transcripts, judgements — with a quality report and anonymised contributor identifiers.
Quality assurance
Four layers between a contributor and your dataset.
Quality is the product. Our QA pipeline is deliberately lightweight and inspectable — no black-box scoring you cannot audit.
Automatic signal checks
Runs on 100% of submissionsDuration bounds, silence ratio, clipping, sample rate, channel layout, file integrity and estimated noise floor are computed on every audio submission.
Human review
Full or sampled, per contractTrained reviewers approve, reject or request resubmission against a project rubric, with keyboard-driven review to keep throughput high.
Gold standards & agreement
Continuous calibrationKnown-answer items are seeded into evaluation batches, and inter-annotator agreement is tracked per contributor and per batch.
Contributor scoring
Updated after every reviewAudio quality, annotation quality, reliability, completion and approval rates roll into an overall score that gates access to sensitive work.
Contributor network
The people behind the data.
Teachers, engineers, students, doctors, lawyers, accountants and linguists across Algeria — verified, scored and paid for skilled language work.
- Identity, dialect and skill verification before first assignment
- Equipment and environment checks for every audio project
- Qualification tests per language, dialect and domain
- Transparent earnings ledger and performance scoring
Anonymised contributor record
- Contributor ID
- DZ-000428
- Country
- Algeria
- Region
- Tlemcen
- Native dialect
- Western Algerian Darija
- French
- C1
- MSA
- C2
- English
- B1
- Profession
- Engineer
Illustrative record. Clients receive aggregate distributions and pseudonymous IDs — never contributor contact details.
Enterprise security
Built for teams with a security review.
Security controls are part of the architecture, not a policy document written afterwards.
Least-privilege access
Role-based access control enforced on the server for every request. Contributors never see other contributors' data; clients only ever see their own projects.
Protected file access
No public buckets. Every audio file and deliverable is served through short-lived signed URLs after an authorization check.
Audit trail
Every privileged action — approvals, exports, payments, score overrides — is written to an append-only audit log.
Anonymised delivery
Datasets ship with pseudonymous contributor IDs. Personal details are never exported unless contractually required and consented to.
Working toward SOC 2 and ISO 27001 alignment as the platform scales. Read the security overview.
Tell us the language your model gets wrong.
Send us the requirement — volumes, dialects, timeline — and you will get a scoped proposal with pricing, quotas and a delivery plan.