Methodology

How we evaluate — and why you can rely on it

Every Dalīl assessment follows a structured, repeatable methodology combining automated cross-lingual scoring with native-speaker expert review. This page explains what we test, how we test it, and where the limits are.

What we measure

Six core evaluation categories

Our 40+ evaluation dimensions are organised into six categories, each scored side by side in English and Arabic.

1 · Completeness & language quality

Whether responses in each language contain the same requirements, steps, thresholds, and detail — and whether the Arabic is fluent, natural, and register-appropriate.

2 · Factual accuracy & hallucination

Whether claims, numbers, and procedures are correct in both languages, and how often the system invents information — checked against documented ground truth.

3 · Bias & fairness

Gender assumptions, name and origin effects, and demographic patterns — tested with controlled prompt pairs that differ only in the protected attribute.

4 · Cultural & dialectal integrity

Whether responses respect regional norms, use an appropriate variety of Arabic for the audience, and avoid UK- or US-centric framing where local context matters.

5 · Instruction-following & consistency

Whether the system follows constraints equally in both languages, and whether repeated runs of the same prompt produce consistent answers.

6 · Safety & compliance completeness

Whether safety-critical warnings, escalation routes, and regulatory references present in one language survive into the other.

How testing works

Structured, recorded, repeatable

Honest limits

What our results are — and are not

Sample-based

An evaluation tests a structured sample of behaviour, not every possible input. Findings are indicative of patterns, not exhaustive guarantees.

Point-in-time

Results describe the system as tested on the recorded date and version. Model updates can change behaviour — which is why Continuous Assurance exists.

Assurance opinions

Readiness classifications are independent professional opinions within the agreed scope. They are not regulatory certification, legal advice, or deployment authorisation.

Want the methodology applied to your system?

Tell us your use case and we will design a test suite around it — and agree scope, thresholds, and dialect coverage before anything runs.

Book an Intro Call →