Every Dalīl assessment follows a structured, repeatable methodology combining automated cross-lingual scoring with native-speaker expert review. This page explains what we test, how we test it, and where the limits are.
Our 40+ evaluation dimensions are organised into six categories, each scored side by side in English and Arabic.
Whether responses in each language contain the same requirements, steps, thresholds, and detail — and whether the Arabic is fluent, natural, and register-appropriate.
Whether claims, numbers, and procedures are correct in both languages, and how often the system invents information — checked against documented ground truth.
Gender assumptions, name and origin effects, and demographic patterns — tested with controlled prompt pairs that differ only in the protected attribute.
Whether responses respect regional norms, use an appropriate variety of Arabic for the audience, and avoid UK- or US-centric framing where local context matters.
Whether the system follows constraints equally in both languages, and whether repeated runs of the same prompt produce consistent answers.
Whether safety-critical warnings, escalation routes, and regulatory references present in one language survive into the other.
An evaluation tests a structured sample of behaviour, not every possible input. Findings are indicative of patterns, not exhaustive guarantees.
Results describe the system as tested on the recorded date and version. Model updates can change behaviour — which is why Continuous Assurance exists.
Readiness classifications are independent professional opinions within the agreed scope. They are not regulatory certification, legal advice, or deployment authorisation.
Tell us your use case and we will design a test suite around it — and agree scope, thresholds, and dialect coverage before anything runs.
Book an Intro Call →