FAQ

Common questions, straight answers

What buyers, risk teams, and technical leads usually ask before commissioning an evaluation.

Which models and systems can you test?

Any system with Arabic–English exposure: foundation models via API (OpenAI, Anthropic, Google, Meta, Mistral and others), vendor products built on them, fine-tuned or open-source models, and RAG or knowledge-base assistants. If we can send it prompts and read its answers, we can evaluate it.

Do we have to share client or production data?

No. The default engagement mode uses synthetic and public domain-representative material only. Redacted-data and client-side modes exist for teams that want testing closer to production — see our Security & Data Handling page.

Can testing run inside our environment?

Yes. Client-side mode runs the evaluation within your infrastructure under your controls. We design the test suite, your team executes it or grants scoped access, and we analyse outputs and report.

Is this a regulatory certification?

No. Dalīl readiness classifications are independent assurance opinions based on the agreed evaluation scope. They are designed to inform procurement, governance, and compliance processes — but they are not certification, legal advice, or deployment authorisation.

How long does an assessment take?

A Readiness Assessment typically takes 5–10 working days from intake; a full Bias & Reliability Audit takes 10–15. The Complimentary Risk Snapshot is delivered within 5 working days, subject to capacity.

What happens if our system is classified as not ready?

You receive the evidence behind every finding, a prioritised remediation roadmap, and the specific conditions that would change the classification. Many clients remediate and re-test within the same quarter.

Can you test internal and fine-tuned models?

Yes — that is what Internal Model Assurance is for: baseline benchmarks for new checkpoints, regression review after fine-tuning or retrieval changes, and release-gate evaluations before material updates ship.

Do you support Arabic dialects?

Yes. Evaluations cover Modern Standard Arabic by default, with Gulf, Levantine, or Egyptian coverage added according to the audience your system serves. Dialect scope is agreed before testing begins.

Can you monitor our system after deployment?

Yes — Continuous Assurance provides scheduled re-evaluation, model update impact assessments, regression alerts, and governance reporting on a quarterly cycle.

A question we haven't answered?

Ask us directly — we respond within two working days.

Get in Touch →