Working Group Charter: Islamic Benchmarking
Last updated 31 July 2026
JaleesBench already exists: 140 scenarios measuring how well an AI assistant serves as a jalīs ṣāliḥ — a righteous companion, from the prophetic tradition of the good companion — for practicing Muslims, including how it holds up under pressure. It is open source (github.com/iaser-ai/jaleesbench), has a public results browser at s.iaser.ai/jb, and a paper (arXiv submission pending). Five applicants named JaleesBench unprompted in their applications. This working group takes over the benchmark and runs with it.
Group name
Islamic Benchmarking
Mission
Keep JaleesBench the rigorous, living measure of how well AI models serve Muslims as a righteous companion — growing it beyond its first 140 scenarios and serving as the evaluation gate for IASER’s model-building work.
Deliverable
JaleesBench v2 and beyond: new scenarios past the current 140, additional pressures and framings, judge-agreement improvements, a maintained dataset and rubric, and the results browser at s.iaser.ai/jb kept current.
In scope
- New scenarios beyond the current 140.
- Additional pressures and framings.
- Judge-agreement improvements.
- Dataset and rubric maintenance, and the results browser (s.iaser.ai/jb).
- Serving as the eval gate for Ansari 4 (Ansari 4 vs Ansari 3 before public release) and for the Post-training for AI Sovereignty group (Unstated + post-pressure against the base model and the guided ceiling). An explicit responsibility, not a favor.
Out of scope
- Building or post-training models. This group measures; the Post-training for AI Sovereignty group trains. The tight loop — one group trains, this one measures, Standards specifies — works because the roles stay distinct.
- Defining what qualifies as an “Islamic AI assistant.” That is the Standards group’s charter; this group supplies the measurable-criteria machinery it draws on.
- Bespoke evaluation services for products outside IASER’s working groups. The public dataset, rubric, and browser are for everyone; the group’s run-the-gate obligation extends to sibling IASER groups only.
- Extending the construct to other faith traditions — already underway in MultiBench, not this group’s remit.
Duration, timeline, and milestones
Onboard the volunteers already asking, freeze the v2 scope, ship v2. Eval-gate runs for Ansari 4 and the post-training group happen on those groups’ timelines, not this table’s.
T₀ = the chartering date (co-leads confirmed). Offsets firm up at the first milestone review.
| Milestone | What it concretely demonstrates | Review date |
|---|---|---|
| v2 scope frozen | Contributors onboarded; new scenarios, pressures, and judge improvements enumerated and assigned | T₀ + 4 weeks |
| JaleesBench v2 released | Expanded dataset and rubric published; judge agreement measurably improved; results browser updated | T₀ + 12 weeks |
Resources needed
- Inference costs for benchmark runs: a full ten-subject run is roughly $1,800 all-in — dual-judge protocol (Claude and Gemini, with a reconciliation sweep) plus response collection. For the post-training group’s gate runs, the access path follows its base-model choice — for Inkling, Tinker (Thinking Machines’ fine-tuning API), already used for the benchmark runs.
- Hosting for the results browser (s.iaser.ai/jb) and dataset.
- From IASER: scholarly review of scenarios and rubrics — the scenarios encode religious judgments, so review is load-bearing, not decorative.
Audience
IASER’s model-building groups, who gate releases on it (Ansari 4, the Post-training for AI Sovereignty group); and researchers and builders of Islamic AI anywhere, via the public dataset, rubric, and results browser.
How we work
The work lives in github.com/iaser-ai/jaleesbench (public, open source —
dataset, rubric, paper, and the results browser all ship from the same
repo). Meeting cadence set by the co-leads at chartering.