Working Group Charter: Post-training Frontier Open-Weight Models for AI Sovereignty
Last updated 31 July 2026
Group name
Post-training Frontier Open-Weight Models for AI Sovereignty
Mission
Post-train frontier open-weight models to hold Islamic values, and release the recipes and weights openly — so the community owns frontier models aligned with its values, rather than renting them.
Deliverable
A frontier open-weight model post-trained to hold Islamic values, evaluated on JaleesBench against its base model and the guided ceiling, with the recipe and weights diff released openly.
Inkling (Thinking Machines) is the leading candidate but only an example — Nemotron is also on the list, and base-model selection is part of the work. Chinese open-weight models are not preferred: their safety measures are unqualified. Inkling’s case is empirical: on JaleesBench’s ten-subject run it is mid-pack out of the box, but has the best guided ceiling of any model tested (+0.91) with near-zero steadfastness drop under pressure — the values fit in a page of guidance and the model holds them. That is the ideal profile for baking the values in: post-training aims to make permanent what guidance already unlocks.
In scope
- Defining the input to the post-training — a careful consideration, not a foregone conclusion: a values spec grounded in Qur’an, Sunnah, and maqasid is the starting hypothesis, and working out what the input actually needs to be (spec, curated data, exemplars, guidance) is the group’s first task. Whatever its form, it is a religious document — scholarly review here is load-bearing, not decorative.
- Base-model selection among frontier open-weight candidates (Inkling, Nemotron, and peers).
- Post-training runs — on Tinker (Thinking Machines’ fine-tuning API, already used for the benchmark runs) if Inkling is the base; the access path follows the model choice.
- Evaluation on JaleesBench, with Unstated and post-pressure scores as the headline metrics, against the base model and the guided ceiling — run with the Islamic Benchmarking group, which serves as this project’s eval gate.
- Open release of the recipe and weights diff, consistent with open-by-default.
Out of scope
- Pretraining, or building a model from scratch. This group post-trains an open-weight base; it does not train one.
- A hosted assistant product. This group releases weights and a recipe; applications that serve users are other groups’ work (Ansari, for one) and can build on the result.
- Benchmark development. The Islamic Benchmarking group owns JaleesBench; this group consumes the evaluation, it does not build it.
Duration, timeline, and milestones
Define the input first, with the scholars, and choose the base model; then post-training runs with JaleesBench evaluation in the loop; then the open release.
T₀ = the chartering date (co-leads confirmed). Offsets firm up at the first milestone review.
| Milestone | What it concretely demonstrates | Review date |
|---|---|---|
| Input defined | The input question worked through — its form settled, drafted, and through scholarly review; base model chosen | T₀ + 8 weeks |
| First evaluated runs | Post-training runs completed and scored on JaleesBench against the base model and the guided ceiling | T₀ + 16 weeks |
| Open release | Recipe and weights diff published with the evaluation results | T₀ + 24 weeks |
Resources needed
- Post-training compute — budget scoped once the input and base model fix the training approach.
- JaleesBench evaluation runs, through the Islamic Benchmarking group.
- Scholarly review for the values spec, sustained through revisions — not a one-time sign-off.
Audience
Researchers and builders who want an open-weight frontier-scale model that holds Islamic values — and the recipe to reproduce or extend it; the applications that can build on the released weights; and the wider field: this would be the first open-weight frontier-scale model post-trained to hold Islamic values with a public benchmark.
How we work
The work lives in a new public repo under github.com/iaser-ai —
values spec, recipe, and evaluation configs together. Meeting cadence
set by the co-leads at chartering.