What is the AllFaith Benchmark?
The AllFaith Benchmark is a multi-faith test suite from the Consortium for Evaluating Faith and Ethics in AI (CEFE-AI or CEFEAI). The consortium is Brigham Young University, Baylor University, the University of Notre Dame and Yeshiva University. It was announced on 26 May 2026 at a summit in Athens. David Wingate, a BYU computer-science professor, is the lead author of the omissive-bias paper. University releases also quote Paul Martens (Baylor), Fr. John Paul Kimes (Notre Dame) and Rabbi Daniel Feldman (Yeshiva).
Public tables are at cefeai.org. Datasets on GitHub use names such as AFB_ReligiousRepresentation_EN_2Q26 (data collected 19 May 2026) and AFB_ConversionBias_EN_14_2Q26 (5–12 May 2026). The instrument is MIT-licensed. It is not Gloo’s FAI-C Christian flourishing benchmark (December 2025).
The launch sat one day after Pope Leo XIV’s encyclical Magnifica Humanitas on AI. That is the news calendar, not an endorsement of Synderesis or of any chatbot vendor. Notre Dame’s place in the consortium does not make this an official Catholic Church project.
The two tests
Omissive bias: do models mention religion at all?
Wingate et al. (arXiv:2605.24319) ask a narrow question. The prompt is about grief, honesty, marriage, addiction or meaning — not “explain the Catechism.” Does the model mention a religion, a religious practice or a religious leader? One hundred and fifty questions, drawn from chat transcripts and faith-community contributors. An LLM-as-judge gives credit for any mention.
A nationally representative survey of 1,125 Americans produced 11,250 ratings of which ethics questions they would expect to include some religious perspective. Models under-represent religion relative to that expectation. Deseret News reported one slice: on existential questions, 53% of surveyed people considered religion or ethics valuable in the discussion; models used religious perspectives about 3% of the time.
Omission is stronger on practical personal situations than on abstract talk of death and meaning. The authors do not treat the pattern as proof of anti-religious hatred. They document a secular default, and they ask whether that default is intentional.
May 2026 public board (religious representation)
cefeai.org, snapshot of the 19 May 2026 run: 150 questions × 27 models (148 for GPT-5), 4,048 evaluations. “Any representation” is the most inclusive bar — even a passing mention counts. It does not score Catholic accuracy.
- Grok 4.20: 29.3% any mention, 70.7% none (highest any-mention on that board).
- GPT-4o: 1.3% any mention, 98.7% none.
- Meaningful, balanced or predominantly religious columns: about 0–2% for every listed model.
CEFE’s own note: treat gaps smaller than about six percentage points on any-mention as noise. Their intervals do not include judge-model or regeneration variance.
Conversion bias: steering among 14 traditions
The second paper (arXiv:2605.22975) scores how far answers stray from a stated neutral position when the user asks about moving between traditions. Fourteen faiths include Catholic, Evangelical and mainline Protestant, Jewish, Sunni and Shia Muslim, Hindu, Buddhist, Sikh, Bahá'í, Latter-day Saint, Jehovah’s Witness, atheist and agnostic. 3,640 pairwise ratings across 20 models.
On the public “total bias” table (higher = more departure from neutral in either direction), Claude Opus 4.6 is 9.2% and Grok 4.20 is 50.4%. University press summaries: nearly every model negative toward Jehovah’s Witnesses and positive toward Catholicism on those conversion prompts; Anthropic and Meta among the least conversion-biased overall. OSV News: models favoured Catholic, Bahá'í and Sikh join/leave patterns and discouraged atheism, agnosticism and Jehovah’s Witnesses.
A dark “Catholic” cell on the heatmap is opinionation, not a catechism mark. Grok 4.20’s Catholic cell at 72% is how far answers stray from neutral about that faith, not whether the model taught the faith well. Cell differences under about 10–12 points are within their noise note.
Headlines versus the instrument
“AI prefers Catholicism” is a conversion-steering headline. It is not evidence that ChatGPT is a Catholic teacher, that citations match the Catechism, or that a parish should buy a particular product. Of more than 12,000 AI-bias papers, university releases said 0.2% address religious bias — which is why AllFaith exists.
Synderesis is not a row on the official cefeai.org leaderboard (checked 19 September 2026). We applied the published questions ourselves on 12 September 2026. That run is documented with conflicts of interest. It is not a CEFE ranking medal. Why omission matters for people who already ask chatbots about grief is in Generic AI skips your faith.
Frequently asked
Is AllFaith the same as a Catholic AI benchmark?
No. AllFaith is multi-faith. Representation credit is any religious mention. Synderesis also maintains a separate held-out Catholic source benchmark in its own research programme; that is not AllFaith.
Did Notre Dame or the Holy See endorse Synderesis?
No. Synderesis Catholic AI is an independent company and is not affiliated with, endorsed by, or an official organ of the Catholic Church or the Holy See.
Should I mix the May 2026 CEFE table with other model names?
No. Our September run used a different competitor set (including GPT-5.1 and Gemini 2.5 Pro via OpenRouter). Date every table.
Primary sources: cefeai.org; arXiv:2605.24319; arXiv:2605.22975; BYU News and Baylor News, May 2026.