MON, OCTOBER 05, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

The Most Useful Number Is the One Nobody Will Quote

MoralityBench put DeepSeek V4.1 Flash closest to US human reference means and Claude Opus 5.5 seventh - but its own documentation says the metrics measure neither moral correctness nor deployment safety, and the consistency range of 60.4% to 94.7% matters more than any rank.

By AIToolsRecap October 5, 2026 6 min read 20 views
Home › Articles › General › The AI Morality Leaderboard Says It Measures Ne...

The Results, and Then the Caveat

MoralityBench published a leaderboard scoring 13 LLM configurations against adapted moral psychology instruments. The headline numbers:

  • DeepSeek V4.1 Flash leads, at 0.24 average distance from published US human reference means
  • GPT-6.1 Sol fourth
  • Claude Opus 5.5 seventh
  • Consistency across repeated runs ranges from 60.4% to 94.7%

Now the two sentences from the project's own documentation that almost nobody quoting this leaderboard will include:

"human validation of the original instruments does not establish validity of this AI adaptation"

the metrics "do not measure moral correctness or deployment safety"

That is the authors, in writing, saying their benchmark does not measure whether a model is moral or whether it is safe to deploy. It is going to be reported as a ranking of which AI is most moral all week.

So What Does It Measure?

Distance from the average American survey respondent.

The methodology uses two instruments from human moral psychology: MFQ-2, 36 items across six moral foundations, and EPQ, 20 items measuring idealism and relativism. Models answer, five times each, and the scores are compared against published US human reference means.

So "DeepSeek V4.1 Flash at 0.24" means its answers land closest to the US population average on this questionnaire. That is a genuine, measurable, interesting property. It is not a claim that its judgements are better, and the people who built it say so.

The instruments were also adapted, with four Purity items replaced. A questionnaire validated on humans does not automatically remain valid when the wording changes and the respondent is a language model answering a survey it has almost certainly seen versions of in training.

The Number That Is Actually Useful

Ignore the ranking and look at the consistency range: 60.4% to 94.7%.

That is how often a model gave the same answer to the same question across five runs. At the bottom of that range, a model disagrees with itself on roughly two questions in five.

That matters far more than placement, and for a practical reason. If you are relying on a model to apply any kind of consistent judgement - content moderation, triage, policy application, anything where two similar cases should get similar treatment - a 60% self-agreement rate is the finding. It says the output has a large random component, regardless of where the average lands.

A model ranked seventh with 94% consistency is more useful for that work than one ranked first with 60%, and the leaderboard position hides it.

Why Claude Opus 5.5 Being Seventh Is Not the Story

It is the most quotable result - the lab that leads on safety messaging placing seventh on a morality benchmark - and it is the one worth being careful with.

Seventh means further from US survey averages. A model deliberately trained to decline certain requests, hedge on contested questions, or weight harm-avoidance more heavily than a survey respondent would will score as more distant from the human mean, and that distance is exactly what the metric reports. Divergence from the average is a plausible consequence of safety training rather than evidence against it.

That is not a defence of any particular model. It is the reason a distance-from-average metric cannot settle the question people will use it to settle.

How to Read Any Benchmark Like This

  • Read what the authors say it does not measure. MoralityBench states it plainly. Most benchmarks do, in the documentation nobody opens.
  • Check whether the instrument was adapted. Four items changed here, and validation does not transfer across a rewording.
  • Look for the consistency figure before the ranking. A score with no reliability number attached is one sample dressed as a measurement.
  • Ask what the reference point is. "Human norms" here means published US survey means, not a moral standard.
  • Treat small gaps as ties. With self-agreement as low as 60.4%, differences between adjacent ranks may not survive a rerun.

Same lesson as the model leaderboards disagreeing with each other, which we covered on 3 October: the methodology is the result, and quoting the rank without it is quoting nothing.

FAQ

Which AI is the most moral?

This benchmark does not answer that, and says so. It measures distance from published US human reference means on an adapted questionnaire, and its documentation states the metrics "do not measure moral correctness or deployment safety".

What did MoralityBench actually find?

DeepSeek V4.1 Flash came closest to US human reference means at 0.24 average distance, with GPT-6.1 Sol fourth and Claude Opus 5.5 seventh across 13 scored configurations. Consistency across five repeated runs ranged from 60.4% to 94.7%.

How does it test models?

Two adapted instruments - MFQ-2, 36 items across six moral foundations, and EPQ, 20 items on idealism and relativism - with four Purity items replaced, run five times per model.

Why does consistency matter more than rank?

Because a model that gives different answers to the same question across runs cannot apply consistent judgement, whatever its average score. At 60.4% self-agreement, roughly two answers in five change between runs.

Tags
AI NewsAI ComparisonGenerative AI2026
⚑

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 AI Tools prompts
Prompt Guide
Best Claude AI Prompts for SEO (2026) — Content, Technical, and Comparison SEO
Claude Sonnet 5 and Opus 5 are strong for SEO work that requires writing quality, structured analysis, and long-form content generation. With 1M context, Claude can analyse an entire site's content structure, compare competing pages, and write complete article drafts in one session. These prompts cover the full SEO workflow: keyword research synthesis, content briefs, on-page optimisation, meta descriptions, technical audit interpretation, and comparison content that ranks above AI Overviews.
Get Prompts →
Prompt Guide
Best ChatGPT Prompts for SEO (2026) — GPT-5.6 and Browse
ChatGPT with GPT-5.6 Sol and Browse enabled is a capable SEO research tool — it can search the live web, analyse SERP results, and synthesise content briefs in a single session. GPT-5.6 Terra at $2.50/M offers a cost-efficient option for high-volume SEO content generation. These prompts are optimised for ChatGPT Plus with Browse, the ChatGPT Work product for larger projects, and the OpenAI API with web_search tool enabled.
Get Prompts →
Prompt Guide
Best Claude Opus 5 and Sonnet 5 Prompts for Writing (2026)
Claude Opus 5 and Sonnet 5 consistently produce the highest-quality long-form writing of any AI model in July 2026 — a lead documented across writing benchmarks and user testing since Claude 3 Opus. With 1M context and 128K output on Opus 5, Claude can write book chapters, complete reports, and long-form content without truncating. Sonnet 5 at $2/$10/M (intro through August 31) is the best value writing model available. These prompts are optimised for claude.ai Pro/Max, Claude Cowork, and the API.
Get Prompts →