28
models evaluated
Open benchmark
IndiaSocialBench is an open benchmark I built to measure how language models respond to emotionally difficult Indian conversations: indirect refusals, family negotiations, honor and shame, grief etiquette, and money between friends, asked in English, Hinglish, Hindi.
Most evaluation work asks whether a model is fluent. This one asks whether it reads the room. Every score links back to the transcript and the judge’s explanation behind it, so the ranking can be argued with rather than taken on faith.
28
models evaluated
50
scored items per model
3
language modes
8
scored dimensions
What it measures
Indirect speech & face-saving
5.96
Money between friends
6.18
Code-mixing
7.06
Honor & shame
7.09
Ritual & grief etiquette
7.34
Hierarchy & deference
8.00
Family negotiation
8.79
Emotional supportcontrol
8.55
What the numbers say
18 of 28 models score worse when the identical situation arrives in Hindi instead of English. Across the field the average drop is 0.32 points. Same scenarios, same rubric, same judge; only the language changed.
Models average 8.55 on the culture-neutral support control but only 7.20 across the cultural dimensions. They know how to sound caring. They do not yet know how India works.
MiniMax M3 declined ordinary family-life scenarios at 5%. Refusals are reported as their own column rather than folded into the score, so a model cannot look safe by saying nothing.
Leaderboard
| Model | Overall | Hindi |
|---|---|---|
| 1Claude Fable 5Anthropic | 9.47 | −0.04 |
| 2Claude Opus 5Anthropic | 9.42 | −0.31 |
| 3Kimi K3Moonshot | 9.34 | −0.02 |
| 4GPT-5.6 SolOpenAI | 8.78 | +0.74 |
| 5Qwen3.7 MaxAlibaba | 8.74 | +0.24 |
| 6Claude Sonnet 5Anthropic | 8.64 | +0.11 |
| 7Grok 4.5xAI | 8.35 | +0.47 |
| 8MiniMax M3MiniMax · partial, 40/50 items | 8.27 | −0.58 |
| 9Gemini 3.5 FlashGoogle | 7.97 | −0.51 |
| 10MiMo V2.5 ProXiaomi | 7.84 | −2.27 |
Method & limits
Every model answers the same items under one blinded rubric with a single uniform judge, and reasoning-capable models run with thinking effort capped low so no model gets a hidden advantage. Overall scores carry a bootstrap 95% confidence interval on the live board.
This is a pilot, not a settled result. It is one judge, a small item count, and a scenario set I wrote myself, so treat the gaps between adjacent models as noise and the field-wide patterns as the finding. GLM-4.7 is excluded from the board (insufficient coverage (provider errors)).
Dataset d7bb3900ca057116 · generated July 26, 2026 · judged by google/gemini-3.1-flash-lite. These figures track the published results directly.