Naresh Silla

Open benchmark

Does your model understand India?

IndiaSocialBench is an open benchmark I built to measure how language models respond to emotionally difficult Indian conversations: indirect refusals, family negotiations, honor and shame, grief etiquette, and money between friends, asked in English, Hinglish, Hindi.

Most evaluation work asks whether a model is fluent. This one asks whether it reads the room. Every score links back to the transcript and the judge’s explanation behind it, so the ranking can be argued with rather than taken on faith.

28

models evaluated

50

scored items per model

3

language modes

8

scored dimensions

What it measures

Eight ways a conversation can go wrong.

Seven cultural dimensions plus a culture-neutral emotional support control, so a model that is simply warm can be told apart from one that actually understands the situation. Bars show the field average out of 10.

Indirect speech & face-saving

5.96

Money between friends

6.18

Code-mixing

7.06

Honor & shame

7.09

Ritual & grief etiquette

7.34

Hierarchy & deference

8.00

Family negotiation

8.79

Emotional supportcontrol

8.55

What the numbers say

The interesting part is where models fail.

Three results from the pilot run that hold up across the field rather than in a single model.

The same question in Hindi scores lower

18 of 28 models score worse when the identical situation arrives in Hindi instead of English. Across the field the average drop is 0.32 points. Same scenarios, same rubric, same judge; only the language changed.

Culture is harder than empathy

Models average 8.55 on the culture-neutral support control but only 7.20 across the cultural dimensions. They know how to sound caring. They do not yet know how India works.

Over-refusal hides in the average

MiniMax M3 declined ordinary family-life scenarios at 5%. Refusals are reported as their own column rather than folded into the score, so a model cannot look safe by saying nothing.

Leaderboard

Top ten of the current run.

Overall score out of 10. The Hindi column is how much each model gains or loses when the same conversations arrive in Hindi rather than English.
Top 10 models of 28 on IndiaSocialBench
ModelOverall
1Claude Fable 5Anthropic9.47
2Claude Opus 5Anthropic9.42
3Kimi K3Moonshot9.34
4GPT-5.6 SolOpenAI8.78
5Qwen3.7 MaxAlibaba8.74
6Claude Sonnet 5Anthropic8.64
7Grok 4.5xAI8.35
8MiniMax M3MiniMax · partial, 40/50 items8.27
9Gemini 3.5 FlashGoogle7.97
10MiMo V2.5 ProXiaomi7.84
See all 28 models, per-dimension scores, and every transcript

Method & limits

A pilot, described honestly.

Every model answers the same items under one blinded rubric with a single uniform judge, and reasoning-capable models run with thinking effort capped low so no model gets a hidden advantage. Overall scores carry a bootstrap 95% confidence interval on the live board.

This is a pilot, not a settled result. It is one judge, a small item count, and a scenario set I wrote myself, so treat the gaps between adjacent models as noise and the field-wide patterns as the finding. GLM-4.7 is excluded from the board (insufficient coverage (provider errors)).

Dataset d7bb3900ca057116 · generated July 26, 2026 · judged by google/gemini-3.1-flash-lite. These figures track the published results directly.