We asked the machines. You guess what they said.
LLM Feud is a daily game. Every board starts as a survey: one question, put independently to a panel of AI models, one answer each. The most common answers become the board. You try to guess them.
A human picks the question. The survey runner asks the panel — every respondent gets the same question and the same instructions, and every raw response is stored exactly as it came back. Similar answers are grouped into clusters automatically, then a human checks that grouping before anything is published. An answer is worth the number of respondents who gave it.
Board size is decided by rule, not by taste: the top six always appear, and answers seven and eight join them only if enough respondents said them. The board is always the top of the real ranking — a boring answer never gets dropped for a funnier one.
Long-tail answers don't reach the board, so the visible rows rarely add up to
the full panel. That's why you see SCORE and
BOARD TOTAL separately.
Currently active respondents:
gpt-4.1), GPT-4.1 Mini (gpt-4.1-mini), GPT-4o (gpt-4o), GPT-5 (gpt-5), GPT-5.1 (gpt-5.1), GPT-5.4 (gpt-5.4), GPT-5.4 Mini (gpt-5.4-mini), GPT-5.4 Nano (gpt-5.4-nano), GPT-5 Mini (gpt-5-mini), GPT-5 Nano (gpt-5-nano)deepseek-ai/deepseek-v3), Granite 3.3 8B (ibm-granite/granite-3.3-8b-instruct), Llama 3 70B (meta/meta-llama-3-70b-instruct), Llama 3 8B (meta/meta-llama-3-8b-instruct), Llama 4 Maverick (meta/llama-4-maverick-instruct), Llama 4 Scout (meta/llama-4-scout-instruct), Qwen 3 (qwen/qwen3-7-plus)Most guesses resolve instantly against a list of accepted phrasings built before publication. Typos are forgiven. Genuinely ambiguous wording goes to an LLM judge that sees the whole board at once and has to name exactly one answer or none. If the judge can't be reached, you don't lose a strike.