# DiaryBench

Source: https://brightonxu.com/works/llm-benchmark/

DiaryBench is a private benchmark for one question: how accurately can a model perform fine-grained emotional analysis over a long sequence of diary entries? It is designed to test not only whether a model can identify meaningful patterns, but whether it can do so without under-interpreting or over-interpreting the evidence. Most LLM benchmarks focus on coding or agentic ability, not on the nuanced interpretation required to track how emotion shifts across time. DiaryBench is meant to fill that gap.

## Model performance board

Jaccard / 100 · highest-effort entry per model. API / Web / Harness; unlabelled entries use the highest effort tier.

| Rank | Model | Mean | Min–max | Runs | Rig |
| --- | --- | ---: | --- | ---: | --- |
| 1 | gpt-6-astra | 64.83 | 63.20–66.70 | 3 | harness |
| 2 | gpt-6-sol-eh | 63.06 | 61.70–64.20 | 3 | harness |
| 3 | grok-4.6 | 55.10 | 54.00–57.20 | 5 | harness |
| 4 | glm-5.3 | 54.28 | 52.20–56.00 | 6 | harness |
| 5 | claude-opus-5 | 53.83 | 52.70–55.50 | 3 | harness |
| 6 | gpt-5.6-terra | 53.61 | 53.00–54.30 | 3 | harness |
| 7 | claude-opus-4.6 | 51.61 | 50.70–52.30 | 3 | harness |
| 8 | stealth/ox-alpha | 51.00 | 50.80–51.20 | 3 | api |
| 9 | gpt-5.6-sol | 50.88 | 47.70–55.50 | 4 | harness |
| 10 | gpt-6-luna-eh | 50.61 | 48.00–52.00 | 3 | harness |
| 11 | kimi-k3 | 49.11 | 47.50–52.30 | 3 | harness |
| 12 | qwen3.8-max | 48.28 | 46.70–50.80 | 3 | harness |
| 13 | omen-alpha | 48.22 | 43.20–51.50 | 3 | harness |
| 14 | gpt-5.5-extra-high | 47.56 | 45.70–49.80 | 3 | harness |
| 15 | gemini-3.8-flash | 47.44 | 46.30–48.20 | 3 | web |
| 16 | gemini-3.7-flash | 47.22 | 46.20–48.20 | 3 | web |
| 17 | glm-5.1 | 46.60 | 39.50–49.20 | 5 | api |
| 18 | deepseek-v4-flash | 46.17 | 42.80–51.20 | 3 | api |
| 19 | qwen3.7-plus | 46.00 | 43.00–49.50 | 3 | api |
| 20 | hy3 | 46.00 | 44.50–47.00 | 3 | api |
| 21 | glm-5.2 | 45.31 | 43.50–46.50 | 3 | api |
| 22 | mimo-v2.5 | 45.06 | 42.70–47.70 | 3 | api |
| 23 | gemini-3.6-flash | 44.89 | 41.70–47.00 | 3 | web |
| 24 | minimax-m3 | 44.67 | 43.20–47.30 | 3 | api |
| 25 | gpt-5.6-luna-max | 44.17 | 42.30–45.70 | 3 | harness |
| 26 | qwen3.6-plus | 44.06 | 42.80–44.80 | 3 | api |
| 27 | kimi-k2.5 | 43.47 | 42.80–44.50 | 5 | api |
| 28 | qwen3.5-plus | 43.17 | 43.20–43.20 | 1 | api |
| 29 | gemini-3.1-pro | 42.61 | 40.70–45.80 | 3 | web |
| 30 | kimi-k2.6 | 42.42 | 39.70–46.30 | 4 | api |
| 31 | claude-sonnet-4.5 | 42.11 | 37.50–46.30 | 3 | web |
| 32 | gemini-3.5-flash | 40.83 | 40.30–41.30 | 3 | web |
| 33 | gemini-2.5-pro | 40.56 | 39.00–41.50 | 3 | web |
| 34 | longcat-2.0 | 39.61 | 36.00–42.20 | 3 | web |
| 35 | gemini-3.5-flash-lite | 38.33 | 37.70–38.70 | 3 | web |
| 36 | gpt-5.6-luna-api | 36.61 | 35.80–37.50 | 3 | api |
| 37 | claude-sonnet-5 | 0.00 | 0.00–0.00 | 3 | web |

## What this measures

The source stays private; I am the author and the only ground truth. Each question presents several plausible readings of a moment, but the relevant evidence may span one passage, a period, or years of diary entries. The fields below define the test format, the answer key, and the scoring rule.

- **Long-context reconstruction**: Connect clues across one passage, a period, or years of diary entries.
- **Fine-grained discrimination**: Separate adjacent, plausible accounts of the same feeling.
- **Chinese register and performance**: Read internet-native and performative emotional language in context.
- **Resist over-interpretation**: Do not mistake a depth-sounding mechanism—or the narrator’s own explanation—for what the text supports.

## Test and scoring

- format 50 multi-select questions about my state of mind at specific moments. Each has 1–3 correct options.
- ground truth I decide what was true. The answer key stays private; web search or training data cannot retrieve it.
- scoring Jaccard overlap |S ∩ C| / |S ∪ C| Across 50 questions. S = selected options; C = correct options. Partial credit is built in; extra choices reduce the score.
## Disclosure

Aggregate scores only. GLM-5.3 helped draft the questions; the private answer key never left the author. claude-sonnet-5 declined all three runs and is scored 0 by policy.

[Download aggregate results (CSV)](https://brightonxu.com/works/llm-benchmark/results.csv). Includes effort variants; no source text, questions or answers.

[GitHub](https://github.com/BrightonXX/DiaryBench)
