WORKS / 01 PRIVATE EVALUATION

DIARYBENCH

DiaryBench is a private benchmark for one question: how accurately can a model perform fine-grained emotional analysis over a long sequence of diary entries? It is designed to test not only whether a model can identify meaningful patterns, but whether it can do so without under-interpreting or over-interpreting the evidence. Most LLM benchmarks focus on coding or agentic ability, not on the nuanced interpretation required to track how emotion shifts across time. DiaryBench is meant to fill that gap.

TRACK 01 / SET 50

Model performance board

37 models · higher is better

full board · 50 questions · visible scale 30–70

FAMILY
RIG
SORT
COMPARE
EVIDENCE
TRAPS

These filters show how much context from my diary a model needs to use. Evidence ranges from one passage, to the same period, to several years, and finally to a later retrospective. Traps marks questions where the wrong options sound psychologically insightful—defense, self-protection, and so on—but do not match what I was actually feeling.

  1. 01
    gpt-6-astra
    64.83 ×3 min 63.20 · max 66.70 · 3 shots
  2. 02
    gpt-6-sol-eh
    63.06 ×3 min 61.70 · max 64.20 · 3 shots
  3. 03
    grok-4.6
    55.10 ×5 min 54.00 · max 57.20 · 5 shots
  4. 04
    glm-5.3
    54.28 ×6 min 52.20 · max 56.00 · 6 shots
  5. 05
    claude-opus-5
    53.83 ×3 min 52.70 · max 55.50 · 3 shots
  6. 06
    gpt-5.6-terra
    53.61 ×3 min 53.00 · max 54.30 · 3 shots
  7. 07
    claude-opus-4.6
    51.61 ×3 min 50.70 · max 52.30 · 3 shots
  8. 08
    stealth/ox-alpha
    51.00 ×3 min 50.80 · max 51.20 · 3 shots
  9. 09
    gpt-5.6-sol
    50.88 ×4 min 47.70 · max 55.50 · 4 shots
  10. 10
    gpt-6-luna-eh
    50.61 ×3 min 48.00 · max 52.00 · 3 shots
  11. 11
    kimi-k3
    49.11 ×3 min 47.50 · max 52.30 · 3 shots
  12. 12
    qwen3.8-max
    48.28 ×3 min 46.70 · max 50.80 · 3 shots
  13. 13
    omen-alpha
    48.22 ×3 min 43.20 · max 51.50 · 3 shots
  14. 14
    gpt-5.5-extra-high
    47.56 ×3 min 45.70 · max 49.80 · 3 shots
  15. 15
    gemini-3.8-flash
    47.44 ×3 min 46.30 · max 48.20 · 3 shots
  16. 16
    gemini-3.7-flash
    47.22 ×3 min 46.20 · max 48.20 · 3 shots
  17. 17
    glm-5.1
    46.60 ×5 min 39.50 · max 49.20 · 5 shots
  18. 18
    deepseek-v4-flash
    46.17 ×3 min 42.80 · max 51.20 · 3 shots
  19. 19
    qwen3.7-plus
    46.00 ×3 min 43.00 · max 49.50 · 3 shots
  20. 20
    hy3
    46.00 ×3 min 44.50 · max 47.00 · 3 shots
  21. 21
    glm-5.2
    45.31 ×3 min 43.50 · max 46.50 · 3 shots
  22. 22
    mimo-v2.5
    45.06 ×3 min 42.70 · max 47.70 · 3 shots
  23. 23
    gemini-3.6-flash
    44.89 ×3 min 41.70 · max 47.00 · 3 shots
  24. 24
    minimax-m3
    44.67 ×3 min 43.20 · max 47.30 · 3 shots
  25. 25
    gpt-5.6-luna-max
    44.17 ×3 min 42.30 · max 45.70 · 3 shots
  26. 26
    qwen3.6-plus
    44.06 ×3 min 42.80 · max 44.80 · 3 shots
  27. 27
    kimi-k2.5
    43.47 ×5 min 42.80 · max 44.50 · 5 shots
  28. 28
    qwen3.5-plus
    43.17 ×1 min 43.20 · max 43.20 · 1 shots
  29. 29
    gemini-3.1-pro
    42.61 ×3 min 40.70 · max 45.80 · 3 shots
  30. 30
    kimi-k2.6
    42.42 ×4 min 39.70 · max 46.30 · 4 shots
  31. 31
    claude-sonnet-4.5
    42.11 ×3 min 37.50 · max 46.30 · 3 shots
  32. 32
    gemini-3.5-flash
    40.83 ×3 min 40.30 · max 41.30 · 3 shots
  33. 33
    gemini-2.5-pro
    40.56 ×3 min 39.00 · max 41.50 · 3 shots
  34. 34
    longcat-2.0
    39.61 ×3 min 36.00 · max 42.20 · 3 shots
  35. 35
    gemini-3.5-flash-lite
    38.33 ×3 min 37.70 · max 38.70 · 3 shots
  36. 36
    gpt-5.6-luna-api
    36.61 ×3 min 35.80 · max 37.50 · 3 shots
  37. 37
    claude-sonnet-5
    0.00 ×3 min 0.00 · max 0.00 · 3 shots
OpenAI Google Anthropic xAI Open source

WHAT THIS MEASURES

Not positive or negative.
More exact than that.

The source stays private; I am the author and the only ground truth. Each question presents several plausible readings of a moment, but the relevant evidence may span one passage, a period, or years of diary entries. The fields below define the test format, the answer key, and the scoring rule.

  1. 01 Long-context reconstruction

    Connect clues across one passage, a period, or years of diary entries.

  2. 02 Fine-grained discrimination

    Separate adjacent, plausible accounts of the same feeling.

  3. 03 Chinese register and performance

    Read internet-native and performative emotional language in context.

  4. 04 Resist over-interpretation

    Do not mistake a depth-sounding mechanism—or the narrator’s own explanation—for what the text supports.

format 50 multi-select questions about my state of mind at specific moments. Each has 1–3 correct options.
ground truth I decide what was true. The answer key stays private; web search or training data cannot retrieve it.
scoring Jaccard overlap |S ∩ C| / |S ∪ C|

Across 50 questions. S = selected options; C = correct options. Partial credit is built in; extra choices reduce the score.