WORKS / 01 PRIVATE EVALUATION
DIARYBENCH
DiaryBench is a private benchmark for one question: how accurately can a model perform fine-grained emotional analysis over a long sequence of diary entries? It is designed to test not only whether a model can identify meaningful patterns, but whether it can do so without under-interpreting or over-interpreting the evidence. Most LLM benchmarks focus on coding or agentic ability, not on the nuanced interpretation required to track how emotion shifts across time. DiaryBench is meant to fill that gap.
TRACK 01 / SET 50
Model performance board
These filters show how much context from my diary a model needs to use. Evidence ranges from one passage, to the same period, to several years, and finally to a later retrospective. Traps marks questions where the wrong options sound psychologically insightful—defense, self-protection, and so on—but do not match what I was actually feeling.
External comparison bars project into the DiaryBench 30–70 window around a shared 50 anchor; raw scores stay visible at right.
- 01gpt-6-astra64.83 ×3 min 63.20 · max 66.70 · 3 shots
- 02gpt-6-sol-eh63.06 ×3 min 61.70 · max 64.20 · 3 shots
- 03grok-4.655.10 ×5 min 54.00 · max 57.20 · 5 shots
- 04glm-5.354.28 ×6 min 52.20 · max 56.00 · 6 shots
- 05claude-opus-553.83 ×3 min 52.70 · max 55.50 · 3 shots
- 06gpt-5.6-terra53.61 ×3 min 53.00 · max 54.30 · 3 shots
- 07claude-opus-4.651.61 ×3 min 50.70 · max 52.30 · 3 shots
- 08stealth/ox-alpha51.00 ×3 min 50.80 · max 51.20 · 3 shots
- 09gpt-5.6-sol50.88 ×4 min 47.70 · max 55.50 · 4 shots
- 10gpt-6-luna-eh50.61 ×3 min 48.00 · max 52.00 · 3 shots
- 11kimi-k349.11 ×3 min 47.50 · max 52.30 · 3 shots
- 12qwen3.8-max48.28 ×3 min 46.70 · max 50.80 · 3 shots
- 13omen-alpha48.22 ×3 min 43.20 · max 51.50 · 3 shots
- 14gpt-5.5-extra-high47.56 ×3 min 45.70 · max 49.80 · 3 shots
- 15gemini-3.8-flash47.44 ×3 min 46.30 · max 48.20 · 3 shots
- 16gemini-3.7-flash47.22 ×3 min 46.20 · max 48.20 · 3 shots
- 17glm-5.146.60 ×5 min 39.50 · max 49.20 · 5 shots
- 18deepseek-v4-flash46.17 ×3 min 42.80 · max 51.20 · 3 shots
- 19qwen3.7-plus46.00 ×3 min 43.00 · max 49.50 · 3 shots
- 20hy346.00 ×3 min 44.50 · max 47.00 · 3 shots
- 21glm-5.245.31 ×3 min 43.50 · max 46.50 · 3 shots
- 22mimo-v2.545.06 ×3 min 42.70 · max 47.70 · 3 shots
- 23gemini-3.6-flash44.89 ×3 min 41.70 · max 47.00 · 3 shots
- 24minimax-m344.67 ×3 min 43.20 · max 47.30 · 3 shots
- 25gpt-5.6-luna-max44.17 ×3 min 42.30 · max 45.70 · 3 shots
- 26qwen3.6-plus44.06 ×3 min 42.80 · max 44.80 · 3 shots
- 27kimi-k2.543.47 ×5 min 42.80 · max 44.50 · 5 shots
- 28qwen3.5-plus43.17 ×1 min 43.20 · max 43.20 · 1 shots
- 29gemini-3.1-pro42.61 ×3 min 40.70 · max 45.80 · 3 shots
- 30kimi-k2.642.42 ×4 min 39.70 · max 46.30 · 4 shots
- 31claude-sonnet-4.542.11 ×3 min 37.50 · max 46.30 · 3 shots
- 32gemini-3.5-flash40.83 ×3 min 40.30 · max 41.30 · 3 shots
- 33gemini-2.5-pro40.56 ×3 min 39.00 · max 41.50 · 3 shots
- 34longcat-2.039.61 ×3 min 36.00 · max 42.20 · 3 shots
- 35gemini-3.5-flash-lite38.33 ×3 min 37.70 · max 38.70 · 3 shots
- 36gpt-5.6-luna-api36.61 ×3 min 35.80 · max 37.50 · 3 shots
- 37claude-sonnet-50.00 ×3 min 0.00 · max 0.00 · 3 shots
WHAT THIS MEASURES
Not positive or negative.
More exact than that.
The source stays private; I am the author and the only ground truth. Each question presents several plausible readings of a moment, but the relevant evidence may span one passage, a period, or years of diary entries. The fields below define the test format, the answer key, and the scoring rule.
- 01
Long-context reconstruction
Connect clues across one passage, a period, or years of diary entries.
- 02
Fine-grained discrimination
Separate adjacent, plausible accounts of the same feeling.
- 03
Chinese register and performance
Read internet-native and performative emotional language in context.
- 04
Resist over-interpretation
Do not mistake a depth-sounding mechanism—or the narrator’s own explanation—for what the text supports.
|S ∩ C| / |S ∪ C|Across 50 questions. S = selected options; C = correct options. Partial credit is built in; extra choices reduce the score.