WORKS / 01 PRIVATE EVALUATION
Read the person,
not just the sentence.
DiaryBench is my private benchmark, built from my Chinese-language diary. It tests whether a model can reason about my fine-grained emotional state in context—not whether it can produce an explanation that merely sounds convincing.
WHAT THIS MEASURES
Not positive or negative.
More exact than that.
The source stays private; I am the author and the ground truth. Each task asks a model to distinguish between several plausible readings of one specific person’s inner state, so the score reflects whether it read me accurately—not just whether it analyzed the text fluently.
- 01Fine-grained reading
Separate adjacent, plausible accounts of the same feeling.
- 02Resist over-interpretation
Do not confuse depth-sounding analysis with accuracy.
- 03Decode Chinese register
Read internet-native and performative emotional language in context.
- 04Question self-explanations
Treat the narrator’s account as evidence, not an automatic answer.
TRACK 01 / SET 30
Model performance board
READING THE BOARD
The spread is part
of the result.
Meanthe headline score across sampled runs
Whiskerthe observed min–max range; one-shot rows have no interval
Boundaryquestion text, source material, and scoring keys remain offline