WORKS / 01 PRIVATE EVALUATION

Read the person,
not just the sentence.

DiaryBench is my private benchmark, built from my Chinese-language diary. It tests whether a model can reason about my fine-grained emotional state in context—not whether it can produce an explanation that merely sounds convincing.

WHAT THIS MEASURES

Not positive or negative.
More exact than that.

The source stays private; I am the author and the ground truth. Each task asks a model to distinguish between several plausible readings of one specific person’s inner state, so the score reflects whether it read me accurately—not just whether it analyzed the text fluently.

task fine-grained emotional reasoning
ground truth author-verified
public output aggregate score only
  1. 01
    Fine-grained reading

    Separate adjacent, plausible accounts of the same feeling.

  2. 02
    Resist over-interpretation

    Do not confuse depth-sounding analysis with accuracy.

  3. 03
    Decode Chinese register

    Read internet-native and performative emotional language in context.

  4. 04
    Question self-explanations

    Treat the narrator’s account as evidence, not an automatic answer.

TRACK 01 / SET 30

Model performance board

21 models · higher is better

visible scale / 30–70 · whisker = min–max

FAMILY
SORT
    OpenAI Google Anthropic Open source

    Only aggregate scores are embedded in this public view.

    READING THE BOARD

    The spread is part
    of the result.

    Meanthe headline score across sampled runs

    Whiskerthe observed min–max range; one-shot rows have no interval

    Boundaryquestion text, source material, and scoring keys remain offline