01
Features that co-move with suspense within a story
Mean within-story Spearman ρ between each feature and LLM suspense (Qwen3-8B) (stories with ≥8 rated bins). Positive means the feature rises when the rating rises. These are associations inside a text, not causes of reader experience.
| feature | family | within-story ρ | share + | pooled ρ (z) | story-level ρ |
|---|---|---|---|---|---|
| sem_arousal_lex | semantic (lexicon) | 0.35 | 97% | 0.35 | 0.55 |
| sem_threat | semantic (lexicon) | 0.31 | 96% | 0.31 | 0.57 |
| sem_action | semantic (lexicon) | 0.25 | 97% | 0.25 | 0.40 |
| nar_verb_rate | narrative (spaCy) | 0.25 | 88% | 0.25 | -0.03 |
| sem_neg_emotion | semantic (lexicon) | 0.23 | 88% | 0.23 | 0.34 |
| sem_valence_lex | semantic (lexicon) | -0.23 | 12% | -0.23 | -0.42 |
| sem_emo_fear | semantic (lexicon) | 0.19 | 81% | 0.18 | 0.42 |
| lex_exclamation_rate | lexical | 0.19 | 89% | 0.15 | 0.10 |
| lex_long_word_rate | lexical | -0.18 | 20% | -0.19 | -0.05 |
| nar_adj_rate | narrative (spaCy) | -0.18 | 20% | -0.18 | 0.04 |
| nar_location_changes | narrative (spaCy) | -0.18 | 21% | -0.13 | 0.22 |
| lex_mean_word_len | lexical | -0.17 | 22% | -0.18 | -0.17 |
| nar_person_introductions | narrative (spaCy) | -0.17 | 18% | -0.14 | 0.13 |
| lex_mean_sentence_len | lexical | -0.16 | 24% | -0.17 | 0.02 |
| nar_motion_verb_rate | narrative (spaCy) | 0.16 | 82% | 0.17 | -0.13 |
| sem_emo_sadness | semantic (lexicon) | 0.16 | 81% | 0.11 | 0.07 |
| nar_location_density | narrative (spaCy) | -0.16 | 23% | -0.13 | -0.14 |
| nar_noun_rate | narrative (spaCy) | -0.15 | 28% | -0.16 | -0.26 |
| nar_entity_density | narrative (spaCy) | -0.14 | 26% | -0.13 | -0.40 |
| mod_surprisal_p90 | model-based | -0.14 | 28% | -0.14 | -0.13 |
| sem_perception | semantic (lexicon) | 0.13 | 74% | 0.14 | 0.14 |
| sem_emo_surprise | semantic (lexicon) | 0.13 | 83% | 0.11 | 0.21 |
| lex_temporal_rate | lexical | 0.13 | 76% | 0.14 | 0.21 |
| nar_finite_verb_rate | narrative (spaCy) | 0.13 | 66% | 0.14 | -0.11 |
| lex_short_sentence_rate | lexical | 0.13 | 78% | 0.14 | 0.07 |
| sem_conflict | semantic (lexicon) | 0.13 | 79% | 0.11 | 0.24 |
| lex_ttr | lexical | -0.13 | 27% | -0.13 | -0.53 |
| sem_emo_anger | semantic (lexicon) | 0.12 | 78% | 0.05 | 0.03 |
| mod_surprisal_sd | model-based | -0.12 | 28% | -0.12 | -0.39 |
| nar_past_progressive_rate | narrative (spaCy) | 0.12 | 74% | 0.13 | -0.06 |
| lex_question_rate | lexical | 0.12 | 74% | 0.11 | 0.03 |
| sem_emo_disgust | semantic (lexicon) | 0.12 | 81% | 0.04 | 0.28 |
| lex_pronoun_i_rate | lexical | 0.11 | 69% | 0.11 | 0.38 |
| mod_surprisal_mean | model-based | -0.10 | 33% | -0.11 | -0.13 |
| lex_dash_rate | lexical | 0.09 | 74% | 0.09 | 0.17 |
| nar_person_distinct | narrative (spaCy) | -0.09 | 33% | -0.08 | 0.06 |
| lex_sd_sentence_len | lexical | -0.08 | 36% | -0.09 | 0.40 |
| nar_person_turnover | narrative (spaCy) | -0.08 | 30% | -0.09 | -0.01 |
| lex_mattr50 | lexical | -0.08 | 32% | -0.08 | -0.03 |
| nar_person_density | narrative (spaCy) | -0.07 | 38% | -0.06 | -0.34 |
| lex_present_tense_hint | lexical | -0.07 | 36% | -0.07 | 0.06 |
| sem_emo_joy | semantic (lexicon) | -0.06 | 39% | -0.06 | 0.01 |
| nar_adv_rate | narrative (spaCy) | 0.06 | 62% | 0.07 | 0.20 |
| lex_pronoun_you_rate | lexical | 0.05 | 59% | 0.05 | -0.02 |
| nar_speech_verb_rate | narrative (spaCy) | 0.05 | 57% | 0.05 | -0.24 |
| mod_surprisal_max | model-based | -0.04 | 41% | -0.04 | 0.16 |
| nar_dialogue_ratio | narrative (spaCy) | 0.04 | 54% | 0.05 | 0.07 |
| lex_negation_rate | lexical | 0.03 | 53% | 0.04 | -0.06 |
| lex_past_tense_hint | lexical | -0.03 | 46% | -0.03 | -0.12 |
| lex_semicolon_rate | lexical | -0.02 | 48% | -0.03 | 0.15 |
| lex_comma_rate | lexical | -0.02 | 45% | -0.02 | 0.32 |
| lex_ellipsis_rate | lexical | 0.01 | 44% | 0.00 | 0.00 |
| nar_subject_shift_rate | narrative (spaCy) | -0.01 | 51% | -0.01 | 0.45 |
| sem_uncertainty | semantic (lexicon) | 0.00 | 52% | 0.01 | 0.08 |
02
Do the operationalizations agree?
Suspense is latent; no single column is ground truth. Mean within-story Spearman ρ between measures (left) and story-level ρ of story means (right). A measure that agrees with the others within stories but not across them is tracking shape, not level.
| LLM suspense (Qwen3-8B) | LLM intensity | Composite lexical proxy | Threat lexicon | Uncertainty | Arousal (lexicon) | Action / motion | Negative emotion | GPT-2 surprisal | |
|---|---|---|---|---|---|---|---|---|---|
| LLM suspense (Qwen3-8B) | 1.00 | 0.85 | 0.40 | 0.31 | 0.06 | 0.35 | 0.28 | 0.23 | -0.10 |
| LLM intensity | 0.85 | 1.00 | 0.42 | 0.33 | 0.02 | 0.40 | 0.29 | 0.26 | -0.07 |
| Composite lexical proxy | 0.40 | 0.42 | 1.00 | 0.69 | 0.30 | 0.84 | 0.57 | 0.34 | -0.10 |
| Threat lexicon | 0.31 | 0.33 | 0.69 | 1.00 | -0.02 | 0.63 | 0.13 | 0.26 | 0.03 |
| Uncertainty | 0.06 | 0.02 | 0.30 | -0.02 | 1.00 | -0.02 | -0.03 | 0.04 | -0.27 |
| Arousal (lexicon) | 0.35 | 0.40 | 0.84 | 0.63 | -0.02 | 1.00 | 0.46 | 0.38 | 0.02 |
| Action / motion | 0.28 | 0.29 | 0.57 | 0.13 | -0.03 | 0.46 | 1.00 | 0.16 | -0.07 |
| Negative emotion | 0.23 | 0.26 | 0.34 | 0.26 | 0.04 | 0.38 | 0.16 | 1.00 | -0.06 |
| GPT-2 surprisal | -0.10 | -0.07 | -0.10 | 0.03 | -0.27 | 0.02 | -0.07 | -0.06 | 1.00 |
| LLM suspense (Qwen3-8B) | LLM intensity | Composite lexical proxy | Threat lexicon | Uncertainty | Arousal (lexicon) | Action / motion | Negative emotion | GPT-2 surprisal | |
|---|---|---|---|---|---|---|---|---|---|
| LLM suspense (Qwen3-8B) | 1.00 | 0.94 | 0.18 | 0.13 | 0.13 | 0.06 | -0.04 | -0.18 | 0.10 |
| LLM intensity | 0.94 | 1.00 | 0.17 | 0.13 | 0.07 | 0.12 | -0.11 | -0.14 | 0.14 |
| Composite lexical proxy | 0.18 | 0.17 | 1.00 | 0.60 | 0.40 | 0.52 | 0.23 | -0.04 | -0.11 |
| Threat lexicon | 0.13 | 0.13 | 0.60 | 1.00 | 0.10 | 0.05 | -0.03 | -0.04 | -0.11 |
| Uncertainty | 0.13 | 0.07 | 0.40 | 0.10 | 1.00 | -0.16 | 0.08 | 0.07 | -0.00 |
| Arousal (lexicon) | 0.06 | 0.12 | 0.52 | 0.05 | -0.16 | 1.00 | -0.10 | -0.04 | -0.04 |
| Action / motion | -0.04 | -0.11 | 0.23 | -0.03 | 0.08 | -0.10 | 1.00 | -0.13 | -0.16 |
| Negative emotion | -0.18 | -0.14 | -0.04 | -0.04 | 0.07 | -0.04 | -0.13 | 1.00 | 0.04 |
| GPT-2 surprisal | 0.10 | 0.14 | -0.11 | -0.11 | -0.00 | -0.04 | -0.16 | 0.04 | 1.00 |
LLM vs composite proxy, by genre
| genre | stories | mean within-story ρ | share positive |
|---|---|---|---|
| adventure | 11 | 0.39 | 100% |
| comic | 11 | 0.31 | 100% |
| detective | 10 | 0.44 | 100% |
| ghost | 13 | 0.43 | 100% |
| horror | 18 | 0.44 | 100% |
| literary | 15 | 0.35 | 100% |
| speculative | 14 | 0.42 | 100% |
03
How reliable is the LLM annotator?
Every rating condition is compared with the main condition (plain prompt, temperature 0, no position information) on the same segments. Position conditions state a narrative position in the prompt: the true one, or a shuffled one. If a shuffled position moves the rating, the model is using position as a prior rather than reading the text.
| model | prompt | temp | position | n | ρ pooled | within-story ρ | exact | shift | |Δ| | pos. slope | r(stated−true, Δ) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Yuu no Sekai | contextual | 0 | none | 560 | 0.82 | 0.75 | 32% | +0.81 | 0.84 | – | – |
| Yuu no Sekai | defined | 0 | none | 560 | 0.95 | 0.92 | 84% | -0.06 | 0.16 | – | – |
| Yuu no Sekai | plain | 0 | shuffled | 560 | 0.90 | 0.85 | 59% | +0.41 | 0.42 | 0.46 | 0.23 |
| Yuu no Sekai | plain | 0 | true | 560 | 0.91 | 0.86 | 57% | +0.42 | 0.44 | 0.32 | – |
| Yuu no Sekai | plain | 0.8 | none | 560 | 0.98 | 0.97 | 94% | -0.05 | 0.06 | – | – |
| Yuu no Sekai | plain | 0.8 | none · s1 | 560 | 0.98 | 0.97 | 94% | -0.05 | 0.06 | – | – |
Two independent samples at the same non-zero temperature agree at ρ = 1.00 (exact agreement 99%, n = 560).
Human raters
No human ratings are pooled yet. The annotation study is live at /annotate; results replace this box once at least a few raters pass the attention and test–retest checks. Until then every “suspense” number on this site is an LLM rating or a lexical proxy, and the human-agreement question (RQ4) is answered only indirectly, by LLM-vs-proxy convergence and by the prior literature cited in the paper.
What the model says it is reacting to
Most frequent words in the model’s free-text “cue” field: with (164), strange (154), about (119), reveals (91), something (89), from (76), begins (69), tension (59), mysterious (59), appears (55), secret (55), approaches (54), door (48), revealed (47), holmes (47), mystery (43), figure (42), into (41), house (39), hidden (36), plan (36), story (35), death (33), dead (32), letter (32).
04
Robustness to segmentation
The same story measured under different segmentations should give the same curve. Curves from each scheme are re-binned to 20 bins and correlated story by story.
| measure | schemes | mean ρ | share ρ > 0.5 | stories |
|---|---|---|---|---|
| Composite lexical proxy | bins20~bins40 | 0.99 | 100% | 92 |
| Composite lexical proxy | bins20~paragraph | 0.92 | 99% | 92 |
| Composite lexical proxy | bins20~window200 | 0.96 | 100% | 92 |
| Composite lexical proxy | bins20~window400 | 0.93 | 100% | 92 |
| Composite lexical proxy | bins40~paragraph | 0.92 | 99% | 92 |
| Composite lexical proxy | bins40~window200 | 0.95 | 100% | 92 |
| Composite lexical proxy | bins40~window400 | 0.93 | 100% | 92 |
| Composite lexical proxy | paragraph~window200 | 0.91 | 100% | 92 |
| Composite lexical proxy | paragraph~window400 | 0.89 | 100% | 92 |
| Composite lexical proxy | window200~window400 | 0.95 | 100% | 92 |
| Threat lexicon | bins20~bins40 | 0.99 | 100% | 92 |
| Threat lexicon | bins20~paragraph | 0.91 | 100% | 92 |
| Threat lexicon | bins20~window200 | 0.95 | 100% | 92 |
| Threat lexicon | bins20~window400 | 0.92 | 100% | 92 |
| Threat lexicon | bins40~paragraph | 0.91 | 100% | 92 |
| Threat lexicon | bins40~window200 | 0.95 | 100% | 92 |
| Threat lexicon | bins40~window400 | 0.92 | 100% | 92 |
| Threat lexicon | paragraph~window200 | 0.90 | 100% | 92 |
| Threat lexicon | paragraph~window400 | 0.87 | 100% | 92 |
| Threat lexicon | window200~window400 | 0.94 | 100% | 92 |
| GPT-2 surprisal | bins20~bins40 | 0.96 | 99% | 92 |