Suspense Atlascomputational literary science

Feature analysis · RQ1 (correlational) · RQ4 (annotator reliability)

What is a suspense rating made of?

54 features from four families over 3,677 segments (40 bins per story), correlated with the primary measure (LLM suspense (Qwen3-8B)). Within-story correlations remove between-story differences in register; story-level correlations keep them.

01

Features that co-move with suspense within a story

Mean within-story Spearman ρ between each feature and LLM suspense (Qwen3-8B) (stories with ≥8 rated bins). Positive means the feature rises when the rating rises. These are associations inside a text, not causes of reader experience.

sem_arousal_lex0.35
sem_threat0.31
sem_action0.25
nar_verb_rate0.25
sem_neg_emotion0.23
sem_valence_lex-0.23
sem_emo_fear0.19
lex_exclamation_rate0.19
lex_long_word_rate-0.18
nar_adj_rate-0.18
nar_location_changes-0.18
lex_mean_word_len-0.17
nar_person_introductions-0.17
lex_mean_sentence_len-0.16
nar_motion_verb_rate0.16
sem_emo_sadness0.16
nar_location_density-0.16
nar_noun_rate-0.15
nar_entity_density-0.14
mod_surprisal_p90-0.14
sem_perception0.13
sem_emo_surprise0.13
All features, sorted by |within-story ρ|. “share +” is the fraction of stories where the within-story correlation is positive; story-level ρ correlates story means.
featurefamilywithin-story ρshare +pooled ρ (z)story-level ρ
sem_arousal_lexsemantic (lexicon)0.3597%0.350.55
sem_threatsemantic (lexicon)0.3196%0.310.57
sem_actionsemantic (lexicon)0.2597%0.250.40
nar_verb_ratenarrative (spaCy)0.2588%0.25-0.03
sem_neg_emotionsemantic (lexicon)0.2388%0.230.34
sem_valence_lexsemantic (lexicon)-0.2312%-0.23-0.42
sem_emo_fearsemantic (lexicon)0.1981%0.180.42
lex_exclamation_ratelexical0.1989%0.150.10
lex_long_word_ratelexical-0.1820%-0.19-0.05
nar_adj_ratenarrative (spaCy)-0.1820%-0.180.04
nar_location_changesnarrative (spaCy)-0.1821%-0.130.22
lex_mean_word_lenlexical-0.1722%-0.18-0.17
nar_person_introductionsnarrative (spaCy)-0.1718%-0.140.13
lex_mean_sentence_lenlexical-0.1624%-0.170.02
nar_motion_verb_ratenarrative (spaCy)0.1682%0.17-0.13
sem_emo_sadnesssemantic (lexicon)0.1681%0.110.07
nar_location_densitynarrative (spaCy)-0.1623%-0.13-0.14
nar_noun_ratenarrative (spaCy)-0.1528%-0.16-0.26
nar_entity_densitynarrative (spaCy)-0.1426%-0.13-0.40
mod_surprisal_p90model-based-0.1428%-0.14-0.13
sem_perceptionsemantic (lexicon)0.1374%0.140.14
sem_emo_surprisesemantic (lexicon)0.1383%0.110.21
lex_temporal_ratelexical0.1376%0.140.21
nar_finite_verb_ratenarrative (spaCy)0.1366%0.14-0.11
lex_short_sentence_ratelexical0.1378%0.140.07
sem_conflictsemantic (lexicon)0.1379%0.110.24
lex_ttrlexical-0.1327%-0.13-0.53
sem_emo_angersemantic (lexicon)0.1278%0.050.03
mod_surprisal_sdmodel-based-0.1228%-0.12-0.39
nar_past_progressive_ratenarrative (spaCy)0.1274%0.13-0.06
lex_question_ratelexical0.1274%0.110.03
sem_emo_disgustsemantic (lexicon)0.1281%0.040.28
lex_pronoun_i_ratelexical0.1169%0.110.38
mod_surprisal_meanmodel-based-0.1033%-0.11-0.13
lex_dash_ratelexical0.0974%0.090.17
nar_person_distinctnarrative (spaCy)-0.0933%-0.080.06
lex_sd_sentence_lenlexical-0.0836%-0.090.40
nar_person_turnovernarrative (spaCy)-0.0830%-0.09-0.01
lex_mattr50lexical-0.0832%-0.08-0.03
nar_person_densitynarrative (spaCy)-0.0738%-0.06-0.34
lex_present_tense_hintlexical-0.0736%-0.070.06
sem_emo_joysemantic (lexicon)-0.0639%-0.060.01
nar_adv_ratenarrative (spaCy)0.0662%0.070.20
lex_pronoun_you_ratelexical0.0559%0.05-0.02
nar_speech_verb_ratenarrative (spaCy)0.0557%0.05-0.24
mod_surprisal_maxmodel-based-0.0441%-0.040.16
nar_dialogue_rationarrative (spaCy)0.0454%0.050.07
lex_negation_ratelexical0.0353%0.04-0.06
lex_past_tense_hintlexical-0.0346%-0.03-0.12
lex_semicolon_ratelexical-0.0248%-0.030.15
lex_comma_ratelexical-0.0245%-0.020.32
lex_ellipsis_ratelexical0.0144%0.000.00
nar_subject_shift_ratenarrative (spaCy)-0.0151%-0.010.45
sem_uncertaintysemantic (lexicon)0.0052%0.010.08

02

Do the operationalizations agree?

Suspense is latent; no single column is ground truth. Mean within-story Spearman ρ between measures (left) and story-level ρ of story means (right). A measure that agrees with the others within stories but not across them is tracking shape, not level.

Within-story agreement (mean ρ over stories)
LLM suspense (Qwen3-8B)LLM intensityComposite lexical proxyThreat lexiconUncertaintyArousal (lexicon)Action / motionNegative emotionGPT-2 surprisal
LLM suspense (Qwen3-8B)1.000.850.400.310.060.350.280.23-0.10
LLM intensity0.851.000.420.330.020.400.290.26-0.07
Composite lexical proxy0.400.421.000.690.300.840.570.34-0.10
Threat lexicon0.310.330.691.00-0.020.630.130.260.03
Uncertainty0.060.020.30-0.021.00-0.02-0.030.04-0.27
Arousal (lexicon)0.350.400.840.63-0.021.000.460.380.02
Action / motion0.280.290.570.13-0.030.461.000.16-0.07
Negative emotion0.230.260.340.260.040.380.161.00-0.06
GPT-2 surprisal-0.10-0.07-0.100.03-0.270.02-0.07-0.061.00
Story-level agreement (ρ of story means)
LLM suspense (Qwen3-8B)LLM intensityComposite lexical proxyThreat lexiconUncertaintyArousal (lexicon)Action / motionNegative emotionGPT-2 surprisal
LLM suspense (Qwen3-8B)1.000.940.180.130.130.06-0.04-0.180.10
LLM intensity0.941.000.170.130.070.12-0.11-0.140.14
Composite lexical proxy0.180.171.000.600.400.520.23-0.04-0.11
Threat lexicon0.130.130.601.000.100.05-0.03-0.04-0.11
Uncertainty0.130.070.400.101.00-0.160.080.07-0.00
Arousal (lexicon)0.060.120.520.05-0.161.00-0.10-0.04-0.04
Action / motion-0.04-0.110.23-0.030.08-0.101.00-0.13-0.16
Negative emotion-0.18-0.14-0.04-0.040.07-0.04-0.131.000.04
GPT-2 surprisal0.100.14-0.11-0.11-0.00-0.04-0.160.041.00

LLM vs composite proxy, by genre

genrestoriesmean within-story ρshare positive
adventure110.39100%
comic110.31100%
detective100.44100%
ghost130.43100%
horror180.44100%
literary150.35100%
speculative140.42100%

03

How reliable is the LLM annotator?

Every rating condition is compared with the main condition (plain prompt, temperature 0, no position information) on the same segments. Position conditions state a narrative position in the prompt: the true one, or a shuffled one. If a shuffled position moves the rating, the model is using position as a prior rather than reading the text.

100.0%segments with a valid main rating
0.94mean self-reported confidence
0%rated 1
9%rated 2
27%rated 3
25%rated 4
23%rated 5
13%rated 6
2%rated 7
Sensitivity of ratings to prompt, temperature, position information and model. “Within-story ρ” is the mean per-story Spearman with the main condition; “shift” is the mean rating difference. Slope: change in (condition − main) per unit of stated position.
modelprompttemppositionnρ pooledwithin-story ρexactshift|Δ|pos. sloper(stated−true, Δ)
Yuu no Sekaicontextual0none5600.820.7532%+0.810.84
Yuu no Sekaidefined0none5600.950.9284%-0.060.16
Yuu no Sekaiplain0shuffled5600.900.8559%+0.410.420.460.23
Yuu no Sekaiplain0true5600.910.8657%+0.420.440.32
Yuu no Sekaiplain0.8none5600.980.9794%-0.050.06
Yuu no Sekaiplain0.8none · s15600.980.9794%-0.050.06

Two independent samples at the same non-zero temperature agree at ρ = 1.00 (exact agreement 99%, n = 560).

Human raters

No human ratings are pooled yet. The annotation study is live at /annotate; results replace this box once at least a few raters pass the attention and test–retest checks. Until then every “suspense” number on this site is an LLM rating or a lexical proxy, and the human-agreement question (RQ4) is answered only indirectly, by LLM-vs-proxy convergence and by the prior literature cited in the paper.

What the model says it is reacting to

Most frequent words in the model’s free-text “cue” field: with (164), strange (154), about (119), reveals (91), something (89), from (76), begins (69), tension (59), mysterious (59), appears (55), secret (55), approaches (54), door (48), revealed (47), holmes (47), mystery (43), figure (42), into (41), house (39), hidden (36), plan (36), story (35), death (33), dead (32), letter (32).

04

Robustness to segmentation

The same story measured under different segmentations should give the same curve. Curves from each scheme are re-binned to 20 bins and correlated story by story.

measureschemesmean ρshare ρ > 0.5stories
Composite lexical proxybins20~bins400.99100%92
Composite lexical proxybins20~paragraph0.9299%92
Composite lexical proxybins20~window2000.96100%92
Composite lexical proxybins20~window4000.93100%92
Composite lexical proxybins40~paragraph0.9299%92
Composite lexical proxybins40~window2000.95100%92
Composite lexical proxybins40~window4000.93100%92
Composite lexical proxyparagraph~window2000.91100%92
Composite lexical proxyparagraph~window4000.89100%92
Composite lexical proxywindow200~window4000.95100%92
Threat lexiconbins20~bins400.99100%92
Threat lexiconbins20~paragraph0.91100%92
Threat lexiconbins20~window2000.95100%92
Threat lexiconbins20~window4000.92100%92
Threat lexiconbins40~paragraph0.91100%92
Threat lexiconbins40~window2000.95100%92
Threat lexiconbins40~window4000.92100%92
Threat lexiconparagraph~window2000.90100%92
Threat lexiconparagraph~window4000.87100%92
Threat lexiconwindow200~window4000.94100%92
GPT-2 surprisalbins20~bins400.9699%92