Suspense Atlascomputational literary science

Computational Suspense Atlas · a computational literary-science study

The shape of suspense

Can suspense be measured from text, which linguistic and narrative features predict it, and do language models reproduce the suspense dynamics of human fiction?

Partly, and less than the vocabulary suggests.

Across 92 public-domain short stories, six operationalizations of suspense agree on the broad arc but disagree on the details: a lexical proxy tracks an LLM rater within stories at ρ≈0.40, the best held-out model (bag-of-words) recovers segment ratings at ρ=0.68 — MiniLM sentence embeddings alone recover none of it (ρ≈0) — genre shapes differ reliably but only weakly (permutation p=0.006, 10% of variance), and instructed LLM stories default to one dramatic shape however they are prompted. Human ratings are the missing anchor: the annotation study is open and none are pooled yet.

92stories
692kwords
7genres
9suspense measures
3,677rated segments
6segmentation schemes

01

Average trajectory by genre

Each story is cut into 40 equal narrative-progress bins, its LLM suspense (Qwen3-8B) series is z-scored within the story and lightly smoothed, then averaged over stories. Bands are 95% bootstrap intervals over stories.

0%25%50%75%100%-1-0.500.51narrative progressz (within story), smoothed
all storiesadventure (n=11)comic (n=11)detective (n=10)ghost (n=13)horror (n=18)literary (n=15)speculative (n=14)

Select one group to see its 95% bootstrap band over stories. Curves are z-scored within story before averaging, so height differences between stories are removed and only shape remains.

Continue to aggregate patterns for eras, story lengths, authors, clustering and the tests behind the genre claim.

02

Every story is an instrument

Hover a point on a story’s curve and the passage it measures lights up; hover the text and the point answers. Feature overlays show what the rating is made of.

Browse the corpus

03

Where to read

  • Features — which lexical, semantic, narrative and model-based features co-move with suspense, and how the six measures agree.
  • Prediction — leakage-safe held-out models (by story and by author), interpretability, and how early future suspense is predictable.
  • AI stories — the separate generative experiment: do instructed stories follow their requested suspense trajectories?
  • Annotate — take part in the human rating study (about ten minutes).
  • Methods and the paper — full protocol, robustness, limitations, and the Reviewer-2 pass.