01
Requested vs measured
For each trajectory instruction, the four generated stories’ measured curves (z-scored, smoothed) are drawn over the idealized target shape (dashed). Agreement is the Spearman ρ between the measured curve and the target; for the “low” instruction, which has no shape, the raw within-story standard deviation is reported instead (lower = flatter).
low · n=4 · mean raw SD 0.77 · mean peak at 70%
a lighthouse keeper on… ρ=0.89, peak 74% · two siblings clear out… ρ=0.71, peak 64% · a night-shift nurse on… ρ=0.60, peak 67% · a botanist tracking an… ρ=0.86, peak 74%
rising · n=4 · mean ρ with target 0.55 · mean peak at 71%
a lighthouse keeper on… ρ=0.61, peak 59% · two siblings clear out… ρ=0.52, peak 74% · a night-shift nurse on… ρ=0.62, peak 74% · a botanist tracking an… ρ=0.46, peak 74%
early peak · n=4 · mean ρ with target -0.25 · mean peak at 71%
a lighthouse keeper on… ρ=-0.34, peak 100% · two siblings clear out… ρ=-0.15, peak 56% · a night-shift nurse on… ρ=-0.17, peak 74% · a botanist tracking an… ρ=-0.34, peak 54%
late climax · n=4 · mean ρ with target 0.44 · mean peak at 85%
a lighthouse keeper on… ρ=0.74, peak 92% · two siblings clear out… ρ=0.33, peak 82% · a night-shift nurse on… ρ=0.44, peak 90% · a botanist tracking an… ρ=0.27, peak 74%
repeated peaks · n=4 · mean ρ with target -0.04 · mean peak at 65%
a lighthouse keeper on… ρ=0.06, peak 64% · two siblings clear out… ρ=-0.33, peak 77% · a night-shift nurse on… ρ=0.27, peak 82% · a botanist tracking an… ρ=-0.14, peak 36%
02
Generated vs human-written dynamics
Two comparisons, on the primary measure. Left: how much the raw rating varies within a story — human stories vs generated stories (a flat generated story has low variance regardless of instruction). Right: the human corpus’ genre curves for reference.
| descriptor | q0.1 | q0.25 | q0.5 | q0.75 | q0.9 |
|---|---|---|---|---|---|
| peak_position | 0.41 | 0.64 | 0.83 | 0.97 | 1.00 |
| peak_height | 1.04 | 1.19 | 1.38 | 1.64 | 1.95 |
| trend_slope | 0.27 | 0.64 | 1.34 | 1.71 | 2.00 |
| n_peaks | 1.10 | 2.00 | 3.00 | 3.00 | 4.00 |
| late_minus_early | -0.04 | 0.42 | 0.86 | 1.18 | 1.33 |
| final_level | -0.31 | 0.05 | 0.46 | 0.98 | 1.41 |
| roughness | 0.20 | 0.24 | 0.27 | 0.30 | 0.33 |
Select one group to see its 95% bootstrap band over stories. Curves are z-scored within story before averaging, so height differences between stories are removed and only shape remains.
03
Caveats specific to this experiment
- The generator (Yuu no Sekai) and the annotator are the same model family; a shared notion of “suspense” could inflate agreement. The lexical proxy is the independent check.
- Twenty stories is a pilot-scale sample; per-trajectory means rest on four stories.
- Generated stories are ~1,400 words; the human corpus spans 1.2k–43k. Comparisons are on normalized progress and within-story variability only.