The final tenth is more action-heavy
The median rises from 54% action around the midpoint to 66% at the ending. Dialogue is the complement.
Screenplay Anatomy Atlas · August 2026
We normalized 2,077 structured screenplays from first scene to last, then measured what changed: action and dialogue balance, scene rhythm, speaking-character arrivals, and who carries the dialogue. These are corpus patterns—not instructions for where your story must turn.
What is actually in here? The included files range from Alien, Chinatown, Fargo, Inception, Moonlight, Parasite, and Whiplash to more than 2,000 other scripts. These are archived screenplay versions—not claims about the final cut or shooting draft.
Three findings first
Every headline below is paired with its distribution later in the atlas. None of the figures identifies a “correct” screenplay or separates produced films from unproduced drafts.
The median rises from 54% action around the midpoint to 66% at the ending. Dialogue is the complement.
Only 22% of the eventual speaking cast appears in the opening tenth; the median reaches 65% by halfway.
As the number of speaking characters grows, the three busiest voices usually account for less of the dialogue. The −0.523 score measures how consistently those two things move in opposite directions; it is not a quality score.
Movement across the draft
Picture each screenplay as ten equal stretches of story text. Action accounts for a median 63.8% in the first stretch, settles near 54.5% around halfway, and reaches 66.1% in the final stretch. The shaded range stays wide: many scripts depart sharply from this contour.
We split every script into ten equal stretches of story text. The line is the middle script; the shaded band contains the middle half.
| Progress | 25th percentile | Median | 75th percentile |
|---|---|---|---|
| 0–10% | 52.6% | 63.8% | 74.3% |
| 10–20% | 43.9% | 55.7% | 66.0% |
| 20–30% | 44.7% | 55.1% | 66.6% |
| 30–40% | 43.6% | 54.7% | 65.8% |
| 40–50% | 43.1% | 54.5% | 65.3% |
| 50–60% | 44.4% | 54.5% | 66.2% |
| 60–70% | 44.5% | 55.7% | 67.3% |
| 70–80% | 44.5% | 56.3% | 68.5% |
| 80–90% | 47.3% | 59.9% | 73.3% |
| 90–100% | 54.0% | 66.1% | 78.5% |
Each bar counts how many scenes fit inside an equal amount of story text. A taller bar means shorter scenes on average—not a longer section.
| Progress | 25th percentile | Median | 75th percentile |
|---|---|---|---|
| 0–10% | 8 | 13 | 17 |
| 10–20% | 7 | 11 | 16 |
| 20–30% | 8 | 11 | 16 |
| 30–40% | 8 | 12 | 17 |
| 40–50% | 8 | 12 | 17 |
| 50–60% | 8 | 12 | 17 |
| 60–70% | 8 | 12 | 18 |
| 70–80% | 9 | 13 | 18 |
| 80–90% | 9 | 14 | 20 |
| 90–100% | 8 | 14 | 20 |
Scene count supplies a second signal. Each bucket contains the same share of narrative words, yet the median ending contains 14 scenes versus 12 near the midpoint. In this representation, the ending is composed of more, shorter scenes. It does not establish faster screen time or editing pace, which require the finished film.
Character arrival
Divide a script into ten equal stretches of story text. In the middle script, only 21.7% of everyone who will eventually speak has appeared after the first stretch. That reaches 65.0% by halfway and 94.4% by the end of the ninth.
Supporting and one-line roles count, so this does not mean protagonists arrive late. It means a blanket “introduce the whole cast immediately” rule does not describe what these scripts actually do.
By each tenth, what share of everyone who will eventually speak has appeared? One-line roles count too.
| Progress | 25th percentile | Median | 75th percentile |
|---|---|---|---|
| 0–10% | 16.1% | 21.7% | 28.1% |
| 10–20% | 28.8% | 35.7% | 44.4% |
| 20–30% | 38.8% | 46.5% | 55.6% |
| 30–40% | 48.5% | 56.0% | 64.6% |
| 40–50% | 57.3% | 65.0% | 73.1% |
| 50–60% | 65.9% | 73.4% | 80.3% |
| 60–70% | 74.1% | 80.6% | 86.6% |
| 70–80% | 82.5% | 87.7% | 92.5% |
| 80–90% | 90.9% | 94.4% | 97.4% |
| 90–100% | 100.0% | 100.0% | 100.0% |
Dialogue center of gravity
In the middle screenplay, the three most talkative characters deliver 57.2% of all dialogue. In scripts with fewer than 20 speaking characters, they carry about 72%; in scripts with 120 or more, that falls to about 41%.
The practical question is not whether a large cast is “wrong.” It is whether the reader can still tell whose choices are driving the scene. The underlying rank score is −0.523; it describes a pattern, not dramatic effectiveness.
Within each cast-size group, what share of all dialogue belongs to its three most talkative characters?
| Speaking characters | Scripts | Median top-three share |
|---|---|---|
| 3–19 | 73 | 71.6% |
| 20–29 | 244 | 65.1% |
| 30–39 | 354 | 62.5% |
| 40–59 | 648 | 58.2% |
| 60–79 | 350 | 52.4% |
| 80–119 | 231 | 49.3% |
| 120+ | 177 | 41.3% |
Corpus anatomy
The center marker is the median: half the included scripts fall above it and half below. The wider bar holds the middle half of scripts. These are word and scene counts, not page counts, because MovieSum does not preserve stable pagination.
Narrative words
Scenes
Typical scene length (words)
Speaking characters
Story held in longest 10% of scenes
Scenes with no dialogue
The last two measures come directly from scene content. In the median script, the longest tenth of scenes carries 38% of all action-and-dialogue words, while 29% of scenes contain no tagged dialogue. They reveal concentration and silence—not whether those choices work.
What a plot summary leaves behind
The median summary keeps only 3.02% of the screenplay’s counted words. That reduction removes performance, staging, rhythm, and most scene-level cause-and-effect. The numbers below help explain what summaries discard; they should not be used to judge screenplay craft.
For scale, the median screenplay contains 20,759 narrative words and its summary contains 656. Longer scripts do not reliably receive proportionally longer summaries.
Unique summary phrases absent from the screenplay
“Absent” means the exact word sequence does not appear in the screenplay. That can be useful paraphrase or an error; this test cannot tell which. The takeaway is simpler: summaries frequently describe events in language the screenplay itself never uses.
Methodology and boundaries
We kept scripts with enough material to compare reliably: at least 20 scenes, 5,000 action-and-dialogue words, three speaking characters, and 20 dialogue blocks.
Each script becomes ten equal stretches of story text, from first scene to last. Every screenplay gets one vote, so a very long script cannot overpower the rest.
This is not a sample of every working screenplay. Word position is not page count or screen time, and a common pattern is not automatically a good choice.
Charts use medians and observed quartiles rather than assuming a bell curve. The downloadable data also includes deterministic 95% bootstrap intervals, rank correlations, and log-normal fit checks.
We inspected 2,200 records, included 2,077, and excluded 123. Exclusion counts can overlap when one file misses more than one threshold.
Analysis version 1.1.0 records the thresholds, random seed, and source fingerprint in the aggregate JSON. No screenplay text, summaries, IMDb identifiers, or title-level measurements are served.
Measure the draft you actually have
Put one draft on these exact distributions with the pacing map—or use the screenplay analyzer for the scene, character, location, and dialogue shape of one draft, or compare two revisions without flattening the screenplay into a generic text diff.