How an audiobook actually gets made
An audiobook arrives as a single unbroken voice, which is exactly what it is not. What you are listening to is a construction: thousands of separate takes, assembled so carefully that the seams disappear, then pushed through a technical specification strict enough to reject the file if a breath is too loud.
The work is invisible on purpose. Here is what is actually happening.
The audio right is a different right
Before anyone records anything, somebody has to own the ability to. Audio is licensed separately from print, and the contract that hands a publisher the right to print a novel does not automatically hand it the right to record one.
This is why an author can have a hardback with one house and an audiobook with another, why some backlist titles simply have no audio edition, and why a bestseller can wait years for one. The book exists; the right was never sold, or was sold to someone who never exercised it. It is the same unbundling of rights that governs everything else in how book publishing works — territory, format, translation and audio are separate goods that happen to share a title.
Those rights are being exercised far more often than they used to be. The Audio Publishers Association put American audiobook revenue at $2.43bn in 2025, up nine per cent on the year, with publishers reporting more than 750,000 active titles — a 43 per cent jump in twelve months. The backlog of unrecorded books is being worked through at speed.
Everything is priced in finished hours
The industry does not count words or pages. It counts the length of the finished recording, and every rate, schedule and budget is quoted per finished hour.
The conversion is remarkably stable, because narration pace is remarkably stable. Read aloud at a natural, unhurried pace and you produce about 140 words a minute, which works out to roughly 9,300 words per finished hour.
| Manuscript | Finished audio | Studio hours behind it |
|---|---|---|
| 40,000 words | about 4.5 hours | 27–31 |
| 80,000 words | about 8.5 hours | 51–60 |
| 120,000 words | about 13 hours | 78–91 |
| 200,000 words | about 21.5 hours | 129–150 |
That third column is the one that surprises people. An experienced narrator needs six to seven hours of work for every finished hour that reaches you; a newcomer can need ten. A standard novel is therefore not a few days in a booth. It is six to eight working weeks of somebody’s life.
Punch and roll, or why there are no retakes
Most solo narration is recorded using a technique called punch and roll, and it explains the strange, seamless quality of the result.
The narrator reads until they make a mistake. Rather than stopping and marking it for later, they immediately roll the recording back a sentence or two, listen to those last words in their headphones, and start speaking again in time — matching their own pace, pitch and breath as the recording punches in over the error.
The correction happens at the moment of the error, in the same voice, in the same room, with the same warmth in the throat. This is why you cannot hear the joins. A take recorded on Thursday cannot be dropped invisibly into a sentence recorded on Monday, because the voice will have moved. Punch and roll never lets the gap open.
It also means the editing is largely done by the time the narrator stands up, which is the entire economic point of the method.
The chain from manuscript to file
Recording is the visible step and roughly a third of the work. The rest is preparation and repair.
| Stage | Who does it | Per finished hour |
|---|---|---|
| Casting | Publisher or rights holder | — |
| Prep: read-through, pronunciation list, character voices | Narrator | 1–2 hours |
| Recording, punch and roll | Narrator, sometimes with a director | 2–3 hours |
| Editing: pickups, breaths, mouth noise | Narrator or engineer | 1–2 hours |
| Proofing against the text | Proof listener | 1–1.5 hours |
| Mastering to spec | Engineer | about 1 hour |
The proofing stage is the one outsiders never guess at. Somebody sits with the manuscript and the audio and follows every word, marking each dropped article, each substituted synonym, each name pronounced two ways in the same chapter. A narrator reading for eight hours will make errors their own ears slid straight past. The proof listener exists because the author’s sentences are the product.
It is close cousin to the work described in how subtitles get made: an invisible craft judged entirely by whether anyone notices it.
The technical floor you never hear
Retailers do not accept a recording because it sounds nice. They accept it because it passes a measurement, and the thresholds are unusually specific.
Audible’s submission standard requires each file to sit between −23dB and −18dB RMS overall, to peak no higher than −3dB, and — the demanding one — to hold a noise floor no louder than −60dB RMS. Files are also required to carry consistent opening and closing credits and to be topped and tailed with a short measure of room tone.
That noise floor is what forces the booth. A quiet spare room is not quiet at −60dB. A refrigerator two rooms away is audible. A passing lorry is a re-record. This single number is the reason narrators build padded closets and record at four in the morning, and it is why “just read it into a phone” produces something a retailer will reject before a human ever hears it.
Who actually gets paid
Two arrangements dominate, and they distribute risk in opposite directions.
A flat per-finished-hour fee pays the narrator on delivery regardless of sales. Rates start in the low hundreds of dollars per finished hour and climb steeply with reputation; an established or well-known voice commands several times the floor. A thirteen-hour novel is therefore a real production cost long before a single copy sells.
Royalty share pays nothing up front and splits the earnings instead. The narrator becomes a co-investor in the title, which suits a first-time author who cannot fund a recording and suits a narrator who believes in the book. It also means an audiobook that does not sell has cost somebody two months of labour and paid them nothing.
The structural asymmetry is familiar from the other performing arts: the person whose voice you actually came for is usually the last to be paid, and paid least reliably. It is the same shape as the split described in music royalties, with fewer intermediaries and the same conclusion.
The synthetic voice question
Machine narration is now good enough for functional text, and retailers have begun accepting it with disclosure. For a manual, a textbook, a reference work, this is straightforwardly useful — those titles were mostly never going to get a human recording at all.
Fiction is a different problem, and not a sentimental one. A narrator makes thousands of interpretive decisions that are not recoverable from the text: which clause carries the irony, where a character’s composure cracks, how long a silence should sit before the next line. The manuscript does not contain that information. The performance is where it is added.
Which returns to the thing worth noticing next time you listen. The seamlessness is not the absence of effort. It is the whole of it, spent on making itself inaudible.
Sources
- ACX audio submission requirements — the RMS, peak and noise-floor thresholds a file must meet to be accepted
- Audio Publishers Association, research surveys — the annual sales survey behind the revenue and active-title figures
- Per finished hour, defined — the industry unit and the studio-hours ratio behind it
