If you were building a website about books, you would build a page for each book. Obviously.
Goodreads did that too. Then it kept going, and that is the interesting part.
We read its robots.txt and found 16 separate sitemap systems, one for each kind of thing on the site. Not sections of a website — different types of noun. Authors. Awards. Genres. Groups. Interviews. Lists. Reading topics. Giveaways. Questions readers asked about a book, and questions readers asked about an author, kept apart as two different things.
And the biggest ones are not books at all.
Results at a glance
Most catalogues have one noun. Goodreads found sixteen, and then built the whole machinery — sitemap, template, URL pattern — for each one.
The single largest page type after authors is the humble quote tag.
16 nouns, not one
Here is the full list Goodreads declares: author, award, blog, genre, giveaway, group, interview, list, quote, quote tag, related work, topic, user, an index, and two separate kinds of community question — one attached to books, one attached to authors.
Behind them sit 527 sitemap files. Sampling the leaves gives a sense of the scale:
| Page type | Sitemap files | Roughly |
|---|---|---|
| Authors | 186 | ~9 million |
| Quote tags | 125 | ~6 million |
| Quotes | 111 | ~5.5 million |
| Topics | 57 | ~2 million |
| Users, questions, lists | 40 | smaller |
Read that table again. A website about books has built roughly six million pages for tags on quotes.
The page nobody would think to build
A quote tag page is exactly what it sounds like. Somebody highlighted a line from a novel, somebody tagged it "scifi", and that tag became a page listing every quote anyone ever tagged the same way.
Goodreads wrote none of this. Readers typed the quotes in and readers applied the tags. The company supplied a template and a URL, and the catalogue built itself.
That is the trick worth stealing. The quote was already on the site, sitting inside a book page where nobody could search for it. Giving it a page of its own did not create new content. It made content that already existed findable.
Small demand, six million times
Now the honest part, because the obvious objection is fair: does anyone actually search for this?
We measured one of the biggest tags and it came back at a few thousand searches a month, with a difficulty of 75 out of 100. That is a small, contested prize.
Which is the whole point. No single quote tag is worth building a page for. Six million of them are. It is the same arithmetic as Zillow's archive of homes nobody can buy: individually trivial, collectively enormous, and only viable because the page costs nothing to produce.
What all those nouns earn
Domain authority of 97 out of 100, and nearly 288,000 websites linking in. That puts Goodreads in the top tier of everything we have measured.
And it makes sense once you see the architecture. A quote is one of the most linkable objects on the internet. People put them in blog posts, essays, newsletters and forum replies, and when they do, they link to wherever the quote lives.
A book page cannot collect that link. A quote page can.
The score it does not get
One number goes the other way, and it is worth reporting.
We ran goodreads.com through our own Website Audit and it came back at 74 out of 100 — one critical issue, four warnings, two notices, 112 checks passed. That is the lowest technical score of any company in this series. Calculator.net, a website with 222 pages, scored 95.
Sixteen page types and tens of millions of URLs is a large surface to keep tidy, and it shows. Whether that is a real cost or an acceptable one depends on how much of the archive Google was ever going to look at closely.
Copy this in an afternoon
You do not need sixteen page types. You almost certainly need more than one. Four steps.
1. List the nouns already inside your pages. Not the sections of your site — the things. A recipe page contains ingredients, techniques, cuisines, occasions and equipment. Each of those is a noun someone searches for.
2. Ask which of them people search for on their own. "Chicken" is a noun people search. "Step 4" is not. This is the filter that stops you generating rubbish.
3. Give the survivors a page, a URL pattern and a sitemap. The page should list everything that shares that noun. If the content already exists inside other pages, you are not writing anything new — you are indexing what you have.
4. Let people tag things, if you have people. Goodreads' six million tag pages exist because readers made them. Tagging is the cheapest content operation there is, and the tags are also the keywords.
What you cannot copy, and what we missed
An audience that types things in. Goodreads has millions of readers who transcribe quotes for free. If nobody is contributing to your site, tags will not appear on their own, and you are back to writing.
Twenty years of accumulation. This did not arrive as a launch. It grew, one page type at a time.
And here is what we could not check:
We sampled, we did not crawl. The page counts come from counting sitemap files and multiplying by the URLs in a sampled leaf. Real totals will differ, and declared URLs are not indexed pages.
There is no separate sitemap for books. Books appear to live in a generic numbered sitemap series rather than in a named entity index, so we cannot give you a book count to compare against the quote count. We would rather say that than guess at it.
No traffic or revenue data. None was available to us and none is claimed.
Before you build a new page type, check whether anyone searches for it. Our Keyword Research tool gives you volume, difficulty and intent for any term, plus the related and long-tail variations — which is exactly how you tell a real noun from one you invented.
- Volume, difficulty, cost per click and intent
- Related terms, long-tail variants and question phrasings
- 3 checks a day, no signup, no card
Frequently asked questions
What are Goodreads' 16 page types?
From its robots.txt on 23 August 2026: author, author community question, award, blog, book community question, genre, giveaway, group, index, interview, list, quote, quote tag, related work, topic and user. Each has its own sitemap index, and together they point at 527 child sitemap files.
How did you estimate the page counts?
By counting the child files in each sitemap index, opening one leaf from each of the largest types, and multiplying. Authors: 186 files at about 48,700 URLs each. Quote tags: 125 files at 50,000, the maximum a sitemap file may hold. Quotes: 111 files at about 49,894. Topics: 57 files at about 34,334. These are estimates from a sample, and declared URLs are not the same as indexed pages.
Why is there no sitemap for books?
We could not find a named one. The index and related work indexes point at a generic numbered sitemap series, which is most likely where book pages live. Because we cannot isolate them, this article does not compare the number of book pages to the number of quote pages, tempting as that comparison would be.
Do quote tag pages actually get searched?
Yes, though modestly per tag. We measured one of the largest tags and it returned a few thousand searches a month at a difficulty of 75 out of 100. The value is not in any one tag; it is that there are roughly six million of them and each costs nothing to produce, since readers supply both the quotes and the tags.
What is Goodreads' backlink profile?
Our Backlink Analyzer returned Domain Rank 97 out of 100, 287,990 referring domains, 68,791,730 dofollow and 17,712,147 nofollow links, 108,839 referring IPs and 47,356 referring subnets on 23 August 2026. Only Booking.com and Canva have matched 97 in this series.
How did it do on a technical audit?
Poorly, relative to the rest of the series. Our Website Audit returned 74 out of 100 for the homepage, with 1 critical issue, 4 warnings, 2 notices and 112 checks passed. For comparison, Calculator.net scored 95 with 222 pages, and wikiHow, Zillow and NerdWallet all scored 88.
Is this just user-generated content?
Partly, and that is the point. The quotes and tags come from readers, but the page types do not. Somebody at Goodreads decided a tag on a quote deserved its own URL, template and sitemap. The contribution is free; the architecture that turns it into findable pages is the work.
Can I check this myself?
Yes, in about ten minutes. Open goodreads.com/robots.txt, list the Sitemap: lines, and notice that each one names a kind of thing rather than a section. Follow a couple of them, count the child files, and open one leaf. That is the entire method behind this article.
Find your second noun
Every company in this series scaled something different. RTINGS scaled measurements. Wise scaled currencies. Zapier scaled pairs and refused to publish most of them. Booking scaled adjectives. Zillow scaled the past. Calculator.net refused to scale at all. wikiHow scaled drawings. NerdWallet scaled named humans. IKEA scaled markets.
Goodreads scaled kinds of thing. It looked at a book page and saw an author, a genre, an award, a list, a discussion, a quote and a tag on that quote — and gave every one of them a URL.
Most websites stop at the first noun. The one they sell. Everything else stays buried inside a page where nobody can search for it.
So look at your most important page and write down every distinct thing mentioned on it. That list is your site's real map, and you are probably only publishing the first line of it.
Then check which of those nouns people actually search for. It takes a minute, and it is the difference between a new page type and a new pile of pages.
Measured on 23 August 2026. Page types come from goodreads.com's robots.txt; counts come from counting sitemap child files and multiplying by the URLs in a sampled leaf, with every sampled figure published in the FAQ above.
Link, keyword and audit figures are free runs of our own tools. Declared sitemap URLs are not indexed pages. No traffic or revenue data for Goodreads was available to us and none is claimed here.