A content inventory is a structured catalog of every URL on your site — title, type, owner, dates, traffic, and status — assembled in one place so you can see what you actually have before you decide what to do with it. It’s the raw dataset every serious content project runs on: migrations, redesigns, pruning, and SEO refreshes all start here. We never start an SEO content audit without one, because you can’t fix a site you haven’t fully mapped.
Content Inventory
A content inventory is a complete, itemized list of all content assets on a website — pages, posts, PDFs, media — capturing identifiers and attributes (URL, title, type, author, last-modified date, status, word count, traffic) so the content can be managed, migrated, or optimized.
Why a Content Inventory Comes First
Most teams skip the inventory and jump straight to opinions: “this post feels old,” “we should add a landing page.” That’s how you end up with 1,400 indexed URLs, no idea which earn traffic, and three pages fighting over the same query. An inventory replaces vibes with a dataset.
The point isn’t bureaucracy — it’s leverage. Once every URL sits in one sheet next to its traffic, last-updated date, and indexation status, decisions get fast and obvious. You can see thin content at a glance, spot orphaned content with zero internal links, and find the dead weight eating your crawl budget.
What a good inventory buys you:
- Visibility. Every asset surfaces — including the forgotten 2019 pages no one remembers publishing.
- Prioritization. High-traffic pages, duplicates, and zombie URLs separate themselves once the data is side by side.
- Migration safety. It becomes the single source of truth for redesigns and CMS moves, so nothing gets lost and redirects get planned, not improvised.
- SEO triage. Missing metadata, redirect chains, and unindexed pages become a punch list instead of a mystery.
- Governance. Ownership, status, and editorial workflow get documented instead of living in someone’s head.
No dashboard theater here: an inventory is just a clean spreadsheet (or database) that turns “the site feels messy” into a list of URLs you can actually act on.
Content Inventory vs. Content Audit
These two get used interchangeably, and that confusion costs teams real time. The inventory is the what. The audit is the so what. You build the inventory first — objective, exhaustive — then layer the audit’s judgment on top. A marketing audit or content audit without an inventory underneath it is just guesswork with a deck.
| Dimension | Content Inventory | Content Audit |
|---|---|---|
| Nature | Objective snapshot | Qualitative judgment |
| Question | What do we have? | What should we do about it? |
| Output | Spreadsheet of every URL + attributes | Scorecards, gap analysis, action plan |
| Who runs it | Content manager, dev, analyst | SEO, content strategist, UX, stakeholders |
| When | First step of any content project | After the inventory, when you need decisions |
| Tools | Crawler, CMS export, GSC, analytics | Inventory data + criteria + business context |
Inventory = what you have. Audit = what you should do about it. They’re sequential, not interchangeable. Build the list, then bring opinions.
How to Build a Content Inventory
You don’t need exotic tooling. A crawler, your analytics, Search Console, and a spreadsheet will carry 95% of this. Here’s the workflow we use.
1. Define scope
Decide what you’re inventorying: the whole site, just /blog/, product pages, a single content type, or a date range. Scope keeps the project finishable. A 50-page site is one afternoon; a 20,000-URL ecommerce catalog needs sampling and a database, not Sheets.
2. Export every URL in scope
Pull URLs from multiple sources and reconcile them — no single source is complete:
- Crawler (Screaming Frog, Sitebulb) for what’s linked and reachable.
- XML sitemap for what you think should be indexed.
- Google Search Console for what’s actually getting impressions.
- CMS export for everything, including unlinked drafts and old pages.
The gaps between these lists are themselves findings — sitemap URLs the crawler never reached are often orphaned content.
3. Capture the baseline fields
For each URL, record the descriptive data. At minimum:
| Field | Source | Why it matters |
|---|---|---|
| URL | Crawl / sitemap | The primary key |
| Page title | Crawl | Duplicate/missing titles surface fast |
| Meta description | Crawl | Spot blanks and stale copy |
| Content type / template | CMS / crawl | Group by post, LP, doc, product |
| Author / owner | CMS | Assigns accountability |
| Publish + last-updated date | CMS | Reveals stale, never-refreshed pages |
| Word count | Crawl | Flags thin content candidates |
| Indexation status | GSC | Find pages Google ignores |
4. Layer in performance data
The inventory becomes a decision tool the moment you bolt metrics onto each row:
- Traffic: sessions and users (GA4).
- Search: clicks, impressions, average position (Search Console); backlinks (Ahrefs/Semrush).
- Engagement: time on page, scroll depth, pages per session.
- Conversions: form fills, goal completions, revenue.
5. Tag a status per URL
Now you can make calls. Tag every row Keep / Update / Consolidate / Remove:
- Keep — performs well, factually current, leave it.
- Update — high potential, stale or under-optimized; queue a content refresh.
- Consolidate — multiple pages cannibalizing one query; merge and 301 to the strongest URL.
- Remove — thin, dead, or duplicate; 410 it or redirect to the closest relevant page, then clean up any 404s.
This is also where you catch a duplicate without user-selected canonical — pages Google clustered and chose a canonical for, against your intent.
Content Inventory Template
Copy this into Sheets or Airtable. One row per URL, one tab per content type if the site is large:
URL,Title,Meta Description,Type,Owner,Published,Last Updated,Word Count,Indexed,Clicks,Impressions,Avg Position,Sessions,Conversions,Backlinks,Status,Action,Notes
/blog/example/,Example Post,An example meta...,Blog,Maya,2022-03-04,2024-11-01,1240,Yes,310,8800,7.4,290,4,12,Update,Refresh stats + add internal links,Outdated 2022 data
/services/pseo/,PSEO Services,Programmatic SEO done right,Landing,Sam,2025-01-10,2026-05-02,890,Yes,540,15200,3.1,610,9,28,Keep,—,Top performer
The columns that earn their keep are Last Updated, Indexed, Avg Position, and Action — that quartet drives almost every refresh, prune, and consolidation decision. Everything else is supporting context.
Where the Inventory Plugs In
A live inventory isn’t a one-off artifact — it’s infrastructure. Use it to enforce a clean SEO site structure, map orphans back into your topic clusters, and route internal links toward your pillar pages. In the AI-Overviews era, where Google synthesizes answers from a handful of trusted pages, a bloated site of thin near-duplicates is a liability. Pruning off a solid inventory is one of the highest-leverage SEO moves left — and it’s the groundwork our programmatic SEO and growth program engagements start from.
Frequently Asked Questions
What is the difference between a content inventory and a content audit?
A content inventory is an objective, complete list of every URL on your site and its attributes — the raw dataset. A content audit applies judgment to that dataset, scoring pages on quality, performance, and relevance to decide what to keep, update, merge, or remove. You build the inventory first, then audit it.
What should a content inventory include?
At minimum: URL, page title, meta description, content type, owner, publish date, last-updated date, word count, and indexation status. Add performance data — clicks, impressions, average position, sessions, conversions, and backlinks — to turn the inventory from a descriptive list into a tool you can actually make pruning and refresh decisions from.
How often should I update my content inventory?
Treat it as a living document, not a one-time export. For high-velocity sites publishing weekly, refresh the inventory quarterly. For slower sites, twice a year is enough, with a full rebuild before any migration, redesign, or major pruning project. Stale inventories produce stale decisions.
What tools do I need to create a content inventory?
A crawler (Screaming Frog or Sitebulb), Google Search Console, an analytics platform like GA4, and a backlink tool (Ahrefs or Semrush). Reconcile their exports in Google Sheets or Airtable. For very large sites — tens of thousands of URLs — move to a database rather than wrestling a spreadsheet past its limits.
Can a content inventory help with AI Overviews?
Yes. AI Overviews pull from a small set of trusted, well-maintained pages, so bloated sites full of thin, duplicative content hurt you. An inventory exposes exactly which pages to consolidate, refresh, or remove, tightening your site into the kind of authoritative, deduplicated source that AI summaries are more likely to cite.