Content pruning – less content, more visibility. How can you organise your website content effectively?

Table of Contents

chevron

Content pruning is the process of organising existing content on a website by updating, merging, redirecting or removing it in order to improve the site’s visibility and quality. The aim is not to reduce the size of the website at any cost, but to retain content that is up-to-date, valuable and meets the specific needs of users.

A large blog does not always mean better results on Google. Older articles may become out of date, several subpages may be competing for the same keyword, and some content may not generate traffic, attract backlinks or support the company’s current offering. As a result, rather than strengthening the website, hundreds of URLs make it harder for search engines to assess its key topics and undermine the visibility of valuable pages.

This approach is particularly useful for website owners, content managers, SEO specialists and those responsible for website development who wish to improve the quality of their content, reduce keyword cannibalisation and better direct traffic to pages that support business objectives. In this article, we explain what content pruning is, how to tell when it’s worth carrying it out, what data to analyse, how to decide whether to update, merge, redirect or remove content, and what mistakes to avoid in order to organise your website without losing valuable traffic.

What is content pruning?

Content pruning is the process of analysing, classifying and organising content published on a website.

Its aim is not to automatically remove old articles. Well-executed pruning involves making one of several decisions for each URL:

  • to leave the content unchanged,

  • updating it or rewriting it from scratch,

  • linking it to another article,

  • to remove it from the website.

In practice, pruning can be compared to caring for a tree. We don’t cut down the whole plant. We remove branches that are dead, weaken the rest of the plant or are growing in the wrong direction.

That is why content pruning should not end with a presentation or a spreadsheet listing the issues. It is only complete once specific decisions have been made and implemented for the URLs under review.

Nor is this a one-off task. A blog evolves, matures and continues to accumulate content, which may vary in quality without a clear content strategy. Depending on the scale of the blog, it is worth repeating this process, for example every 3, 6 or 12 months.

Why can a large number of articles be a problem?

Each article should have a specific purpose (or several). For example, they might:

  • answer users’ questions,

  • to support a particular product or service category,

  • develop a relevant thematic silo that helps build expertise in a particular niche,

  • build the author’s or brand’s credibility,

  • generate traffic that leads to conversions.

If it is not possible to clearly define the purpose of a piece of content, it is likely to become a candidate for pruning over time.

Large-scale blogs usually suffer from a number of recurring problems.

Keyword cannibalisation

Several articles answer almost the same question. Instead of a single comprehensive page, several mediocre subpages are created, between which Google has to choose. As a result, these subpages appear intermittently in the top 10 or do not appear on the first page of search results at all.

For example, the website may feature separate articles such as:

  • „Which collagen should I choose?”,

  • „Which collagen is best?”,

  • „Good collagen – a ranking and advice”,

  • „What should you look out for when buying collagen?”.

These are not always four separate topics. Sometimes it is a single topic split into several sections, e.g. „How to choose collagen? Part 1” and „How to choose collagen? Part 2”. Such content should be on a single subpage and form a single, complete article.

The ageing of information

Not all content becomes outdated at the same rate. A mathematical equation remains ‘evergreen’ (content that is always relevant) practically forever. An article on tax law, technology, products or supplementation may require frequent updates. This is due to the emergence of new legislation or research that changes the approach to particular issues.

Old content that continues to generate traffic but contains out-of-date product information or recommendations is particularly damaging to the brand.

The blurring of the website’s specialisation

The content on the website should support its area of activity. If a supplement manufacturer publishes random recipes or travel advice, it becomes more difficult for both users and Google’s algorithm to understand in which field the brand wishes to be an expert or authority.

This does not mean that every article has to directly sell a product. It should, however, fit in with the website’s thematic context and reinforce its expertise.

Difficulties with internal linking

The more random content there is, the harder it is to establish a clear hierarchy and, consequently, logical, semantic internal linking between articles.

How should you create content to ensure effective internal linking?

  • Content should be created in a silo-based approach, starting with a single overarching piece of content, which is then expanded upon

  • plan linking within silos to reinforce the content within them,

  • avoid creating orphan pages (unlinked „orphan” subpages).

The cost of maintaining low-quality content

Search engine bots must crawl and index every published URL. If the content is not of high quality, not only will it never make it into the top 10 search results, but it will also use up our crawl budget.

A single poor-quality article isn’t a major problem. However, hundreds or thousands of such subpages start to pose a challenge for search engine crawlers, as it takes them a considerable amount of time to crawl and index them.

When is it a good idea to carry out content pruning?

It is worth considering content pruning when at least one or more of the following indicators are present:

  • Content pruning is particularly effective for websites containing several hundred or several thousand articles. However, it can also be carried out to good effect on smaller websites,

  • the texts on the website have not been updated for at least a year,

  • subsequent articles are not delivering the expected increase in visibility,

  • the visibility of the entire blog or a particular section is gradually declining,

  • Articles on similar topics appear on the website, which may indicate content cannibalisation,

  • Most of the traffic is generated by a small percentage of articles,

  • the blog covers topics unrelated to the website’s current content or subject matter,

  • the website contains old news items, announcements and short pieces with no substantive value,

  • The structure of the website, as well as the types of categories, products and services, have changed.

A good starting point is to compare the age of the content with the traffic it generates. It may turn out, for example, that articles published before a certain date account for 50% of the entire blog, but are responsible for only 5% of the clicks. This result does not necessarily mean that they should all be deleted. However, it does show where to begin a more detailed analysis.

What data should you collect before starting a content audit?

Decisions on content removal should not be based solely on intuition. A combination of quantitative and qualitative data is required, which we meticulously compile before every content pruning exercise for our clients.

Data from Google Search Console

For each address, it is worth collecting data covering at least the last 12 months:

  • clicks,

  • views,

  • the middle position,

  • the number of enquiries,

  • search queries that generate the highest visibility.

A lack of clicks does not always mean a lack of potential. An article may have a high number of views further down the rankings and may, above all, need to be improved in terms of quality.

Data from the crawl

A crawl of the website should provide, amongst other things:

  • response code,

  • canonical tag,

  • title and description,

  • headings,

  • the number of words,

  • click depth,

  • the number of internal links,

  • date of publication,

  • date of update,

  • information about the author,

  • belonging to a specific category.

This data enables us to compile a comprehensive report containing information that supports the analysis and helps us draw conclusions about the content on the website.

Value of the content

Please check whether the subpages under review have:

  • external links,

  • visibility history,

  • mentions or traffic from other channels.

Deleting a subpage that has good external links pointing to it may result in a loss of accumulated value. Where a subpage has links pointing to it, we usually leave it at its old URL.

Business data

Not every high-quality article necessarily generates a lot of organic traffic.

The content may:

  • lead to conversions,

  • support the sales department,

  • reduce the number of enquiries to customer services,

  • to help compare products,

  • build trust in an important area.

Therefore, SEO data should be cross-referenced with conversions, clicks on products, sign-ups, downloads, an assessment of the business value of the content being analysed, or other events relevant to the business.

Semantic similarity analysis

When dealing with a large number of texts, it is worth using embeddings or other similarity analysis methods.

A semantic map can help you find:

  • duplicate topics,

  • groups of articles addressing the same intent,

  • content located outside the main clusters,

  • silos that are either too small or too scattered.

Data from the GSC, crawling, content extraction and embeddings can be combined into a single analytical process.

Don’t delete the article just because it’s old

One of the most common mistakes is to lay down an arbitrary rule, which might go something like this:

We either delete or update all articles that are more than three years old.

Such a limit may be too lenient for a fast-moving industry and too restrictive for evergreen content. It all depends on the industry you’re in.

A better approach is to compare the age of your own articles with that of the pages currently ranked in the TOP 10.

Content Age Index

The age metric can be defined as the age of the content that Google takes into account when determining rankings. The timeliness of content is particularly important in sensitive YMYL (Your Money Your Life) sectors.

An example of the calculation process:

  1. Select 30–50 representative phrases from different sections of the website.

  2. Download the TOP 10 results for each search term.

  3. Check the date of the last update for each subpage.

  4. If there is no update date, use the publication date.

  5. Calculate the median age and the 75th percentile.

Why the 75th percentile? It allows us to establish the threshold above which the oldest quarter of the websites in the TOP10 are found.

An article that is older than the indicator does not necessarily have to be deleted. However, it becomes a candidate for further analysis.

Age of the article Interpretation
Younger than the indicator suggests The date is probably not the main issue
Close to the index It’s worth checking the quality, relevance and up-to-date nature of the content
Older than the indicator Candidate for updating, rewriting, merging or deletion
Many times older than the TOP10 results High priority for analysis and removal/updating, particularly in a fast-moving industry

Four key decisions in content pruning

Once the data has been collected, each address should be assigned to one of the four groups.

1. Keep

The content remains in its current form when:

  • generates high-quality traffic,

  • ranks highly,

  • carries out the user’s intention,

  • has gained external links,

  • supports sales or other business objectives,

  • fits in with the main thematic area.

„Retain” does not mean „forget”. It is worth setting a date for a follow-up review to check that the information is still up to date.

2. Update or rewrite

An update is a good choice when an article:

  • it has views, but ranks lower,

  • responds to meaningful enquiries,

  • contains out-of-date information,

  • it is too shallow,

  • does not directly address the intention,

  • does not contain any sources, data or expert input,

  • has a history that is too valuable to lose.

An update shouldn’t just involve changing a few sentences and the date.

Depending on the problem, you will need to do the following:

  • further research,

  • change/expansion of the structure,

  • verification by an expert,

  • removal of outdated sections,

  • improving internal linking,

  • adjusting the system to better match the user’s intentions.

3. Combine

„We use ”article merging’ when several articles cover the same or very similar topics and often target the same search terms. This allows us to eliminate keyword cannibalisation and establish a single subpage that ranks for a given keyword.

The process should include:

  1. selecting the best destination address,

  2. collecting valuable excerpts from all the pages,

  3. the creation of a single, comprehensive piece of content,

  4. redirecting the remaining URLs using a 301 redirect,

  5. updating internal links,

  6. removing old URLs from the sitemap.

We usually select the target URL based on visibility history, link profile, URL match and the quality of the current content, as this consolidation also supports the optimisation of content around a single, main URL.

4. Delete

Content removal is generally justified when a page:

  • does not generate traffic or page views,

  • there are no external links,

  • does not respond to substantive enquiries,

  • does not fit in with our current activities,

  • contains very poor or out-of-date content,

  • does not contain any content worth using in another article,

  • does not fulfil a key business function.

Depending on the situation, the following may be used:

  • 404 or 410, when a page has no counterpart and does not pass on any values,

  • 301, when there is a similar landing page that can generate traffic and improve visibility in search results

  • noindex, when content is needed by users but should not appear in search results; this also applies to situations where certain content, such as a landing page offering a discount, is necessary for business purposes but does not need to compete on Google.

You should not redirect all deleted articles to the main page or a random category. The redirect should make sense to the user.

The order in which things are done matters

The safest order of implementation is:

1. Merge

First, we merge duplicates and related topics. This ensures we don’t remove content that could strengthen a more important article.

2. Removal

Once you have identified the valuable sections and URLs, you can remove content that has no traffic, no links and no business justification.

3. Rewriting/expanding

Finally, we rewrite or expand the content with the greatest potential – that is, content that appears in the results for many search queries but does not rank among the top results in search engines.

This sequence can be summarised as follows:

„Merging preserves value, pruning tidies up the site’s structure, and rewriting/expansion drives growth.”

An example decision matrix

Our proprietary table can help standardise and support the classification of content for content pruning.

Traffic and visibility Quality of content Business links Recommendation
High High High Keep
High Low High Update
Low High High Improve optimisation and linking
Low Low High Rewrite or merge
Low Low Low Delete
Averages across several similar URLs Miscellaneous High Merge
No movement, but strong links Low Medium or high Merge or a 301 redirect
Exercise, but a topic outside my area of specialism Average Low Individual business assessment

The matrix should not operate automatically. However, it helps to build an initial set of recommendations and reduce the number of addresses requiring manual analysis.

What should you do once you’ve finished pruning?

Pruning should not leave any gaps. Its aim is to create a cleaner structure for further development and to establish a regular content pruning process that will prevent the accumulation of outdated and low-quality content.

Build thematic silos

Each article should have a specific place in the hierarchy:

  • main topic page,

  • cluster articles,

  • more specific questions,

  • relevant categories, products or services.

If the content cannot be assigned to any relevant area, it is worth asking again whether it is necessary.

Organise your links

Once the articles have been deleted and merged, you should:

  • correct links pointing to old URLs,

  • add internal links to subpages that are not currently linked internally,

  • remove links to 404 pages,

  • Check the sitemaps and every page in the link structure to ensure that no old URLs remain and that new ones have been added.

Implement a content update policy

A different review frequency can be set for each content group, and regular reviews usually improve content quality and search engine visibility.

Content type Example inspection frequency
Legal and tax news Every 3–6 months
Product and tool rankings Every 3–6 months
Health and supplementation Every 6–12 months
Technical guides Every 6–12 months
Evergreen content Every 12–24 months
Company information After any change to the offer or organisation

For large websites, it is worth carrying out pruning every 1–3 months. Websites with up to 1,000 subpages should be reviewed every 6 months. Smaller websites can be reviewed once or twice a year.

This is merely a starting point. The appropriate frequency should be determined by the pace of change in a particular sector. In practice, it is worth scheduling regular reviews as part of monthly maintenance tasks and basing them on a fixed timetable, whilst monitoring the results using analytical tools.

Create content featuring experts

It is worth combining content pruning with a change to the standards for new publications.

Expertise should not be limited to the author’s biography. An expert should provide specific knowledge, verify the answer and take responsibility for its content. It is best to use the services of experts who are widely recognised within the industry and who have published articles on industry portals. The greater their recognition, the better for us.

One possible format is the „Expert Opinion”. This involves one question from a user and one specific answer signed by an expert. This structure makes it easier to scale content without requiring the expert to write long articles themselves.

How can you measure the impact of content pruning?

The results of pruning should not be assessed solely on the basis of the overall increase in traffic during the first few weeks following the removal of content. Traffic may temporarily drop, but ultimately, following the next algorithm update, the page should be re-ranked, thereby restoring the upward trend.

It is worth keeping an eye on:

  • the number of URLs in the index,

  • the number of broken internal links,

  • the proportion of important sections crawled,

  • visibility of the entire blog,

  • the number of phrases in the TOP 3 and TOP 10,

  • visibility of individual silos,

  • cannibalisation,

  • activity on preserved and updated articles,

  • transitions from the blog to the product range,

  • assisted conversions,

  • visibility in Direct Answers and AI-generated answers.

The most common mistakes made during content pruning

Why aren’t subpages deleted solely on the basis of traffic?

A lack of clicks does not automatically mean a lack of potential. An article may still generate page views, contain good links or have sales value.

Why aren’t subpages deleted solely on the basis of their publication date?

An old article may still be the best answer. A new piece of text may not meet the user’s needs from the outset.

Why don’t we redirect all subpages to the home page?

A 301 redirect should lead to a page with a relevant topic. Otherwise, it does not help the user and may be treated as an inappropriate redirect.

Why do we merge content first and only then delete it?

The deleted content may contain sections, data or links that would be worth transferring to a more comprehensive article.

Why is it necessary to update internal links?

Leaving hundreds of internal links pointing to deleted pages detracts from the user experience and hinders crawling.

Why shouldn’t you publish random texts?

If the new plan isn’t based on silos, intentions and business objectives, the blog will quickly revert to its previous state.

Why do we have to wait for the results of content pruning?

Implementing changes does not necessarily lead to an immediate increase in traffic. The search engine needs time to revisit the URLs, process the redirects and assess the revised structure.
Author
Kamil Brzozowski

Kamil Brzozowski

Co-founder & SEO Expert

SEO Expert & Co-founder Odyseo. SEO specialist with the industry since 2014. He has gained experience by leading his own projects and also working as SEO Team Leader at agencies Semcore and SAMOSEO.

He has had and continues to have the pleasure of conducting SEO projects for many well-known companies such as AuraHerbals.pl, AptekaGemini.pl, Apteka NowaFarmacja.pl, Mydlarnia 4szpaki.pl, Bank Zachodni WBK, Bakalland, MTU24.pl, Wójcik Fashion, KUPLIO, ATAS, CStore.pl, Raszczyk.com.pl.

It focuses on long-term SEO in line with Google guidelines, combining technical optimisation with UX and content development. His goal is to increase conversions and key site events. I develop strategies, perform audits and develop content based on effective content marketing.

More about the author

Author

Enter a search term

Explore our case studies

Find out how we increased traffic by 300% for our client.

Our offer

We design our comprehensive SEO campaigns to increase the visibility and caloric organic traffic to your website, which directly translates into increased profits.

We know SEO inside out and have created campaigns for some of the biggest companies in Poland. Our SEO campaigns are transparent, effective and created in a partnership atmosphere.

We acquire powerful links which, in combination with an SEO-optimised website, will yield measurable results.

Content marketing is our hobby. We have created blogs for some of the biggest businesses in Poland. In addition, Michal has over 130 blogs of his own, which he develops regularly.

Using the synergy of SEO and PPC, we create complete campaigns that will attract new customers with their gravitas.

Blog

Related articles

Discover our latest blog posts on SEO and marketing.

Precise growth trajectory - time for your move!

Tired of ineffective SEO efforts? It's time for a strategy that delivers real results. Contact us and we'll analyse your situation and offer tailored solutions. With us, you'll gain full transparency, effective action and real growth.

Write to us