Short answer: You cannot fully stop scrapers from copying a public feed, but you can make copying pointless and easy to trace. Publish summaries instead of full text if theft is a real problem, add a source link to every feed item, keep canonical tags and internal links in your content, and consider a short feed delay so your pages are found first. For copies that still hurt you, use takedown notices to the host and removal requests to search engines. Avoid blocking measures that also lock out real readers.
What feed scraping looks like
An RSS feed is designed to be read by software. That is its strength: readers, podcast apps, newsletters and automation tools all rely on it. The same openness means that anyone can point a script at your feed and republish every new item on another site, usually automatically and without permission. These sites are often called autoblogs or scraper sites.
Typical signs that your feed is being scraped:
- Exact copies of your articles appear on unknown domains within minutes or hours of publication.
- Your analytics or server logs show hotlinked images loading on other domains.
- You receive pingbacks or backlinks from low-quality sites that contain your full text.
- Searching for a distinctive sentence from a new article returns other sites alongside yours.
It helps to keep perspective. Most scraper sites get very little traffic, and search engines are generally good at identifying the original. The practical harm is usually small, but it is worth a few simple defences, especially for publishers whose content is their product. The risks for the sites that do this are covered in autoblogging from RSS: SEO, legal and quality risks.
Legitimate reuse versus scraping
Before you change anything, separate the uses you want from the ones you do not. Much of what reads your feed is welcome:
- RSS readers and apps used by your subscribers, which show your content to the people who chose to follow you.
- Aggregators and news apps that show headlines and summaries and send readers to your site.
- Automation tools that post your new articles to your own social pages or email list.
- Partners with whom you have a syndication agreement, which the guide to content syndication with RSS covers in detail.
Scraping is different: full copies published elsewhere without permission, usually without a clear link back, often surrounded by ads. Every defence below is chosen to hurt that use while leaving legitimate readers untouched. The legal side of reusing other sites’ content is discussed in whether it is legal to create feeds from other websites.
Defence 1: decide between full text and summaries
The single most effective measure is to publish a summary instead of the full article. A scraper that republishes a two-sentence excerpt gains almost nothing, and the copy points readers to your site.
The trade-off is real, though. Many subscribers prefer full-text feeds because they can read everything in their reader, and some will unsubscribe if you switch. Consider:
- Keep full text if your feed audience is loyal and scraping has not caused measurable harm.
- Switch to summaries if scraped copies outrank you, if your content is expensive to produce, or if you depend on page views.
- Write proper excerpts rather than truncating the first paragraph, so summary-only readers still get a useful preview.
The full comparison is in full text or summary: choosing what your feed shows.
Defence 2: make every copy point back to you
Scrapers rarely edit what they copy. Use that to your advantage by putting attribution inside the feed content itself:
- A source line in every item. Add a short footer such as “This article first appeared on [site name]” with a link to the original URL. In WordPress this can be done with a small snippet or plugin, as shown in adding a footer or links to every feed item.
- Internal links in your articles. Links to your own related articles, written as absolute URLs in the feed, travel with the copied text and send any readers back to your site.
- Canonical tags on your pages. Your own pages should declare themselves as canonical. Scrapers do not always strip or change canonical references in copied content, and even when they do, a clear canonical on the original helps search engines decide.
- Your name in images. A small, unobtrusive credit on original graphics identifies them wherever they end up.
These measures do not stop copying, but they turn many copies into weak backlinks and make the original obvious to readers and search engines alike.
Defence 3: publish first, syndicate second
When a scraper copies your article within seconds, both versions exist almost simultaneously, and in rare cases the copy is discovered first. A short feed delay gives your own page a head start:
- Delay the feed by 15 to 60 minutes after publication, so search engines can find the original through your sitemap and internal links first. WordPress users can follow how to delay your WordPress RSS feed.
- Keep your XML sitemap up to date so new URLs are discoverable immediately.
- Link to new articles from your home page or blog index as soon as they go live.
Keep the delay short. Subscribers and automation tools expect timely updates, and a delay of several hours makes your feed feel stale.
Defence 4: technical controls, used carefully
Blocking is tempting but risky, because scrapers and legitimate readers use similar technology. Some measures are reasonable:
| Measure | Helps against | Risk to real readers |
|---|---|---|
| Rate limiting the feed URL | Aggressive polling bots | Low, if limits are generous |
| Blocking specific abusive IP addresses | A known scraper server | Low, but scrapers change addresses |
| Blocking by user agent | Lazy scripts | Medium: easy to fake, easy to block the wrong tool |
| Hotlink protection for images | Copies loading your images | Medium: can break images in some readers |
| Password or token-protected feeds | Everything unauthorised | High for public feeds; fine for member content |
If you add protection at the server or CDN level, make sure the feed still returns the right headers and caches sensibly; serving an RSS feed correctly explains what readers expect. After any change, test the feed in two or three popular readers.
Defence 5: find copies and request removal
For copies that matter, for example ones that rank for your key topics, use formal channels:
- Find the copies. Search for exact sentences from your articles in quotation marks, or use a plagiarism checking service for important pieces.
- Document them. Save the URL, a screenshot and the date you found it, together with proof of your original publication date.
- Contact the host. Look up the hosting provider of the copying site and send a copyright complaint through its abuse process. In the United States this is typically a DMCA notice; other countries have comparable procedures.
- Ask search engines to remove the copy. Search engines offer legal removal request forms for copyright infringement, which can remove infringing pages from results.
- Prioritise. Chasing every tiny scraper is not worth the time. Focus on copies that compete with you for traffic.
This is general information, not legal advice. For serious or repeated infringement, speak to a lawyer familiar with copyright in your jurisdiction.
What not to do
- Do not remove your feed. The subscribers, apps and tools that rely on it are worth far more than the damage scrapers do.
- Do not break the feed for everyone. Heavy blocking, captchas or JavaScript challenges on the feed URL lock out RSS readers.
- Do not hide content in tricks. Invisible text or deliberately broken markup can confuse readers and search engines more than scrapers.
- Do not panic about rankings. Search engines usually rank the original. Strengthen your own pages rather than trying to erase every copy.
Feeds for your readers, on your terms
Feeds works from the reader’s side: it helps people follow websites through clean feeds, including sites that publish no RSS, and it reads public pages like a visitor, identifying itself as FeedsBot. Its feeds carry titles, images, summaries and dates with a link to the original item, so readers are sent to the source rather than given a copy to republish. Publishers who want to understand how such tools see their pages can simply create a feed from their own blog and look at the preview.
Related reading
- RSS feed best practices for publishers: a 24-point checklist
- Why every company blog should publish a clean RSS feed
- How to measure RSS feed subscribers and readership
The bottom line
Scrapers are an irritation more than a threat for most publishers, and the answer is not to hide your feed. Decide deliberately between full text and summaries, put a source link in every item, keep canonical tags and internal links, add a short delay if copies appear before your pages are found, use gentle technical limits, and reserve takedown requests for copies that genuinely compete with you. Your subscribers keep a good feed, and scrapers get very little from it.
GYIK
Can I stop people from copying my RSS feed?
Not completely, because a public feed can be read by anyone. You can make copying far less useful by publishing summaries, adding source links to every item and keeping canonical tags, and you can request removal of copies that harm you.
Should I switch my feed to summaries because of scrapers?
Only if scraping causes real harm, such as copies outranking you. Full-text feeds are popular with subscribers, so weigh their preferences against the damage before switching.
Do scraper sites hurt my SEO?
Usually not much, because search engines are generally good at identifying the original. Clear canonical tags, internal links and quick indexing of your own pages make that even more likely.
Is it safe to block bots from my RSS feed?
Be careful. RSS readers, podcast apps and automation tools are bots too. Prefer rate limits and blocking specific abusive addresses over broad user agent or challenge-based blocking, and test the feed in real readers afterwards.
How do I get a scraped copy of my article removed?
Document the copy and your original publication, send a copyright complaint to the copying site’s hosting provider, and use search engines’ legal removal request forms. For repeated or serious cases, get legal advice.


