modifiedTime && Skip to content
All posts

An AI Crawler Took Down a WooCommerce Store via Its Cart

A crawler followed about 51,000 add-to-cart and remove links in one day. Every visit skipped the cache. Here is why and how we blocked it.

Radoslav Stefanov

In August one of the WooCommerce stores we host kept going down during a normal day. The site kept hitting its memory limit and the server killed its PHP workers each time, while traffic from real shoppers looked completely ordinary. When we went through the access logs, almost all of the load came from one crawler, Meta’s AI crawler, which identifies itself as meta-externalagent. In a single day it had requested about 51,000 cart URLs.

WooCommerce puts actions like “add to cart” and “remove item” into ordinary links:

/cart/?remove_item=abc123&_wpnonce=...
/?add-to-cart=452

A crawler that follows every link it finds keeps discovering new combinations of these, so it never runs out of URLs to request. Cart and checkout pages are different for every visitor, so they are never served from the page cache, and every one of those requests made WordPress build a full page. On this store a cart page used around 275 MB of memory to build, and a handful of them arriving at the same time was enough to fill the site’s 1 GB limit.

What we changed

We now turn these requests away at the proxy in front of every site, before they reach WordPress. A request is refused only when it comes from a crawler on an explicit list, such as meta-externalagent, GPTBot and ClaudeBot, and it is also a cart action, meaning a cart or checkout page or a URL with add-to-cart, remove_item, undo_item or a coupon parameter. We don’t match the word “bot” in general, and we left out the crawler Facebook uses for link previews, so shared product links still show a preview. The path rules are anchored to the start of the URL, because a product like “cartoon mug” at /cartoons/ must never be blocked.

In nginx it looks roughly like this:

map $http_user_agent $is_crawler {
    default 0;
    ~*(meta-externalagent|gptbot|claudebot) 1;
}
map $request_uri $is_cart_action {
    default 0;
    ~^/(cart|checkout)(-[0-9]+)?(/|\?|$) 1;
    ~[?&](add-to-cart|remove_item|undo_item)= 1;
}
map "$is_crawler$is_cart_action" $block_bot_cart {
    default 0;
    11 1;
}
# inside the server or location block:
if ($block_bot_cart) { return 403; }

Since then the crawler gets a quick 403 from the proxy and the store has stopped running out of memory. Search engines still crawl product and category pages as before, and only the cart actions are refused, which shouldn’t be indexed anyway.

If you run a WooCommerce store, search your access logs for add-to-cart and remove_item requests from crawlers.

Related update: Protection From Bots Filling Your Cart