Another massive distributed HTTP flood is currently hitting git.friendi.ca and causing slow responses. Operationally, this has to be treated as an application-layer DDoS attack, although we cannot determine from the logs whether disruption or aggressive scraping is the original intent.

In a fresh five-minute sample, Anubis saw 56,457 requests carrying client IPs from 50,470 different addresses:


  • 48,519 requests via IPv4
  • 7,938 requests via IPv6
  • 43,621 unique IPv4 addresses
  • 7,407 unique IPv6 addresses


We compared all addresses against the Tor Project’s official, target-specific exit list for our server on port 443. There were zero Tor exit nodes in the sample.

The traffic is spread across many networks and regions and appears more consistent with a large residential-proxy or compromised-device network than with Tor. We are investigating additional capacity protection that does not lock out legitimate users.

!Friendica Support

This entry was edited (Wednesday, August 12, 2026, 7:15 AM)
in reply to utzer

Update: The flood is still ongoing and intensified again today.

Between approximately 17:50 and 22:03 CEST, Anubis recorded at least 141,090 challenges from 136,393 different IP addresses, peaking at roughly 35,800 challenges per hour.

Most of the current traffic is still targeting the friendica-addons pull-request list. Our temporary rate limit is successfully keeping this traffic away from the backend: the rest of git.friendi.ca remains fast and accessible, normal Git access works, and server load is currently back to normal.

The affected pull-request page may continue to return “Too Many Requests” while the flood persists.

in reply to utzer

@utzer what is the user agents, looking at your access server logs, you should be able to see if even you had 1million different ip hitting at the same time, if all the user agents are the same, then you have trace back all million ip between whois and dig, to figure out if they are tied together, if all same user agent and time each ip hit and pulls same amount of data, that leans more to coordinated DDoS, if differing user agents, differing hit times and data captures, then massive scraping
did I make any sense
in reply to pasjrwoctx👽

@pasjrwoctx👽 Thanks, I checked the user agents. The pattern looks highly coordinated.

Out of 164,435 Anubis challenge requests, there were only 160 distinct user-agent strings. The top ten account for 96.79% of all requests and were used almost perfectly evenly:

  • Firefox 120 / Windows — 16,066
  • Chrome 118 / Windows — 16,052
  • Firefox 121 / macOS — 15,949
  • Edge 119 / Windows — 15,917
  • Chrome 119 / Windows — 15,874
  • Chrome 119 / macOS — 15,872
  • Edge 120 / Windows — 15,863
  • Chrome 120 / macOS — 15,853
  • Firefox 121 / Windows — 15,812
  • Chrome 120 / Windows — 15,761

These are old browser versions from around 2023. Combined with almost one IP per request, this looks like a single coordinated tool rotating through a fixed list of spoofed browser identities and a very large residential/proxy network. It does not prove whether the objective is scraping or disruption, but it is clearly not ordinary independent crawling.

in reply to utzer

@pasjrwoctx👽

More details regarding user-agent distribution

Retained log window: 2026-08-14 17:50–22:42 CEST
Challenge requests: 164,435
Distinct user-agent strings: 160
Share represented by the top ten: 96.79%

The almost perfectly even distribution among the first ten identities is particularly striking.

 COUNT  USER AGENT
------  ------------------------------------------------------------

 16106  Mozilla/5.0 (Windows NT 10.0; Win64; x64)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/118.0.0.0 Safari/537.36

 16097  Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:120.0)
        Gecko/20100101 Firefox/120.0

 15993  Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:121.0)
        Gecko/20100101 Firefox/121.0

 15959  Mozilla/5.0 (Windows NT 10.0; Win64; x64)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/119.0.0.0 Safari/537.36 Edg/119.0.0.0

 15924  Mozilla/5.0 (Windows NT 10.0; Win64; x64)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/119.0.0.0 Safari/537.36

 15914  Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/119.0.0.0 Safari/537.36

 15899  Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/120.0.0.0 Safari/537.36

 15887  Mozilla/5.0 (Windows NT 10.0; Win64; x64)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/120.0.0.0 Safari/537.36 Edg/120.0.0.0

 15855  Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:121.0)
        Gecko/20100101 Firefox/121.0

 15800  Mozilla/5.0 (Windows NT 10.0; Win64; x64)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/120.0.0.0 Safari/537.36

  2871  Mozilla/5.0 (compatible; crawler)

   709  Mozilla/5.0 (X11; Linux x86_64)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/148.0.7778.0 Safari/537.36

   413  Mozilla/5.0 (Windows NT 10.0; Win64; x64)
        AppleWebKit/537.36 (KHTML, like Gecko)
        Chrome/144.0.0.0 Safari/537.36

    86  Chrome/135.0.0.0 — Windows
    40  Chrome/116.0.0.0 — Windows
    40  Chrome/107.0.0.0 — Windows
    39  Chrome/148.0.0.0 — macOS
    34  Chrome/150.0.0.0 — Windows
    33  Chrome/131.0.0.0 — Windows
    32  Chrome/133.0.0.0 — Windows
    25  Edge/135.0.0.0 — Windows
    24  Chrome/146.0.0.0 — macOS
    23  Chrome/145.0.0.0 — Windows
    23  Chrome/110.0.0.0 — Windows
    22  Chrome/126.0.0.0 — Linux
    22  Chrome/123.0.0.0 — macOS
    21  Chrome/146.0.0.0 — Windows
    21  Chrome/109.0.0.0 — Windows
    21  Chrome/136.0.0.0 — macOS
    20  Chrome/99.0.4844.51 — Windows
    20  Chrome/149.0.0.0 — Windows
    20  Chrome/145.0.0.0 — macOS
    19  Chrome/133.0.0.0 — macOS
    18  Chrome/105.0.0.0 — Windows
    17  Chrome/134.0.0.0 — Windows
    17  Chrome/111.0.0.0 — Windows
    17  Safari/15.3 — macOS
    17  Chrome/131.0.0.0 — macOS
    16  Edge/99.0.1150.30 — Windows
    16  Chrome/104.0.5112.81 — Windows
    16  Firefox/135.0 — macOS
    16  Safari/18.0 — macOS
    16  Safari/15.5 — macOS
    15  Chrome/100.0.4896.75 — Windows
    15  Chrome/149.0.0.0 — macOS
    14  Firefox/137.0 — Windows
    14  Chrome/147.0.0.0 — Windows
    14  Chrome/124.0.0.0 — Windows
    14  Chrome/117.0.0.0 — Windows
    14  Chrome/112.0.0.0 — Windows
    14  Chrome/108.0.0.0 — Windows
    14  Chrome/101.0.4951.67 — Windows
    14  Safari/26.0 — macOS
    14  Safari/17.0 — macOS
    14  Chrome/124.0.0.0 — macOS
    13  Chrome/104.0.0.0 — Windows
    13  Safari/18.4 — macOS
    13  Chrome/150.0.0.0 — macOS
    12  Firefox/133.0 — macOS
    12  Chrome/135.0.0.0 — macOS
    11  Chrome/148.0.0.0 — Windows
    11  Chrome/103.0.0.0 — Windows
    10  Edge/101.0.1210.47 — Windows
    10  Safari/18.3.1 — macOS
    10  Chrome/134.0.0.0 — macOS
     9  Chrome/106.0.5249.119 — Windows
     9  Chrome/147.0.0.0 — macOS
     8  Chrome/142.0.0.0 — Windows
     7  Firefox/137.0 — Ubuntu Linux
     7  Firefox/125.0 — Linux
     7  Chrome/132.0.0.0 — Linux
     7  Chrome/151.0.0.0 — Windows
     7  Chrome/130.0.0.0 with CCleaner — Windows
     7  Chrome/139.0.0.0 with Safari WebKit string — macOS
     6  Chrome/137.0.0.0 — Linux
     6  Chrome/114.0.0.0 — Linux
     6  Firefox/153.0 — Windows
     6  Edge/151.0.0.0 — Windows
     5  Chrome/131.0.0.0 — Linux
     5  Chrome/134.0.0.0 — ChromeOS
     5  Chrome/106.0.0.0 — Windows
     5  Chrome/128.0.0.0 — macOS
     5  Chrome/151.0.0.0 — Android
     4  Firefox/136.0 — Ubuntu Linux
     4  Chrome/129.0.0.0 — Linux
     4  Firefox/140.0 — Windows
     4  Firefox/135.0 — Windows
     4  Chrome/138.0.0.0 — Windows
     4  Edge/136.0.0.0 — Windows
     4  Chrome/129.0.0.0 — Windows
     4  Safari/17.6 — macOS
     4  Safari/17.5 — macOS
     4  Chrome/147.0.0.0 — Android
     3  Chrome/136.0.0.0 — Linux
     3  Chrome/130.0.0.0 — Linux
     3  Chrome/124.0.0.0 — Linux
     3  Chrome/133.0.0.0 — ChromeOS
     3  Obsidian/1.8.10, Electron/34.2.0 — Windows
     3  Chrome/58.0.3029.110 — Windows
     3  Edge/138.0.0.0 — Windows
in reply to utzer

@utzer ok so you could block those user agents I have something like # Fake browser detection
RewriteCond %{HTTP_USER_AGENT} (Chrome/[0-9]{3}|Chrome/1[3-9][0-9]|Chrome/150|Firefox/1[3-9][0-9]|Safari/60[0-9]|Version/17)
[NC]RewriteCond %{HTTP_ACCEPT_LANGUAGE} ^$
RewriteRule ^ - [G,L]

RewriteCond %{THE_REQUEST} "GET\shttp"
[NC]RewriteRule ^ - [G,L] and then # 1. BLOCK BAD BROWSER NAMES / BOT FRAMEWORKS

RewriteCond %{HTTP_USER_AGENT} (CCBot|SearchEngineBot|Pandalytics|UCBrowser|ZoneProjectBot|Embarcadero\sURI\sClient|Xenu\sLink\sSleuth|siteradar|SignalsBot|fun-cert-watch|SERankingBacklinksBot|Pinterestbot|CMS-Checker|HeadlessChrome|Puppeteer|SeznamBot|Sogou|8LEGS|HTTrack|cherrypicker|AhrefsBot|BLEXBot|DotBot|MJ12bot|PetalBot|SemrushBot|BuiltWith|Viewer/99|Python|aiohttp|curl|Wget|libwww|Go-http-client|GeedoShopProductFinder|DuckDuckBot|node|IMJ-CompanyPage-Scraper|baidu|RootEvidence|NetAPI\sv1|Scrapy|Bingbot|SummalyBot|got|HUNT-Bot|CibraxScanner|SalesOS-CompanyVerifier|RecordedFuture|SurdotlyBot|panscient\.com|Xiaomi|Android.*Firefox|GPTBot|ClaudeBot|BardBot|LLMScraper|Firecrawl|Crawl4AI|ia_archiver|archive\.org_bot|Google|wp2shell|okhttp|Cortex-Xpanse|Mozilla\.5\.0\.compatible;\.MSIE\.10\.0;\.Windows\.NT\.6\.1;\.Trident/6\.0|facebookexternalhit|facebookexternalua|Version\.13\.0\.3\.Mobile\.15E148\.Safari\.604\.1|iPhone.*Version/13\.0\.3|Safari/604\.1|cms-scanner|OAI-SearchBot|ChatGPT-User|PerplexityBot|Perplexity-User|Amazonbot|Applebot-Extended|Meta-ExternalAgent|Meta-ExternalFetcher|cohere-ai|DeepSeek|Bytespider|Diffbot|Omgilibot|Omgili|Google-Extended|Google-CloudVertex|MistralAI-User|OAI-AdsBot|YouBot|anthropic-ai|NosibleBot) [NC,OR]
RewriteCond %{THE_REQUEST} "^[A-Z]{3,9}\s+https?://"
[NC]RewriteRule ^.*$ - [G,L]

# 2. BLOCK SPAM WEBSITES (REFERRERS)
RewriteCond %{HTTP_REFERER} (baidu\.com|bsky\.(net|com|app)|facebook\.com|meta\.com|threads\.(com|net)|instagram\.com|google\.com|googleusercontent\.com|youtube\.com|x\.com|t\.co|twitter\.com|x\.ai|bing\.com|yahoo\.com|yandex\.com|duckduckgo\.com|microsoft\.com|amazon\.com|brave\.com|semalt\.com|buttons-for-website\.com|darodar\.com|blackhatworth\.com|ilovevitaly\.com|priceg\.com|ranksonic\.com)
[NC]RewriteRule ^.*$ - [G,L]

###############################################
# SAFE BOT & SCRAPER KILLER (FRIENDICA-COMPATIBLE)
###############################################

# Kill obvious scanners by User-Agent
RewriteCond %{HTTP_USER_AGENT} (nmap|nikto|acunetix|sqlmap|fimap|nessus|openvas|arachni|wpscan|dirbuster|fuzzer)
[NC]RewriteRule ^ - [G,L]

# Kill requests with directory traversal attempts
RewriteCond %{QUERY_STRING} (\.\./|\.\.\\|%2e%2e|%5c)
[NC]RewriteRule ^ - [G,L]

# Kill malformed absolute URLs
RewriteCond %{THE_REQUEST} "^[A-Z]{3,9}\s+https?://"
[NC]RewriteRule ^ - [G,L]

# Kill empty User-Agent ONLY if NOT federation
RewriteCond %{HTTP_USER_AGENT} ^$
RewriteCond %{REQUEST_URI} !^/\.well-known/
[NC]RewriteRule ^ - [G,L] which has cut down a lot of bad traffic from getting 200, and getting hit with a 410 when it knocks on my door, I would start there, because blocking ips is fun and easy, but if they are being spoofed, they will just spoof more, and worse when they get recycled to actual valid users then you lose traffic, I have found the agent blocking is more effective, they are not willing to rewrite every script to adjust for that so for now it seems to be the fastest block, and if it is getting that hard, run it through cloudflare on the free side proxied for a while anyways to help slowdown and divert the bad traffic, it will keep your server happier, I know not everyone is a fan of that, but somtimes you have to change the route to stay on the road

in reply to utzer

@utzer @pasjrwoctx👽

Das hat einige Begleiterscheinungen, die man nicht möchte.

Prüf aus dem Log das tatsächliche Header-Profil des Floods (Accept-Language leer/nicht leer, Anzahl eindeutiger Sec-CH-UA/Accept-Kombinationen). Wenn da ein konstanter Unterschied zu echten Browsern ist, kann ich dir eine gezielte Regel bauen, die genau diesen Header-Set blockt statt der UA-Version. Das ist der einzige UA-nahe Filter, der hier sauber zwischen Flood und echten Nutzern trennt.

in reply to utzer

@utzer @pasjrwoctx👽 Der Flood rotiert UAs bewusst gleichverteilt durch 160 Identitäten – das ist ein Tool, das darauf gebaut ist, dass man es am UA erkennt. Hier wirst du also mit UA-Blocking nur die nächste der 160 Identitäten sehen, nicht weniger Traffic.

Was helfen kann ist TLS-Fingerprint und die Challenge-Ökonomie (Anubis-PoW-Difficulty).

in reply to utzer

@utzer Hat der Flood ein leeres Accept-Language? Der schnellste Test: einmal eine Stichprobe aus euren Anubis-Logs ziehen und schauen, ob bei den Top-UAs auch Accept-Language: leer ist. Wenn ja, reicht beim nginx
if ($http_accept_language = "") { return 444; }
444 = Verbindung sofort verwerfen, keine Antwort
Das musst du aus dem Log lesen, nicht raten oder probieren!
in reply to tom s

@tom s Es ist tatsächlich ein neuer dominanter User-Agent nachgerückt:

Android 6 / Nexus 5 / Chrome 65

In einer Stunde kamen damit 527 Requests von 526 unterschiedlichen IPv4-Adressen. Das sieht also weiterhin nach demselben rotierenden Proxy-Netz aus.

Das Headerprofil hat sich allerdings geändert:

vorher: Accept-Language: en-US,en;q=0.9
jetzt: Accept-Language: en-US,en;q=0.5

Priority: u=0, i ist gleich geblieben, wird aber auch von echten Browsern verwendet. Darauf kann ich daher nicht sauber filtern.

TLS wird bereits im Reverse-Proxy-Nginx terminiert. Anubis sitzt dahinter und erhält nur noch normales HTTP. Anubis kann in diesem Aufbau deshalb keine TLS-/JA3-Fingerprints ermitteln. Dafür müsste ich den ClientHello vor der TLS-Terminierung am Nginx separat erfassen.

in reply to utzer

@utzer tja, q=0.9 -> q=0.5 Das bedeutet meist, dass der Betreiber beobachtet und anpasst. Egal auf was an dieser Stelle.

nginx kann euch einen groben TLS-Fingerprint direkt ins Access-Log schreiben, das ist mehr als nix.

log_format tlsinfo '$remote_addr $time_local "$request" $status '
                   'proto=$ssl_protocol cipher=$ssl_cipher curves=$ssl_curves '
                   'ua="$http_user_agent" al="$http_accept_language"';
access_log /var/log/nginx/access.log tlsinfo;

nginx -t und reload.

Mein Chatfenster mach das leider ein wenig kaputt.

Chrome 65 ist gut gewählt, aber so kannst Du rausfinden, ob es überhaupt einer ist.

Eins noch, löst das Ding die Anubis-Challenges?
Im SLOG_LEVEL=DEBUG-Log von Anubis siehst Du das. Wenn er sie nicht löst, verbrennt er nur Anubis-CPU und es kann weg.

in reply to tom s

@tom s Danke für die Hinweise. Ich habe das jetzt in einer kurzen Stichprobe getestet:


  • UA-Filter vorübergehend deaktiviert
  • Anubis auf DEBUG gestellt
  • separates TLS-Light-Logging in Nginx aktiviert


Ergebnis für den angeblichen Android-6-/Chrome-65-Client:

Requests:              13
Unterschiedliche IPs:  13
Challenges ausgestellt: 11
Challenges gelöst:       0

Die 13 Verbindungen ergaben zunächst zwölf verschiedene TLS-Profile. Die Unterschiede bestanden allerdings ausschließlich aus zufälligen GREASE-Werten. Nach deren Normalisierung hatten alle Verbindungen exakt dasselbe TLS-Profil.

Der TLS-Fingerprint allein eignet sich trotzdem nicht für eine sichere Sperre: Dasselbe normalisierte Profil wurde in der Stichprobe auch von aktuellen Chrome-142-, Chrome-148- und Chrome-149-Clients verwendet.

Die Kombination ist jedoch eindeutig verdächtig: Der User-Agent behauptet Android 6 / Chrome 65, verwendet aber einen modernen TLS-1.3-/HTTP/2-Stack und über alle IP-Adressen hinweg dasselbe vollständige HTTP-Headerprofil.

Auch der User-Agent Mozilla/5.0 (compatible; crawler) löste keine seiner zwölf Challenges.

Der aktuelle Flood verbraucht in dieser Stichprobe also hauptsächlich Anubis-Ressourcen und gelangt nicht bis zu Forgejo.

UA-Filter und normales Anubis-Logging sind inzwischen wieder aktiv.

in reply to utzer

@utzer hmm, der Zeitraum ist eigentlich viel zu kurz.

Warum Der Anubis bei diesem bisschen Rauschen schon umfallen soll, ist sehr merkwürdig. Irgendwo ne Fehlconfig am Werk.
ca. 10 Challenges/s, das ist kein Flood im wirklichen Sinne.
DEBUG-Logs und METRICS_BIND-Metriken suchen -> Speicher-Anstieg, offene Verbindungen, OOM? Irgendwas hakt da.


/etc/nginx/conf.d/uaold13.conf:

map "$ssl_protocol $http_user_agent" $uaold13 {
    default 0;
    "~^TLSv1\.3 .*Chrome/[0-6][0-9]\."   1;
    "~^TLSv1\.3 .*Chrome/[0-9]\."        1;
    "~^TLSv1\.3 .*Firefox/[0-5][0-9]\."  1;
    "~^TLSv1\.3 .*Edge/1[0-7]\."         1;
}

Im server-Block:

if ($uaold13) { return 444; }

doas nginx -t && doas /etc/init.d/nginx reload

Das könntest Du als temporäre Spielerei aufnehmen. Würde halt nicht lange halten. Am besten erst eine Stunde oder Zwei nur loggen statt 444, um zu sehen, wen Du triffst!

Das ist also hier vermutlich nur ein Vorspiel.
Werden zukünftig Challanges gelöst, hat der Schlingel auf eine JS-Version umgeschwenkt. Erst dann wirds lustig. Bis dahin musst Du das Anubis genauer durchschauen.
Eindeutiges könntest Du auf DENY statt CHALLENGE stellen, das kann ich aus der Ferne nicht beurteilen.

in reply to utzer

@utzer guter Punkt. Für nginx sind ca. 185/s zwar ok, aber dahinter kann das enger werden.

Ok, würde vorschlagen, den uaold13 scharf stellen und zudem noch:

bots:
  - name: generischer-crawler
    user_agent_regex: 'Mozilla/5.0 \(compatible; crawler\)'
    action: DENY
  - name: antike-ua
    user_agent_regex: 'Chrome/([0-6][0-9]|[0-9])\.'
    action: DENY

Damit fällt schon mal Last ab.
Nun gilt es den Zähler von gelösten Challenges im Auge zu behalten.

Für das Beurteilen von Verbindungen, einfach mal nen Tag lang anschauen, was so passiert, z.B.

nft add rule inet filter input ct state new tcp dport 443 
counter

Wäre ein Anfang.

in reply to tom s

@utzer
hab das zuvor ein wenig unglücklich ausgedrückt. ich meinte die input rule mal in ruhigen, normalen Zeiten laufen lassen. so 4x am Tag für jeweils ne Stunde. Aber vielleicht kennst Du die normale Auslastung ja selber schon.

Je nachdem, was Deine Messung ergibt, könnte man sowas beispielsweise einbauen: Rate: 50/s und Burst: 150

table inet floodfilter {
    chain input {
        type filter hook input priority filter; policy accept

        ct state established,related accept

        # Limit auf neue Verbindungen zu 443
        ct state new tcp dport 443 limit rate 50/second burst 150 accept
        ct state new tcp dport 443 drop
        
    }
}

Ist nur ne Idee, um den Traffic zu deckeln.

in reply to utzer

Further update: I have replaced the temporary path-specific rate limit with an exact User-Agent filter.

The ten evenly rotated browser identities responsible for 96.79% of the Anubis challenges, plus the explicit `Mozilla/5.0 (compatible; crawler)` identity, are now rejected by Nginx before reaching Anubis or Forgejo.

The pull-request and issue pages are available normally again, regular web and Git access remain unaffected, and backend load is close to zero.

The crawler has since introduced a new identity:
Android 6 / Nexus 5 / Chrome 65

This identity produced 527 requests from 526 different IPv4 addresses within one hour, confirming the rotating proxy-network pattern. However, the overall volume reaching Anubis has dropped from roughly 35,800 challenges per hour at the peak to about 1,000 per hour, so the exact filtering is currently working very well.

This remains a temporary mitigation because User-Agent strings can be changed at any time.

in reply to utzer

@utzer @utzer ok so you could block those user agents I have something like
 # Fake browser detection
RewriteCond %{HTTP_USER_AGENT} (Chrome/[0-9]{3}|Chrome/1[3-9][0-9]|Chrome/150|Firefox/1[3-9][0-9]|Safari/60[0-9]|Version/17) [NC]
RewriteCond %{HTTP_ACCEPT_LANGUAGE} ^$
RewriteRule ^ - [G,L]

RewriteCond %{THE_REQUEST} "GET\shttp" [NC]
RewriteRule ^ - [G,L]  and then # 1. BLOCK BAD BROWSER NAMES / BOT FRAMEWORKS

RewriteCond %{HTTP_USER_AGENT} (CCBot|SearchEngineBot|Pandalytics|UCBrowser|ZoneProjectBot|Embarcadero\sURI\sClient|Xenu\sLink\sSleuth|siteradar|SignalsBot|fun-cert-watch|SERankingBacklinksBot|Pinterestbot|CMS-Checker|HeadlessChrome|Puppeteer|SeznamBot|Sogou|8LEGS|HTTrack|cherrypicker|AhrefsBot|BLEXBot|DotBot|MJ12bot|PetalBot|SemrushBot|BuiltWith|Viewer/99|Python|aiohttp|curl|Wget|libwww|Go-http-client|GeedoShopProductFinder|DuckDuckBot|node|IMJ-CompanyPage-Scraper|baidu|RootEvidence|NetAPI\sv1|Scrapy|Bingbot|SummalyBot|got|HUNT-Bot|CibraxScanner|SalesOS-CompanyVerifier|RecordedFuture|SurdotlyBot|panscient\.com|Xiaomi|Android.*Firefox|GPTBot|ClaudeBot|BardBot|LLMScraper|Firecrawl|Crawl4AI|ia_archiver|archive\.org_bot|Google|wp2shell|okhttp|Cortex-Xpanse|Mozilla\.5\.0\.compatible;\.MSIE\.10\.0;\.Windows\.NT\.6\.1;\.Trident/6\.0|facebookexternalhit|facebookexternalua|Version\.13\.0\.3\.Mobile\.15E148\.Safari\.604\.1|iPhone.*Version/13\.0\.3|Safari/604\.1|cms-scanner|OAI-SearchBot|ChatGPT-User|PerplexityBot|Perplexity-User|Amazonbot|Applebot-Extended|Meta-ExternalAgent|Meta-ExternalFetcher|cohere-ai|DeepSeek|Bytespider|Diffbot|Omgilibot|Omgili|Google-Extended|Google-CloudVertex|MistralAI-User|OAI-AdsBot|YouBot|anthropic-ai|NosibleBot) [NC,OR]
RewriteCond %{THE_REQUEST} "^[A-Z]{3,9}\s+https?://" [NC]
RewriteRule ^.*$ - [G,L]

# 2. BLOCK SPAM WEBSITES (REFERRERS)
RewriteCond %{HTTP_REFERER} (baidu\.com|bsky\.(net|com|app)|facebook\.com|meta\.com|threads\.(com|net)|instagram\.com|google\.com|googleusercontent\.com|youtube\.com|x\.com|t\.co|twitter\.com|x\.ai|bing\.com|yahoo\.com|yandex\.com|duckduckgo\.com|microsoft\.com|amazon\.com|brave\.com|semalt\.com|buttons-for-website\.com|darodar\.com|blackhatworth\.com|ilovevitaly\.com|priceg\.com|ranksonic\.com) [NC]
RewriteRule ^.*$ - [G,L]

###############################################
# SAFE BOT & SCRAPER KILLER (FRIENDICA-COMPATIBLE)
###############################################

# Kill obvious scanners by User-Agent
RewriteCond %{HTTP_USER_AGENT} (nmap|nikto|acunetix|sqlmap|fimap|nessus|openvas|arachni|wpscan|dirbuster|fuzzer) [NC]
RewriteRule ^ - [G,L]

# Kill requests with directory traversal attempts
RewriteCond %{QUERY_STRING} (\.\./|\.\.\\|%2e%2e|%5c) [NC]
RewriteRule ^ - [G,L]

# Kill malformed absolute URLs
RewriteCond %{THE_REQUEST} "^[A-Z]{3,9}\s+https?://" [NC]
RewriteRule ^ - [G,L]

# Kill empty User-Agent ONLY if NOT federation
RewriteCond %{HTTP_USER_AGENT} ^$ 
RewriteCond %{REQUEST_URI} !^/\.well-known/ [NC]
RewriteRule ^ - [G,L] 

which has cut down a lot of bad traffic from getting 200, and getting hit with a 410 when it knocks on my door, I would start there, because blocking ips is fun and easy, but if they are being spoofed, they will just spoof more, and worse when they get recycled to actual valid users then you lose traffic, I have found the agent blocking is more effective, they are not willing to rewrite every script to adjust for that so for now it seems to be the fastest block, and if it is getting that hard, run it through cloudflare on the free side proxied for a while anyways to help slowdown and divert the bad traffic, it will keep your server happier, I know not everyone is a fan of that, but somtimes you have to change the route to stay on the road

This website uses cookies. If you continue browsing this website, you agree to the usage of cookies.