Last updated: 2026-09-07T00:13:48Z | Log entries analyzed: 319128 | Model: glm-5.2:cloud | Enrichment: Censys
What changed
NOTE: This script runs the day after data collection. The report reflects the site state as of 2026-09-06, not the current day.
Report day: 2026-09-06. Incremental run. The site saw 406 new IPs.
Scanner traffic reached 1,101 requests, 2.22x the trailing 7-day average of 496. AI crawler traffic surged to 1,081, 3.53x the average of 306. Search crawler traffic hit 374, 3.94x the average of 95. Visitor traffic rose to 891, 1.88x the average of 474. Fediverse traffic fell to 690, 0.33x the average of 2,109. Other crawler traffic dropped to 95, 0.35x the average of 273. RSS reader traffic held near baseline at 328, 1.34x the average of 245.
New scanner IPs of note:
- 51.77.70.237: 446 requests, 274 404s, first seen 2026-09-06 04:43:50.842+00. The IP used a truncated Chrome UA that ends at
AppleWebKit/537.36with no browser engine or version string. - 169.40.142.138: 240 requests, 126 404s, first seen 2026-09-06 04:43:51.199+00. Same truncated Chrome UA. The start time and UA match 51.77.70.237, which suggests coordination.
- 68.183.9.16 and 138.68.86.32: 31 requests each, first seen 2026-09-06 06:35:10. Both used the
l9scan/2.0UA from leakix.net. Censys enrichment for 138.68.86.32 shows ASN14061 DigitalOcean in Germany, rDNSb69efeaf93.scan.leakix.org, ports 22/SSH and 80/HTTP, GreyNoise classification malicious. Enrichment for 68.183.9.16 was unavailable. - 91.148.245.81: 34 requests, 25 404s,
Go-http-client/1.1UA, first seen 2026-09-06 21:17:15.573+00.
New recon URIs on the report day include config file probes: /includes/config_mail.php, /config/local.php, /config.php~, /config/production.php, /config/packages/mailer.yaml, /site/config.php, /conf/config.php — 4 hits each from 2 IPs. A GitHub workflow file /.github/workflows/npm-publish.yml drew 5 hits from 3 IPs. JavaScript bundle recon appeared: /static/js/runtime-main.js, /static/js/runtime~main.js, /static/js/vendors.js, /static/js/0.chunk.js, /static/js/1.chunk.js, /assets/js/app.js, /src/main.js — 2 hits each from 1 IP. Environment file recon continued with /environment.json, /assets/env.json, /static/env.json, /assets/settings.json, /assets/app-config.json. Admin panel probes hit /panel and /cp.
New user agents:
Claude-User (claude-code/2.1.263; +https://support.anthropic.com/)— ai_crawler class, 3 requests from 2 IPs. Known Anthropic agent with an updated version string.Claude-User (claude-code/2.1.257; +https://support.anthropic.com/)— ai_crawler class, 1 request from 1 IP. Known Anthropic agent, different version.AdsBot-Google-Mobile(compatible; +http://www.google.com/mobile/adsbot.html) — other_crawler class, 1 request from 1 IP. Known Google ad crawler.CragCrawler/1.0 (+https://cragsoftware.com)— ai_crawler class, 1 request from 1 IP. Unknown crawler. It may need a classifier rule if it returns.
Canary hits: the /.env decoy returned 200 to 438 requests from 200 distinct IPs through the observation window. The /.git/config decoy returned 200 to 204 requests from 123 IPs. Fourteen .env variant paths each drew between 37 and 170 hits. The most recent canary trigger was 2026-09-06 23:13:05.016+00 on /.env.
Classification gaps: no data available.
Traffic summary
The observation period spans 2026-03-10 to 2026-09-07, covering 319,128 requests from 45,785 unique IPs. The site served 7.06 GB of bytes. Median request duration was 4.55 ms, with the 95th percentile at 28.22 ms.
Fediverse link-preview fetches dominate at 48.4% of all traffic (154,383 requests from 8,384 IPs). Visitors account for 16.1% (51,309 requests). Scanners produce 14.5% (46,167 requests). AI crawlers generate 8.6% (27,471 requests). Other crawlers add 5.3% (17,060). RSS readers contribute 4.0% (12,784). Search crawlers make up 2.7% (8,723). Owner traffic is 0.4% (1,231).
The signal-to-noise ratio is rough. If you count fediverse previews and scanner probes as noise, legitimate human and feed traffic is about 20% of requests. The rest is automated fetches of one kind or another.
Resource consumption
Bandwidth figures reflect bytes transferred. Conditional and cached responses show 0 bytes, so totals undercount logical content size.
Fediverse fetches consumed 4.54 GB (64.3% of total bandwidth) across 154,383 requests. Visitors consumed 1.31 GB (18.6%). AI crawlers took 482 MB (6.8%). Other crawlers used 381 MB (5.4%). Search crawlers consumed 167 MB (2.4%). Scanners used 121 MB (1.7%). RSS readers took 27 MB (0.4%).
The fediverse bandwidth share is disproportionate because each link-preview fetch pulls the full HTML of /. These are one-shot requests, but they add up across 8,384 instances.
Temporal patterns
Scanner traffic concentrates in the late evening and overnight UTC. Peak scanner hours are 23:00 (3,688 requests), 00:00 (5,092), 03:00 (4,559), and 21:00 (2,860). The lowest scanner activity is at 09:00 (460) and 11:00 (301).
Visitor traffic follows a daytime pattern. It peaks at 14:00 (2,637) and 17:00 (2,792), with a steady floor around 1,700-2,000 overnight. The visitor curve is flatter than the scanner curve.
Fediverse traffic has a sharp midday spike. Hours 11:00 (21,526), 12:00 (16,873), 13:00 (13,748), and 14:00 (15,270) account for a large fraction of all fediverse requests. Hour 10:00 also shows 6,607. This pattern suggests link shares happen during European afternoon hours.
AI crawler traffic is uniform across hours, with no strong time-of-day preference. It ranges from 722 to 1,687 requests per hour.
By day of week, scanner traffic peaks on Saturday (13,551 requests) and Sunday (8,217). Tuesday is also high at 7,181. Visitor traffic peaks on Monday (9,227) and Thursday (8,528). Fediverse traffic is heaviest on Monday (40,474), which dwarfs every other day.
Content & visitors
The most-read post is /posts/2026-06-27-running-ornith-locally-with-opencode-and-claude-code/ with 2,429 hits from 2,354 visitors. The root page / drew 3,369 hits from 2,148 visitors. The weekly report post /posts/2026-06-20-weekly-bulletproof-report/ attracted 1,695 hits. The /posts/observatory/ page drew 947 hits from 748 visitors.
Referrer data shows Google as the top source with 1,816 referrals. Twitter sent 406. Hacker News sent 394. Reddit sent 389. Facebook sent 366. DuckDuckGo sent 272. Mastodon sent 98. Bing sent 91. The site’s own domain (rud.is) accounts for 68, and infosec.exchange sent 51.
HTTP/3 adoption among visitors is low: 3,088 of 51,309 visitor requests used HTTP/3 (6.0%). The owner class shows 918 of 1,231 requests on HTTP/3 (74.5%). AI crawlers almost never use HTTP/3 — only 2 of 27,471 requests. Fediverse fetches are almost entirely HTTP/1.1 (154,322 of 154,383). RSS readers show 11,201 on HTTP/2 and 5 on HTTP/3.
Browser families among visitors: Chrome leads with 36,595 requests from 21,607 IPs. Safari follows with 5,213 from 1,903 IPs. Firefox accounts for 4,360 from 1,415 IPs. Edge contributes 1,630 from 941 IPs. Opera adds 221 from 161 IPs.
RSS feed activity shows 11,973 requests from the rss_reader class across 85 unique IPs. The feed started on 2026-04-20 and remains active. A small but consistent subscriber base exists.
Agent-artifact access patterns
The site serves six machine-readable formats for agents. Post .md files dominate agent-artifact traffic with 4,207 requests from 2,255 IPs, which transferred 17.2 MB. The llms.txt index drew 203 requests from 149 IPs, which transferred 1.36 MB. The identity.json file received 6 requests. The llm.txt variant got 4 requests. The ai.json metadata file and llms.html index each received 2 requests.
Post .md files get the most use across all traffic classes. The top .md file is /posts/observatory.md with 184 requests (108 visitor, 28 AI crawler, 13 search crawler). The /posts/2026-06-27-running-ornith-locally-with-opencode-and-claude-code.md file drew 182 requests (99 AI crawler, 45 visitor, 25 search crawler). AI crawlers are the primary consumers of .md content for most posts, which aligns with the design intent.
The llms.txt index sees traffic from search crawlers and AI crawlers. Baiduspider made 187 requests and accessed 51 artifacts. PetalBot made 177 requests and accessed 36 artifacts. Amazonbot made 147 requests and accessed 51 artifacts. GPTBot made 95 requests and accessed 48 artifacts. ClaudeBot made 48 requests and accessed 47 artifacts.
The less-used artifacts (ai.json, llms.html, identity.json, llm.txt) each received fewer than 10 requests total. These formats exist but see little use among agents.
Compared to HTML page views, agent-artifact traffic is modest. The top HTML page drew 2,429 hits. The top .md file drew 184. The llms.txt index drew 203. Agents request the .md and llms.txt formats, but the volume is a fraction of HTML traffic.
AI crawler activity
ChatGPT-User (OpenAI) leads with 3,837 requests from 1,600 IPs, across 395 unique pages. ClaudeBot (Anthropic) follows with 3,355 requests from 154 IPs across 449 pages. Bytespider (ByteDance) made 2,967 requests from 909 IPs across 320 pages. Amazonbot (Amazon) made 2,482 requests from 439 IPs across 583 pages. OAI-SearchBot (OpenAI) made 1,641 requests from 182 IPs. GPTBot (OpenAI) made 1,473 requests from 69 IPs. PetalBot (Huawei) made 1,435 requests from 22 IPs. PerplexityBot made 1,325 requests from 32 IPs. Meta’s crawler made 1,201 requests from 667 IPs. Applebot made 966 requests from 563 IPs.
On the report day, AI crawler traffic was 3.53x the 7-day average. This is a notable spike for a small personal blog. The bandwidth cost over the full period is 482 MB, which is manageable but not trivial.
Smaller AI crawlers include CCBot (678 requests), AIWebIndex (443), Velen Web Crawler (436), DuckAssistBot (389), and a long tail of niche crawlers. DeepSeekBot made 93 requests. xAI-SearchBot made 76. YouBot made 54. Cohere made 6. MistralAI-User made 8.
Search engine crawlers
Bingbot made 2,500 requests from 387 IPs across 419 pages, last seen 2026-09-06 23:29:10.522+00. Googlebot made 2,015 requests from 393 IPs across 382 pages, last seen 2026-09-07 00:02:44.946+00. Baiduspider made 1,803 requests from 357 IPs across 439 pages. DuckDuckBot made 1,078 requests from 71 IPs across 153 pages. Yandex made 580 requests from 283 IPs across 166 pages.
Search crawler traffic on the report day was 3.94x the 7-day average. The major engines are crawling the site. Googlebot and Bingbot both hit the site within the last hour of the observation window. Baiduspider covers more unique pages (439) than Googlebot (382), which is unusual for a site that targets an English-speaking audience.
Fediverse activity
Fediverse link-preview traffic is the largest single traffic class. The top instances by request count:
- mastodon.n41.lat: 126 requests from 1 IP, 26 unique URIs
- evy.pet (Misskey): 125 requests from 2 IPs, 17 unique URIs
- social.louis-vallat.dev (Misskey): 124 requests from 2 IPs, 12 unique URIs
- mastodon.distos.org: 120 requests from 1 IP, 26 unique URIs
- federation.network (Misskey): 111 requests from 1 IP, 8 unique URIs
Fediverse traffic dropped to 0.33x the 7-day average on the report day. This is a 67% decline from the trailing average of 2,109 requests. The cause is unmeasured. The midday UTC spike pattern in the hourly data suggests this traffic correlates with European daytime social activity.
Scanner & recon activity
The top scanner by lifetime volume is 185.219.151.78 with 9,629 requests and 8,629 404s across 7,813 unique URIs. It used a Chrome 113 UA and was active from 2026-06-26 to 2026-06-28. The IPv6 address 2a01:239:248:1300::1 made 6,003 requests with 2,712 404s, with the python-requests/2.34.2 UA.
On the report day, the new scanner 51.77.70.237 made 446 requests in 29 seconds (04:43:50 to 04:44:19), with 274 returning 404. The IP 169.40.142.138 made 240 requests in the same window. Both used a truncated Chrome UA that ends at AppleWebKit/537.36 with no browser engine or version. The matching start times and identical UA suggest a coordinated burst scan.
The leakix scanner pair (68.183.9.16 and 138.68.86.32) used the l9scan/2.0 UA. Censys enrichment for 138.68.86.32 shows ASN14061 DigitalOcean in Germany, rDNS b69efeaf93.scan.leakix.org, ports 22/SSH and 80/HTTP, and GreyNoise classification malicious.
Recon URI patterns show WordPress probing (/wp-admin/admin-ajax.php, /blog/wp-json/batch/v1, /wordpress/wp-json/batch/v1), PHPUnit exploitation (/vendor/phpunit/phpunit/src/Util/PHP/eval-stdin.php), and environment file hunting (/.env, /api/.env). The /nuclei.svg path with a query parameter suggests a Nuclei template validation probe. Authentication endpoint recon hit /signin, /api/auth/signin, and /account.
The canary paths are working. The /.env decoy returned 200 to 438 requests from 200 distinct IPs. The /.git/config decoy returned 200 to 204 requests from 123 IPs. Fourteen .env variant paths each drew hits. The canary paths served a combined 282 KB of decoy content. No real environment data was exposed.
IP persistence
The longest-persisting IP is 70.8.208.88, seen on 65 days with 630 total requests, first seen 2026-06-26 and last seen 2026-09-06. The IP 43.155.125.36 appeared on 30 days with 38 requests. The IP 92.247.181.45 appeared on 28 days with 136 requests. The IP 116.128.185.53 appeared on 28 days with 35 requests.
Several IPs in the 220.196.160.x range show persistence across 18-24 days. These are likely a single network or organization with addresses that rotate. IP persistence does not equal individual identity because IPv6 privacy extensions and NAT hide the true count.
Serendipity
The canary setup is the most clever element in this data. The site operator configured /.env, /.git/config, and 14 .env variant paths to return 200 with decoy content. Scanners that grab these files get garbage. The canary paths drew 1,642 total hits from hundreds of distinct IPs across the observation period, which means the decoy absorbed real scanner attention. The bytes served are small — 81,920 bytes for /.env across 438 hits — so the bandwidth cost is negligible. This is a well-designed passive defense for a static site.
Observations
The coordinated scanner pair 51.77.70.237 and 169.40.142.138 fired 686 requests in under 30 seconds with matching start times and identical truncated UAs. This is a burst scan, not a sustained campaign. The truncated UA (which ends at AppleWebKit/537.36 with no Chrome or Safari version) is a distinctive fingerprint. It can be added to the scanner classifier.
CragCrawler/1.0 appeared once from one IP. It is unknown to the classifier. It may need a rule if it returns. The two new Claude-User versions (claude-code/2.1.257 and 2.1.263) are known Anthropic agents with updated version strings. The classifier will match these without changes if it keys on the Claude-User prefix.
The fediverse traffic drop on the report day (0.33x average) is the largest volume decline in the dataset. The cause is unmeasured. If this decline persists, it will reduce the site’s dominant traffic class and its largest bandwidth consumer.
Generated by observatory.sh — Caddy logs → DuckDB → Censys → Ollama → Astro