Last updated: 2026-07-10T00:12:44Z | Log entries analyzed: 135714 | Model: glm-5.2:cloud | Enrichment: Censys
Site Observatory — ai.rud.is
What Changed
Report day: 2026-07-09 (incremental run, 462 new IPs). This report reflects the state of the site as of the previous day — analysis ends 2026-07-09, not the current day.
Scanner traffic spiked hard: 235 requests today vs a 7-day average of 64.4 (ratio 3.65). The rest of the traffic classes are below their trailing averages — fediverse at 0.65× (1,343 vs 2,080), visitor at 0.66× (577 vs 868), other_crawler at 0.76×, search_crawler slightly above at 1.22×, and AI crawlers roughly flat at 0.93×. A quieter day for everything except people poking at the server.
Newly-seen scanners:
- 100.57.165.217 — 180 requests today, 176 of them 404s, all within an 11-second window (10:42:40–10:42:51 UTC). Chrome 131 UA on Linux. Classic burst-scan: 177 unique URIs, almost all misses. No enrichment available for this IP.
- 46.101.1.225 — 31 requests, 23 404s, UA
Mozilla/5.0 (l9scan/2.0.731313e21353e23393e2237313; +https://leakix.net). LeakIX scanner, first seen 17:33 UTC. DigitalOcean range (no enrichment line provided, but 46.101.x is DO). - 107.150.120.129 — 2 requests, UA includes
Assetnote/1.0.0. Assetnote recon, light touch. - 129.146.112.148, 34.90.254.162, 34.91.129.93 — all Scrapy/2.16.0, 1–2 requests each. Oracle Cloud IPs doing low-volume crawl recon.
- 141.11.62.23 — 2 requests, 2 404s, Firefox 149 UA. Likely a manual or scripted probe.
- 2804:8c64:10c1:2000:8233:ca7a:d8c6:998c — single request, UA
Mozilla/5.0 crawler. Brazilian IPv6, one hit. - 123.245.85.163 — single request,
Go-http-client/1.1. Chinese IP, one-shot probe.
No new recon URIs were observed for the first time today — the recon landscape is stable, which means scanners are rehashing the same paths they always do.
Canary hits: Both decoy paths are still being served 200s and still getting hit. /.env saw 124 hits from 57 distinct IPs (last 2026-07-09 19:31:35 UTC, 23,202 bytes served). /.git/config saw 58 hits from 41 IPs (last 2026-07-09 17:34:18 UTC, 696 bytes). These are active canary triggers — the decoys are doing their job, absorbing scanner attention with minimal bandwidth cost.
No classification gaps detected — no bot-like UAs slipped through to visitor this run.
Traffic Summary
The full observation window spans 2026-03-10 through 2026-07-10: 135,714 total requests from 20,153 unique IPs, serving 2.81 GB. Median response time is 4.28 ms (p95: 25.91 ms) — this is a fast static site on Caddy and it shows.
The signal-to-noise ratio is rough. Fediverse link-preview fetches alone account for 48% of all traffic (65,189 requests from 5,533 IPs). Legitimate visitor traffic is 21.5% (29,165 requests). Everything else — AI crawlers (8.7%), scanners (8.4%), other crawlers (5.7%), RSS readers (3.7%), search crawlers (3.0%) — is automated. The owner’s own traffic is 0.9%.
Status codes: 77.4% 200, 13.6% 404, 5.1% 308 (trailing-slash redirects), 3.2% 304 (conditional gets), a handful of 405s (271), and 12 500s across the entire window.
Resource Consumption
| Traffic Class | Bytes Served | % of Total | Requests |
|---|---|---|---|
| fediverse | 1,844,594,102 | 65.7% | 65,189 |
| visitor | 418,569,755 | 14.9% | 29,165 |
| ai_crawler | 219,632,826 | 7.8% | 11,783 |
| other_crawler | 187,999,655 | 6.7% | 7,764 |
| search_crawler | 79,921,296 | 2.8% | 4,063 |
| owner | 32,253,302 | 1.1% | 1,231 |
| scanner | 17,271,012 | 0.6% | 11,467 |
| rss_reader | 8,795,865 | 0.3% | 5,052 |
Fediverse fetches are the dominant bandwidth consumer at 1.84 GB — these are link-preview renders pulling the full HTML of / every time someone shares a URL on Mastodon. Scanner traffic is remarkably cheap: 11,467 requests for only 17 MB, because almost everything they hit returns 404 (zero or minimal body). Bandwidth figures reflect bytes actually transferred; conditional/cached responses (304s) show 0 bytes, so logical content size is undercounted.
Temporal Patterns
Fediverse traffic has a pronounced midday spike in UTC: hours 10–14 see the highest fediverse volume (4,065 at 10:00, peaking at 7,657 at 11:00, 6,229 at 13:00). This maps to European afternoon and early US morning — when people are actively sharing links on Mastodon. Fediverse drops to its lowest around 02:00–06:00 UTC (244–371 requests/hour).
Visitor traffic has a different cadence: it peaks in the late evening UTC (3,329 at 23:00, 3,731 at 00:00, 2,148 at 22:00) and has a secondary morning bump (1,767 at 05:00, 1,710 at 07:00). The evening peak suggests a North American audience — 23:00 UTC is 19:00 Eastern.
Scanner traffic is bursty and unpredictable but has notable spikes at 08:00 (1,394), 13:00 (1,316), 18:00 (1,367), and 19:00 (1,543) UTC. These are individual campaign bursts, not a steady cadence. The 08:00 spike likely corresponds to the 100.57.165.217 burst-scan at 10:42 (though the hourly aggregation would put it in the 10:00 bucket — the 08:00 spike is probably a different campaign).
AI crawler traffic is remarkably flat across all hours (307–730 requests/hour), consistent with distributed crawling that runs 24/7 without regard for time zones.
Day-of-week: Tuesday is the worst scanner day (5,687 requests, driven by the June 2 burst of 5,292). Saturday is the highest visitor day (6,737), and Monday leads fediverse (11,898). Thursday is the top AI crawler day (1,964).
Content & Visitors
The most-read post by visitor count is /posts/2026-06-27-running-ornith-locally-with-opencode-and-claude-code/ — 594 hits, 570 visitors. The observatory page itself (/posts/observatory/) drew 501 hits from 385 visitors. The ollama usage posts (both the original and the enhanced version) remain popular, as does the Apple container machine post and the opencode-go-usage post.
Referrers are dominated by Google (568 + 32 = 600 hits), followed by DuckDuckGo (88), then self-referrals from rud.is (42 + 23 = 65). Mastodon.social sent 59, phanpy.social 15, and infosec.exchange 13 — the fediverse is a real referral source. One curiosity: 12 referrals from https://app.qiuchong.com.cn:443/phpmyadmin/index.php?route=/database/structure{{REFERRERS}}db=zhizhen — a spoofed referrer from a phpMyadmin instance, almost certainly a scanner or bot trying to create backlink noise.
HTTP/3 adoption: 2,243 requests used h3 (1.7% of all traffic). Among visitors specifically, 1,270 of 29,165 requests used HTTP/3 (4.4%). The owner’s own traffic is overwhelmingly h3 (918 of 1,231). AI crawlers barely touch it (2 requests), and fediverse is almost entirely HTTP/1.1 (65,187 of 65,189). RSS readers are mostly HTTP/2 (4,647 of 5,052). The TLS negotiation data shows 67.7% of requests have an empty negotiated field — likely HTTP/1.1 over TLS without ALP.
Browser families among visitors: Chrome leads (16,554 requests, 6,528 IPs), Safari (5,253), Firefox (4,516), Edge (729), Opera (184), and 1,929 “Other.”
RSS feed activity: 4,625 requests from 31 unique IPs in the rss_reader class, spanning 2026-04-20 to 2026-07-10. That’s a small but consistent subscriber base — roughly 31 feed readers polling regularly. AI crawlers also hit the feed 63 times from 58 IPs.
Agent-Artifact Access Patterns
The site serves six machine-readable artifact types. Here’s what’s actually being used:
| Artifact | Requests | Unique IPs | Bytes Transferred |
|---|---|---|---|
| post-md | 1,499 | 813 | 4,764,321 |
| llms.txt | 68 | 55 | 246,340 |
| llm.txt | 2 | 2 | 4,230 |
| llms.html | 2 | 2 | 841 |
| ai.json | 2 | 2 | 518 |
| identity.json | 2 | 2 | 376 |
Post .md versions are the clear winner — 1,499 requests transferring 4.76 MB. The top .md files are /posts/2026-04-04-ollama-usage.md (72 requests), /posts/2026-05-23-starlog-and-the-case-of-the-missing-feed.md (70), and /posts/2026-05-23-making-airudis-legible-to-machines.md (65). AI crawlers account for the plurality of .md hits on most posts, but visitors pull a surprising number too — on /posts/observatory.md, visitors (25) actually outnumber AI crawlers (12).
llms.txt has real traction: 68 requests from 55 IPs, 246 KB transferred. The other four artifacts (llm.txt, llms.html, ai.json, identity.json) each have exactly 2 requests from 2 IPs, all first seen 2026-06-12 09:24:03 and last seen 2026-07-01 03:06:14 — these look like a single agent’s initial discovery crawl, not ongoing use.
Who’s consuming agent artifacts? Amazonbot leads with 68 requests accessing 35 distinct artifacts — the most thorough agent consumer. ClaudeBot accessed 32 artifacts across 32 requests (100% artifact hit rate per request). GPTBot hit 28 artifacts in 31 requests. Bytespider (both variants) accessed 19–21 artifacts each. PetalBot hit 22. OAI-SearchBot only 17. Applebot 22. Meta’s crawler 17. CCBot 17.
Notably, Barkrowler (other_crawler, 85 requests, 21 artifacts) and SemrushBot (53 requests, 31 artifacts) are consuming agent artifacts more aggressively than some AI crawlers. MJ12bot accessed 30 artifacts in 35 requests. Even Baiduspider pulled 31 artifacts in 65 requests. The llms.txt convention is being picked up by traditional SEO crawlers, not just LLM trainers.
The scanner TLM-Audit-Scanner/1.0 hit 9 artifacts in 72 requests — it’s not just doing path brute-force, it’s also cataloging the agent-facing surface. And Scrapy/2.16.0 accessed 22 artifacts across 45 requests — someone is building a corpus.
For context, the top HTML page (/) got 2,183 hits from 1,134 visitors. The .md versions collectively got 1,499 requests. That’s a meaningful ratio — for every ~1.5 HTML page views, there’s one markdown fetch. The agent-facing surface is not a vanity feature; it’s getting real use.
AI Crawler Activity
Fifteen distinct AI crawlers have hit the site. By volume:
| Crawler | Requests | Unique IPs | Unique Pages |
|---|---|---|---|
| Bytespider (ByteDance) | 2,489 | 854 | 230 |
| Other AI | 2,220 | 927 | 319 |
| ClaudeBot (Anthropic) | 1,837 | 71 | 269 |
| Amazonbot (Amazon) | 1,303 | 402 | 341 |
| OAI-SearchBot (OpenAI) | 894 | 103 | 208 |
| GPTBot (OpenAI) | 623 | 28 | 327 |
| PetalBot (Huawei) | 565 | 22 | 193 |
| PerplexityBot (Perplexity) | 542 | 15 | 235 |
| Meta Crawler (Meta) | 539 | 349 | 226 |
| Applebot (Apple) | 506 | 331 | 212 |
| CCBot (Common Crawl) | 251 | 15 | 174 |
| Timpibot | 6 | 5 | 2 |
| Diffbot | 4 | 4 | 2 |
| YouBot (You.com) | 2 | 1 | 2 |
| Cohere Crawler | 2 | 1 | 2 |
Bytespider leads in raw volume (2,489 requests from 854 IPs — a massive IP spread suggesting distributed crawling infrastructure). But ClaudeBot is the most concentrated: 1,837 requests from only 71 IPs, hitting 269 unique pages. That’s ~26 requests per IP — deep, methodical crawling from a small fleet. GPTBot is even more concentrated: 623 requests from 28 IPs, 327 unique pages — 22 requests/IP with the broadest page coverage of any crawler.
Amazonbot has the widest page coverage (341 unique pages from 402 IPs) — it’s doing the most thorough index of the site. PerplexityBot is efficient: 542 requests from 15 IPs, 235 pages.
For a small personal blog, this is a lot of AI attention. The 11,783 AI crawler requests represent 8.7% of all traffic and 219 MB of bandwidth. The “Other AI” bucket (2,220 requests from 927 IPs) is worth watching — that’s a long tail of unidentified AI agents that may need classifier rules if patterns emerge.
Search Engine Crawlers
| Crawler | Requests | Unique IPs | Unique Pages | Last Seen |
|---|---|---|---|---|
| Googlebot | 970 | 137 | 278 | 2026-07-09 18:15:08 |
| Bingbot | 967 | 263 | 183 | 2026-07-09 18:55:24 |
| Baiduspider | 887 | 297 | 293 | 2026-07-09 20:44:14 |
| Other Search | 641 | 111 | 243 | 2026-07-10 00:09:51 |
| DuckDuckBot | 488 | 68 | 70 | 2026-07-09 14:21:32 |
| Yandex | 110 | 84 | 20 | 2026-07-06 04:21:28 |
Google and Bing are actively indexing — both were last seen on the report day. Baiduspider is surprisingly aggressive (887 requests, 293 pages) and was the most recent (20:44 UTC). DuckDuckGo is present but light (488 requests, only 70 unique pages — it’s not deep-crawling). Yandex has been absent since July 6. The site is being well-crawled by the major Western engines and Baidu; Yandex is largely ignoring it.
Fediverse Activity
Fediverse traffic is the single largest traffic class at 48% of all requests. The top 20 instances by request volume are a mix of Mastodon and Misskey:
- transfem.social (Misskey) — 62 requests, 11 unique URIs
- evy.pet (Misskey) — 59 requests, 9 URIs
- federation.network (Misskey) — 58 requests, 8 URIs
- mastodon.n41.lat (Mastodon) — 56 requests, 13 URIs
- patrickflynn.me (Mastodon) — 55 requests, 13 URIs
- is-a.cat (Mastodon glitch) — 55 requests, 13 URIs
- social.medienzentrum-hdh.de — 55 requests, 13 URIs
- mastodon.distos.org — 55 requests, 13 URIs
Most instances fetch 13 unique URIs (the typical link-preview pattern: / plus a handful of post URLs). The 13-URI pattern suggests someone shared a post that links to several other posts, and the instance fetched all of them for card rendering. The Misskey instances tend to fetch fewer URIs (1–9), suggesting simpler share patterns.
All of these are automated http.rb or Misskey fetchers — benign by definition, but they account for 1.84 GB of bandwidth. That’s the cost of being shareable on the fediverse: every share triggers a full-page fetch.
Scanner & Recon Activity
The all-time scanner leaderboard is dominated by a few large campaigns:
- 45.148.10.95 — 3,816 requests, 1,704 404s, 636 unique URIs, UA
TLM-Audit-Scanner/1.0. Active 2026-06-02 only, an 11-hour burst. No enrichment available. - 185.219.151.78 — 1,914 requests, 1,763 404s, 1,196 unique URIs. Firefox 129 UA (spoofed). Active 2026-06-27 to 06-28, a 3-hour window. High URI diversity suggests a content discovery sweep.
- 195.178.110.199 — 1,456 requests, 752 404s, 636 unique URIs. Chrome 131 UA. Active 2026-05-10 to 06-02.
- 185.177.72.38 — 662 requests, 661 404s, 662 unique URIs (every request a different URI, every one a 404).
curl/8.7.1. Active 2026-03-11, a 74-second burst. Pure brute-force path enumeration. - 172.94.9.253 — 530 requests, 471 404s, 30 unique URIs. Firefox 124 UA. Per Censys enrichment: [benign], ASN213790 LimitedNetwork-AS (Limited Network LTD, United Kingdom), no rDNS, no observed ports, no service labels. The “benign” verdict is likely because Censys hasn’t observed malicious activity from this IP, but the behavior — 530 requests with 89% 404 rate across 30 URIs over a 6-week window — is scanner-like. This is a UK hosting provider, not residential.
Today’s new scanner 100.57.165.217 did 180 requests in 11 seconds — 177 unique URIs, 176 404s. That’s a burst scan at ~16 requests/second, which is aggressive but not unusual for automated recon tools.
The recon URI list is the usual suspects: /wp-admin/admin-ajax.php (103 hits from 1 IP — a single-minded WordPress probe), /.env (67 hits from 41 IPs — broad .env hunting), /.git/config (49 hits from 45 IPs), /backend/.env and /api/.env (42 and 41 hits respectively). The .env variants are the most distributed — 41–45 distinct IPs each, meaning many different scanners are all looking for the same thing. The /nuclei.svg?VaOD6=x path (102 hits from 1 IP) is a Nuclei scanner fingerprint probe.
The /CDGServer3/SystemConfig path (32 hits from 1 IP) is interesting — that’s a specific recon for a CDG document management system, not a generic web app. Someone is looking for a specific vulnerability class.
IP Persistence
Top persistent IPs (approximation only — IPv6 privacy extensions and NAT mean IP ≠ identity):
- 47.160.48.231 — 19 days seen, 111 total requests, 2026-04-21 to 2026-07-08. Likely a regular reader or a monitoring service.
- 43.155.125.36 — 19 days, 25 requests, 2026-06-07 to 2026-07-09. Tencent Cloud IP, light but consistent.
- 116.128.185.53 — 17 days, 21 requests. Chinese IP, sparse but persistent.
- 70.8.208.88 — 14 days, 206 requests, 2026-06-26 to 2026-07-09. 206 requests over 14 days is ~15/day — a daily reader or an automated checker.
- 92.247.181.45 — 16 days, 82 requests. Russian IP, consistent but low volume.
The 220.196.160.x range appears multiple times (30, 40, 165, 218) — these are likely the same network or organization with multiple egress IPs, appearing across 9–12 days each. Chinese ISP range, consistent with a reader behind a rotating IP pool.
Observations
-
The canary paths are earning their keep.
/.envand/.git/configreturning 200 have absorbed 182 scanner hits from 98 distinct IPs across the observation window, costing a combined 23,898 bytes. That’s less than 24 KB to keep scanners busy with decoys instead of probing real content. The .env canary is particularly effective — 57 unique IPs have hit it, and it’s still being discovered by new scanners as of the report day. -
Agent artifacts are not just for LLMs anymore. The
llms.txtfile and post.mdversions are being consumed by SEO crawlers (SemrushBot, MJ12bot, Barkrowler), search engines (Baiduspider, Googlebot, Bingbot), and even scanners (TLM-Audit-Scanner, Scrapy) — not just AI trainers. Thellms.txtconvention has leaked into the broader crawler ecosystem. Amazonbot is the most thorough agent-artifact consumer (35 artifacts), but ClaudeBot has the highest artifact-to-request ratio (32/32 = 100%). The four niche artifacts (llm.txt, llms.html, ai.json, identity.json) have exactly 2 hits each from the same 2 IPs — they were discovered once and never revisited. They may not be worth maintaining if the goal is agent adoption. -
Fediverse is the site’s biggest bandwidth tax. 65,189 requests generating 1.84 GB — 65.7% of all bandwidth — for link previews that most readers will never see. This is the cost of being shareable on Mastodon. The site serves the full HTML of
/to every instance that fetches a preview, and there are thousands of instances. A lighter-weight OpenGraph endpoint or a cached preview response could cut this dramatically, but for a personal blog the bandwidth is still well within Hetzner’s generous limits. It’s more of an observation than a problem.
Generated by observatory.sh — Caddy logs → DuckDB → Censys → Ollama → Astro