Saturday 2026-09-12 was the busiest day my honeypot has had: 19,995 requests, against a prior week of 124 to 1,114 a day. The weekly summary email reported 13,967 "Galah AI responses", which reads like an expensive day for an LLM-backed honeypot. It was not. That metric counts cache hits and ceiling denials; the live model made 384 calls all day, for about 0.50 USD.
19,363 of those requests came from one cohort: 15 source addresses, the same seven user-agent strings, the same 231-path wordlist of credential and configuration files, the same fixed probe path. It had visited on five earlier days since 2026-08-23, never above 460 requests. Whether one operator or several customers of one scanning service sit behind it, I cannot tell, so I call it "the scanner".
Two of its three routes into my site were Cloudflare's own address space: the shared egress that Cloudflare's WARP client and Gateway use, and the single IPv6 address Cloudflare stamps on every Worker that fetches one Cloudflare zone from another. Behind that one address, a header Cloudflare sets on Worker subrequests named 870 distinct workers.dev accounts in a day. Neither lane can be blocked without blocking legitimate WARP users or every Worker on the platform.
The third route gave the game away. Five hosting-provider machines at Hetzner, DigitalOcean and Contabo sent 3,713 requests, every one carrying a Referer of the form https://<script>.<account>.workers.dev/proxy?modify&proxyUrl=https://rehagen.net/<path>. That is the example link printed by the landing page of Cloudflare's create-cloudflare starter template, which ships an unauthenticated proxy route. My apex domain answers every path with a 301. Add two documented behaviours, one in the Workers runtime and one in Go's HTTP client, and the redirect walks out of the proxy and lands on my origin from the scanner's real client, carrying the Worker URL it was just bounced from. More than 3,600 distinct Worker hostnames showed up that way across that day and the next.
I reconstructed that mechanism from documentation and source; I did not reproduce it. For anyone behind Cloudflare the practical points are short: log the cf-worker header at your origin, tag Worker subrequests at the edge with cf.worker.upstream_zone instead of blocking them, and read Cloudflare's country field as where traffic exited Cloudflare, not where an operator sits.
The campaign is not mine alone. GreyNoise documented forged AI-crawler identities aimed at .env, .aws/credentials and .git/config in late August, and Known Agents lists 35 targeted credential paths under an AI-bot spoofing campaign it flags as active (checked 2026-09-13 and 2026-09-15). Two things did not appear in that reporting: the starter-template proxy pool, and requests for wrangler.toml and .dev.vars, which hold a Workers developer's own account IDs and local secrets.
One cohort, three ways in
Cohort membership needs two predicates together: a source address in one of the lanes in Table 1, and one of exactly seven user-agent strings. Both Cloudflare lanes are shared egress, so neither predicate is enough alone; on 09-12 the address predicate admitted 23 requests that the user-agent predicate excluded.
| Lane | Addresses | Requests | Header evidence |
|---|---|---|---|
| Cloudflare client egress (WARP or Gateway) | 9 in 104.28.0.0/16 | 9,382 | No cf-worker; Referer is the apex URL of the same path on all 9,382 |
| Cloudflare Workers cross-zone subrequests | 2a06:98c0:3600::103 | 6,268 | cf-worker on all 6,268, naming 870 *.workers.dev accounts and 7 custom zones |
| Hosting providers | 5: two Hetzner, two DigitalOcean, one Contabo | 3,713 | Referer is a *.workers.dev proxy URL naming the same path on all 3,713 |
104.28.0.0/16 is Cloudflare space absent from Cloudflare's published proxy ranges, and the WARP and Gateway egress docs describe shared client egress without naming the range. Darktrace and GreyNoise Labs have both documented scanning from it. WARP is likely, not certain.
The same strings and header profiles appear on all three lanes; the lanes differ in transport, not behaviour. Table 2 lists the strings with 09-12 cohort counts; they sum to 19,363.
| Requests | User agent |
|---|---|
| 2,862 | Mozilla/5.0 (Linux; Android 14; Pixel 8) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/135.0.6422.113 Mobile Safari/537.36 |
| 2,792 | Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:137.0) Gecko/20100101 Firefox/137.0 |
| 2,776 | Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) |
| 2,770 | Mozilla/5.0 (Macintosh; Intel Mac OS X 14_5) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.4 Safari/605.1.15 |
| 2,769 | Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/135.0.0.0 Safari/537.36 |
| 2,739 | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/135.0.0.0 Safari/537.36 |
| 2,655 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ChatGPT-User/1.0; +https://openai.com/bot) |
Googlebot and ChatGPT-User are impersonation: OpenAI's published IP list for ChatGPT-User has no prefix in 104.28.0.0/16 and no IPv6 prefix, and the same addresses sent the other six strings within the same seconds.
The busiest address sent 6,294 requests over 23 hours in 57 bursts, median 115 requests over 60 paths in 57 seconds, starts a median 16 minutes apart. The probe path /aZ9xQ7kL3m opened 52 of the 57, and across all 456 cohort bursts that day it sat in the first three requests of 208. A scheduler re-scanning every quarter hour fits; it is a hypothesis, not an observation. The cohort sent only GET requests and never asked for /robots.txt.
The 231 paths are a secrets list: 25 .env variants, 17 WordPress backup and log paths, 14 platform configuration files including wrangler.toml, .dev.vars, fly.toml and vercel.json, and the usual Git metadata, cloud credential, database dump and MCP configuration files. wrangler.toml and .dev.vars were not in SecLists or nuclei-templates when I searched on 2026-09-13; Cloudflare's docs say .dev.vars holds local development secrets and wrangler.toml holds account IDs and bindings. A scanner reaching targets through Workers is asking them for their Workers credentials.
The Workers fleet behind one address
Cloudflare's HTTP header reference states: "In cross-zone subrequests from one Cloudflare zone to another Cloudflare zone, the CF-Connecting-IP value will be set to the Worker client IP address '2a06:98c0:3600::103' for security reasons." The same page says CF-Worker "is set to the name of the zone which owns the Worker making the subrequest". Because the honeypot stores headers, the fleet is countable: 877 distinct cf-worker values on 09-12, 870 of them *.workers.dev account labels, a median of 7 requests per account that day. 849 of the labels were first seen that day, 603 of them between 01:00Z and 04:00Z. A 2025 Community thread describes the same pattern and notes the address cannot be blocked by IP at the WAF.
The proxy pool and the redirect that leaked through it
- All 3,958 requests from the five hosting-provider addresses across 09-12 and 09-13 carry a Referer. 3,946 have the form
https://<script>.<account>.workers.dev/proxy?modify&proxyUrl=https%3A%2F%2Frehagen.net%2F<path>, and in 3,945 the encoded path equals the path requested from me; the other 12 use a different proxy script's/?url=form. - Those Referers name 3,629 distinct hostnames, 3,627 account labels and 3,599 script names, none of them among the 871
cf-workeraccounts seen across the two days. /proxy?modify&proxyUrl=is the example link printed by the landing page of Cloudflare's create-cloudflare "common" template. Itsproxy.jshas no allowlist and no authentication, and passes the incoming request tofetch()as the init argument:
const proxyUrl = url.searchParams.get('proxyUrl');
const modify = url.searchParams.has('modify');
if (!proxyUrl) {
return new Response('Bad request: Missing `proxyUrl` query param', { status: 400 });
}
// make subrequests with the global `fetch()` function
let res = await fetch(proxyUrl, request);
- Cloudflare's Request documentation states: "The default for a new
Requestobject isfollow. Note, however, that the incomingRequestproperty of aFetchEventwill have redirect modemanual." Withmanual, "the3xxredirect response will be returned to the caller as-is." https://rehagen.net/<any path>answers 301 withLocation: https://jon.rehagen.net/<path>(curl, 2026-09-13).- Go's
net/httpclient follows 301 by default, and the comment onrefererForURLinclient.goreads: "refererForURL returns a referer without any authentication info or an empty string if lastReq scheme is https and newReq scheme is http. If the referer was explicitly set, then it will continue to be used." - On the 104.28 lane the Referer is
https://rehagen.net/<same path>on all 9,382 requests.
The mechanism that fits all seven: the scanner's client requested a pool Worker's /proxy URL with proxyUrl set to the apex path. The template fetched the apex with redirect mode manual, inherited from the incoming request, got my 301 and returned it unchanged. The client followed Location to jon.rehagen.net directly, outside the proxy, with Referer set to the Worker URL it had just left. On the 104.28 lane the same client fetched the apex directly and followed the same 301, which is observation 7.
Where the evidence stops:
- The pattern matches steps 3 to 6 on 3,945 of 3,958 requests. The alternative, 3,600 fabricated decoy Referers in the template's own link format, has no supporting evidence and no purpose I can see.
- Browsers and some HTTP wrappers also follow redirects and set Referer from the previous hop. Go is a likely client, not a demonstrated one.
- The five hosts ran the client or are its immediate egress; a VPN exit or forward proxy would look the same, and Shodan lists a Squid service on one of them. Confidence is medium-high.
- Whether the 3,629 hostnames are the operator's accounts, a rented proxy service, or third parties' default deployments, I cannot determine. I am not reproducing any label; some may belong to uninvolved people.
What it cost and what the LLM layer bought
384 live model calls on a 20,000-request day, about 0.50 USD, with a week-long response cache and a per-path ceiling absorbing the rest. 1,126 cohort requests received a response carrying a planted artifact; none reused an identifier in the window, and none followed a link from a generated response. Reuse elsewhere would be invisible to me. It is one observation, not a comparison arm, and it suggests wordlist scanners do not read what they get.
Where this sits in the public record
Searched 2026-09-13; each negative is bounded by what I searched. /aZ9xQ7kL3m has no GitHub hits beyond my own tracker and is absent from SecLists and nuclei-templates; because ffuf accepts a caller-supplied calibration string, a constant probe does not prove a private tool, but it is a stable correlation indicator. No public scanner ships the seven-string set. GreyNoise's 2026-08-28 report describes 824 IPs with forged AI-crawler identities, none matching OpenAI's lists and none requesting /robots.txt; hrbrmstr documents a similar kit with a different rotation. I found no report of the template's /proxy route used as a proxy pool.
What an operator behind Cloudflare can do
Blocking is the wrong tool for the two Cloudflare lanes. The useful moves are about visibility:
- Log
cf-workerat the origin; Cloudflare's header reference shows the nginxlog_formatline. One unblockable address becomes a list of zones you can count and report. - Match on
cf.worker.upstream_zoneat the edge, not on the header, which WAF rules run before. Transform Rules have matched on it since June 2025, so a header stampedworkerwhen it is non-empty andcloudflare-clientwhen the source ASN is 13335 gives your logs the lane split. Tag and log; do not block AS13335; validate on owned traffic first. - Read
cf-ipcountryas the exit country, not the operator's. - Keep an apex redirect; it hands you the address that follows it, when the client follows redirects and egresses directly.
What I am changing
- Promote
cf_workerand a parsed Referer host to first-class columns. - Turn the three cohort properties (probe path, exact user-agent set, the
wrangler.tomland.dev.varspair) into a classification after a false-positive check on other weeks. - Report Cloudflare egress as "Cloudflare-fronted" instead of as countries, and add the edge tag above, tag-only.
- Split "AI responses" into live calls, cache hits and ceiling denials, so the next summary email does not overstate a day by a factor of 36.
Methods and limits
The cohort rests on exact user-agent matching within known egress ranges; it cannot separate one operator from several customers of one tool, and it would miss the same operator with a different string set. Every Cloudflare-lane figure depends on cf-connecting-ip, cf-worker and cf-ipcountry as Cloudflare sets them, application headers rather than TCP peer addresses. The redirect mechanism is reconstructed from the Cloudflare Request API docs, the workers-sdk template at commit 164e4fb (dated 2026-09-12; proxy.js unchanged since 2023 and still present at main on 2026-09-15) and Go's client.go at master on 2026-09-13, not reproduced. Nothing here contacted the scanner's infrastructure.
On 2026-09-15 I reported the 871 subrequest accounts and the 3,627 proxy-pool labels to Cloudflare Trust and Safety as an abuse report, and raised the template's unauthenticated /proxy route in workers-sdk issue #15660. Whether a demo template should ship an open proxy is Cloudflare's call. This dataset shows it in use at scale.