A CloudFront Function (JS 2.0 runtime) designed to protect a static website hosted on AWS S3 (no PHP). It filters incoming requests before they reach the origin or cache.
Requests with no User-Agent header, an empty value, or whitespace-only value return 404. This check runs first, before URI decoding, and cannot be bypassed.
Requests matching automated-scan patterns return 404:
- URI extensions:
.php*,.sql,.bak,.phtml,.config,.ya?ml,.toml,.conf,.key,.pem,.axd,.boto,.s3cfg,.npmrc,.htpasswd,.tfstate - Common scanner folders:
/admin,/wp-admin,/phpmyadmin,/backup,/wp-content,/wp-json,/api(all/api/*paths are security scans against this static site), etc. - Any dotfile/dot-directory path (
/.env,/.git,/.docker/,/.netrc,/.yarnrc,/.aws/credentials,/.ssh/id_rsa,/.well-known/..., etc. — this site serves no content under a dot-prefixed path, no exceptions), known credential-scan filenames (/secrets.json,/config.json,/service-account.json, etc.), and/ip
- A truncated Windows UA that stops right after
AppleWebKit/537.36instead of continuing with the real Chrome/Safari tail - Any UA containing
chrome/withoutapplewebkitimmediately before it — every real Chromium browser emitsAppleWebKit/537.36 (KHTML, like Gecko)right before theChrome/token, so its absence marks a hand-built UA - A full build/patch
Chrome/version (e.g.130.0.6723.70) on Chrome 113+ — post-UA-reduction Chrome only ever reportsmajor.0.0.0, so a real build/patch number there is a stale, pre-freeze template (self-identifying crawlers usingcompatible;, e.g. Bingbot, are exempted) - A
Chrome/major version below 149 — logs.db shows the site's real audience only from 149 up; the "asset-loading" traffic on 145–148 is a single rotating-UA cloud fleet (Tencent/Huawei/GCP/AWS ranges). Exempted: self-identifying crawlers withcompatible;(Bingbot, Googlebot…), Samsung Internet (ships a lagging Chromium), and Feeder (feeder.co, an RSS service with a hardcodedChrome/106) - An
Edg/(desktop Edge) major version below 150 — same stale-UA-fleet pattern as Chrome, checked independently since a scraper can fake either token. Same exemptions as the Chrome floor.
Requests with a Firefox/ major version below 139 return 404. Exempted: major 115, Mozilla's actively-maintained legacy ESR train (Windows 7/8.1/macOS 10.12-10.14 support, extended through March 2027).
Requests from known-malicious IP ranges return 404 on every path, regardless of User-Agent — same robots.txt/sitemap.xml/feed.xml exceptions as bad actors and blocked bots below. Currently blocks Techoff SRV Limited's ranges: 45.148.10.0/24, 93.123.109.0/24, 195.178.110.0/24.
Requests matching 60+ known bot/scraper user-agent patterns return 404 on every path — except /robots.txt (a real 200 disallow-all body), /sitemap.xml (a real 200 empty <urlset> body), and /feed.xml (a real 200 empty Atom <feed> body), instead of a 404, so a blocked scraper checking any of these gets a correct answer. The same exception applies to any other bad actor (security-scan URI or spoofed/stale-browser UA) landing on those paths.
Blocked patterns include: scrapers (Scrapy, DataForSEO, Bytespider, etc.), old browser tokens (Trident, Presto), generic HTTP clients (python-requests, aiohttp, got), and more, matched case-insensitively against the User-Agent header.
A request for a directory-style path with no trailing slash (e.g. /about) gets a real 301 to the same path with / appended (e.g. /about/), instead of the origin's 302. This runs after all bot/security filtering above, so a bad actor never reaches it. It's skipped for:
- Paths that already end in
/(including/) - Asset-looking paths — anything whose final path segment contains a
.(.jpg,.css,.js,.pdf, etc.) — dotfile paths like/.well-known/...never reach this check, since they're already blocked by dotfile-path filtering above
All other requests are forwarded to the origin unchanged.
function.js keeps only a one-line pointer comment for anything longer than a
sentence; the full reasoning (evidence, edge cases, why a pattern is shaped
the way it is) lives here instead, to keep the deployed file under
CloudFront's 10 KB function-size limit. Each heading below matches the
identifier the code points at.
Only ~3% of URIs contain a %-escape (per logs.db); the rest skip the
decodeURIComponent call entirely. Decoding loops (capped at 3 rounds) rather
than running once, so a double-encoded probe like /admin%252F.env — which a
single decode only turns into /admin%2F.env, leaving the / hidden from the
dotfile/prefix checks — still ends up fully decoded before matching.
/robots.txt, /sitemap.xml and /feed.xml get a real disallow-all / empty
sitemap / empty feed instead of a 404 for bad actors and blocked bots alike —
a correct, on-brand "you're not welcome here" rather than a generic miss.
isBadActor runs path traversal, then dotfile paths, then security scans,
then truncated/malformed Chrome UAs, then outdated Chrome/Edge/Firefox UAs,
ordered most- to least-frequent per logs.db so common cases short-circuit
before rarer, costlier checks run.
Any path containing /. (a dotfile or dot-directory segment) is blocked
unconditionally, no exceptions — this includes /.well-known/...: the site
has no ACME HTTP-01 challenge (certs are provisioned another way) and serves
nothing else under a dot-prefixed path, so there's nothing there worth
excepting. A single generic rule covers /.env, /.git, /.docker/,
/.netrc, /.yarnrc, /.aws/credentials, /.ssh/, and any other
credential/config dotfile a scanner might probe for, without needing a
per-filename entry.
Combined into a single precompiled regex: one pass over the URI covers
extensions, folder prefixes, and known credential-scan filenames (dotfiles
are handled separately by dotfile-path). The trailing
.json group is not a blanket .json$ rule — /about/data/*.json and
/pagefind/*.json are real, legitimately-served site data — so only known
credential-scan filenames (secrets.json, config.json,
service-account.json, etc.) are matched there.
actuator is in the folder-prefix group — Spring Boot's Actuator endpoints
(/actuator/configprops, /actuator/env, etc.) are only ever probed by
scanners against this static site, never served legitimately.
api is in the folder-prefix group — every /api/* path against this static
site is a security scan (the origin is an S3 static website with no API
backend). Probes like /api/v2/config, /api/v1/users, etc. are automated
scanners looking for API endpoints that don't exist here.
read-document is in the folder-prefix group — a new scan pattern first seen
2026-09-17 (38 requests to /read-document in one day), presumably probing
for a document-reading/SSRF-style endpoint that doesn't exist on this static
site.
swp (vim swap file) and a bare trailing ~ (generic editor backup, e.g.
wp-config.php~) are both classic backup-file scan suffixes appended after
a real filename/extension, so they're matched separately from the \.ext$
group rather than folded into it.
@fs, @vite and @id are Vite dev-server internal endpoints (@fs in
particular serves arbitrary files off the host filesystem, e.g.
/@fs/home/ec2-user/.aws/credentials) — this is a static site with no Vite
dev server behind it, so any request for these paths is a scanner, full stop.
Any path starting with __vite (e.g. /__vite_rsc_findSourceMapURL/,
/__vite_ping) is matched as a bare prefix rather than a folder segment,
since Vite's RSC-internal endpoints don't share the @-prefixed naming of
@fs/@vite/@id above — same rationale: no Vite behind this site.
A literal .. anywhere in the (already-decoded) URI path is blocked
unconditionally, rather than only matching specific traversal targets
(.env, .aws/credentials, etc.). A static site never legitimately needs
.. in a path segment, and this also catches multi-encoded evasion attempts
(e.g. ..%25252f..%25252f...) — the literal .. survives even when the
attacker double/triple-encodes the surrounding slashes to dodge a single
decodeURIComponent pass.
A real browser always continues past AppleWebKit/537.36 with
(KHTML, like Gecko) Chrome/... Safari/.... A string that stops dead right
after AppleWebKit/537.36 is a bot with a copy-pasted, incomplete UA, not a
real Chrome/Edge. The shared literal fragments (UA_OPEN, CLOSE_APPLEWEBKIT,
WINDOWS_PLATFORM) are factored out for readability and composed into a
RegExp once at parse time, not rebuilt per request.
An exact-template match on OS/engine string plus a Chrome major-version range
used to be treated as "spoofed" here, but real Chrome (which freezes its UA
to major.0.0.0) produces this exact template too — logs.db showed the two
most common UAs in real traffic matching it. Structure alone can't tell real
Chrome from spoofed Chrome (see CLAUDE.md: never block Chrome solely on the
.0.0.0 minor/patch version) — see min-chrome-major for
a separate, evidence-backed floor on the version number itself.
Every real Chromium browser emits AppleWebKit/537.36 (KHTML, like Gecko)
immediately before the Chrome/ token, so a UA with chrome/ but no
applewebkit is a hand-built/incomplete UA, not a browser — catches
malformed strings the exact-template regexes above don't cover (e.g.
Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/120.0.0.0).
Floor set from logs.db (June–Sept 2026, Chrome 152 current) by counting
distinct IPs that load the site's real assets (main.css,
bootstrap.bundle.min.js) per Chrome major, then checking who they are.
Majors 145–148 look organic at first glance (~300 asset-loading IPs) but
are one bot fleet: the same Tencent/Huawei-cloud /16s (116.204, 1.92,
113.44, 81.70, 43.138…) appear on all four majors, rotating UAs, plus
GCP/AWS monitoring ranges on 148. The site's actual audience (FR/BE/CH/CA,
~70 % of real sessions) is essentially absent below 149 in the last 30 days
and shows up at 149 exactly (7 IPs), then 41 at 150, 95 at 151. Below 145
the picture is the same as it always was: no asset loads, or a single
narrow IP/country cluster.
Cost/benefit at 149: blocks ~92 % of non-crawler "Chrome ≥ 99" requests
(~30 k/month — mostly the Chrome/120 and Chrome/148 fleets) for zero
observed real-audience sessions in the last month. Real collateral found
over three months and handled via
chrome-floor-exemptions: Samsung Internet and
Feeder. Not worth exempting: Opera on 147/148 (5 FR IPs, but Opera tracks
Chromium within 1–2 majors so those were fresh at the time), one Electron
app on Chrome/124.
This floor is ~3 majors behind current and needs raising periodically: the fleet will eventually move its UAs up, and every Chrome release (monthly, or faster) widens the gap. Re-run the per-major asset-loading-IP query before each bump.
Desktop Edge (Edg/) is Chromium underneath, so it gets the same
stale-fleet treatment as min-chrome-major, checked
independently rather than folded into the Chrome floor — a scraper can hold
one token fixed while bumping the other. Per the user's own logs.db
analysis, real Edge sessions only appear from major 150 up; below that is
the same rotating cloud-fleet pattern as sub-149 Chrome. Reuses
chrome-floor-exemptions rather than a separate
list — NewsBlur's hardcoded fetcher UA (see below) carries both a stale
Chrome/147 and a stale Edg/147, so the same exemption has to cover both
tokens or it breaks.
UAs skipped by min-chrome-major and min-edge-major (checked as one regex, same reason as blocked-bot-regex):
compatible;— self-identifying crawlers: every Bingbot variant inlogs.db(7 UA shapes, ~6.9 k requests) carries it and reportsChrome/116; Googlebot, YouBot, meta-webindexer, Google-InspectionTool likewise. Crawlers that are unwanted are blocked by name in blocked-bot-regex regardless.samsungbrowser/— Samsung Internet ships a Chromium several majors behind Chrome (SamsungBrowser/30→Chrome/143,/28→130,/27→125); a handful of real FR/BE/NL users over three months.feeder.co;— Feeder, an RSS service, polls/feed.xmlwith a hardcodedChrome/106desktop UA (~1.2 k requests from 10 IPs, active daily). Blocking it would silently drop its subscribers.newsblur.com— NewsBlur's Feed Fetcher and Page Fetcher self-identify withNewsBlur ... - https://www.newsblur.com/site/... (Mozilla/5.0 ... Chrome/147.0.0.0 ... Edg/147.0.0.0); the domain in the prefix is the identifying token, the parenthetical is a hardcoded browser-like UA that will drift stale over time same as Feeder's.chrome-lighthouse— Google Lighthouse (and PageSpeed Insights, which runs it) appendsChrome-Lighthouseto whatever UA it's invoked with, which is often an old hardcoded Chrome major. The token is public and documented, not a secret, so this isn't security-by-obscurity — it's safe to exempt because unwanted crawlers are still blocked by name in blocked-bot-regex regardless of what Chrome version they claim, and a scraper wanting past the floor could just claim a current Chrome major instead of bothering to fake this token.
Firefox auto-updates, so a stale major version is a scraper with a hardcoded UA, not a real user. 139 shipped in June 2025 and is below every Firefox release still in general support (current ESR is 140+); the one still-maintained release below it is carved out separately, see firefox-esr-115-exemption.
Firefox ESR 115 is normally end-of-life, but Mozilla has repeatedly extended its security updates for Windows 7/8.1 and old macOS versions (currently through March 2027) — a real, still-patched browser used by a legacy-OS population, not a scraper. Its major is exempted by exact match rather than folded into the min-version floor, so it doesn't drag the floor down for everything else. Tor Browser also reports an ESR major (115 or higher), so privacy users on the current Tor ESR base remain unaffected either way.
The same regex also exempts googleimageproxy (ggpht.com's GoogleImageProxy,
e.g. Mozilla/5.0 (Windows NT 5.1; rv:11.0) Gecko Firefox/11.0 (via ggpht.com GoogleImageProxy)) — Google's proxy that fetches remote images for Gmail/
Google services. It hardcodes an ancient Firefox/IE-era UA rather than a
real browser's, but it's a legitimate fetcher identified by its own token, not
a scraper impersonating a browser, so it's exempted the same way as the
Chrome floor's named exemptions in
chrome-floor-exemptions.
Floor set at the user's request rather than derived from logs.db (unlike
min-chrome-major/min-edge-major,
which are log-backed). Safari majors track the OS/App Store, not an
auto-update cadence, so a hardcoded old major is a reasonable scraper signal
the same way a stale Firefox major is (see
min-firefox-major) — Safari 18 shipped September 2024,
so anything below it is over a year stale. Note Apple's version-numbering
jump: Safari went from the 18.x line straight to 26 (aligned to the iOS/
macOS release year), so "below 18" and "18 or above" is the entire floor —
there's no 19–25 range to worry about.
Matched on version\/(\d+)\., not safari\/: the Safari/ token in a
Safari UA is a fixed WebKit build number (605.1.15 desktop, 604.1
mobile) that doesn't track the browser release, so it can't be used as a
version signal — see safari-floor-exemptions for
why other WebKit-based browsers need excluding from this check.
UAs skipped by min-safari-major:
compatible;— same self-identifying-crawler reasoning as chrome-floor-exemptions.crios\/,fxios\/,opios\/— Chrome/Firefox/Opera for iOS are Apple-mandated to use WebKit (hence theSafari/604.1tail) but don't carry aVersion/token in their real UA, so these never actually match; listed defensively in case a future variant adds one.edgios\/— Edge for iOS does carry bothVersion/(WebKit's, stale and unrelated to Edge's own release) andEdgiOS/(its real version), so without this exemption a current Edge-for-iOS user gets floored by aVersion/number that was never meant to signal anything.duckduckgo,ucbrowser\/— other WebKit-based iOS/Android browsers with the same stale-Version/-token shape as Edge for iOS.
Plain substrings matched against the (already lowercased) User-Agent header,
as ONE regex literal. Written out literally rather than built at runtime
from an array: a literal is compiled when the script is parsed, whereas
new RegExp(list.map(escape).join('|')) re-does the escape calls, a map, a
join and a pattern compile on every script evaluation — pure compute we were
paying for. Alternatives are ordered most- to least-frequent per logs.db so
common bots exit early (non-matching UAs still try every alternative).
To add a bot: append |your-token (escaping ., (, ) and / as
\., \(, \), \/) and add a UA sample to the blockedAgents fixture in
function.test.js.
seojuice-searchbot/ (SEOJuice, seojuice.io) is blocked because it does
not obey robots.txt.
coccocbot (Coc Coc, coccoc.com, a Vietnamese search engine crawler) is
blocked at the user's request.
hubspot crawler (HubSpot's own crawler, hubspot.com) is blocked at the
user's request.
domain-harvester/ (github.com/esc-city/domain-harvester) is blocked at
the user's request.
mapthenetbot/ (mapthenet.org) is blocked at the user's request.
expansel-monitor/ (expansel.com) is blocked at the user's request.
fogbot/ is blocked at the user's request.
newsletterformresearchbot/ is blocked at the user's request.
srchs-research-bot/ is blocked at the user's request.
aionbot/ is blocked at the user's request.
tiktokspider (no /-separated version in its UA token) is blocked at the
user's request.
opentheboxbot/ is blocked at the user's request.
veryhip/ (veryhip.com) is blocked at the user's request.
cms-security-auditor/ is blocked at the user's request, despite
self-identifying as an "authorized self-check" — the origin has no
allowlist for it, so it's treated like any other unsolicited scanner.
censysinspect/ (Censys internet-scanning bot, about.censys.io) is
blocked at the user's request.
petalbot (Huawei's search crawler, webmaster.petalsearch.com) is no
longer blocked: it honors the robots.txt disallow list, so it's allowed
through like any other well-behaved crawler.
publicwwwbot/ (PublicWWW, publicwww.com, a source-code search engine
crawler) is blocked at the user's request.
wp2shell (a WordPress vulnerability scanner/exploit tool) is blocked at
the user's request.
webatlabot (https://webatla.com/bot) is blocked at the user's request.
ssi-nutch/ (SSI's broad web crawler, https://ssi.inc/, run by
adi@ssi.inc) is blocked at the user's request. Matched on the ssi-nutch/
name token only, version number dropped, per the standard pattern in
blocked-bot-regex.
variableratio-publicassetresearch/ is blocked at the user's request.
Matched on the variableratio-publicassetresearch/ name token only, version
number (1.0) dropped, per the standard pattern in
blocked-bot-regex.
baiduspider (Baidu's search crawler, including the Baiduspider-render
variant, baidu.com/search/spider.html) is blocked at the user's request.
Matched on the baiduspider name token only (no trailing /, since the
-render variant has no slash before its version number), per the standard
pattern in blocked-bot-regex.
HaloBot/1.0 is blocked at the user's request. Matched on the halobot/
name token only, version number (1.0) dropped, per the standard pattern in
blocked-bot-regex.
Flowb0t-ContentEngine/1.0 is blocked at the user's request. Matched on the
flowb0t-contentengine/ name token only, version number (1.0) dropped, per
the standard pattern in blocked-bot-regex.
ClarityBot/0.1 (https://clarity.surf/bot) is an AI crawler that does not
consult robots.txt, and its bot info page is missing — blocked at the
user's request. Matched on the claritybot/ name token only, version number
(0.1) dropped, per the standard pattern in
blocked-bot-regex.
undici (the Node.js HTTP client) is blocked at the user's request — it does
not fetch robots.txt and the traffic observed operates from DigitalOcean,
LLC · Toronto, CA, a hosting provider rather than a residential/ISP network.
Matched on the undici name token only, per the standard pattern in
blocked-bot-regex.
Linkwarden (Server-Side Fetch) (the self-hosted link manager,
linkwarden.app, fetching link previews/archives on behalf of its users) must
not be blocked — it's legitimate self-hosted server-side traffic, not a
scraper.
Synapse (bot; https://github.com/matrix-org/synapse) must not be
blocked — Synapse is an open-source Matrix homeserver fetching URL previews
on behalf of its users, not a scraper. It isn't currently matched by
blocked-bot-regex (no generic bot substring is
matched, only specific named tokens), noted here so it stays excluded if the
regex is ever extended.
event.viewer.ip (the client IP CloudFront Functions exposes on every viewer
request/response) is matched against a blocklist of known-malicious network
ranges via blockedIpRangeRegex — a plain string-prefix regex, same style as
blocked-bot-regex, not integer/bitmask CIDR math. This
only works because every current range is /24 (octet-aligned): the regex
alternatives are literal first.second.third. prefixes, so a match on
45.148.10. covers exactly 45.148.10.0–45.148.10.255. Folded into the
same condition as bad-actor-response-mapping
so a blocked IP gets the same /robots.txt//sitemap.xml//feed.xml
treatment as any other bad actor, regardless of what User-Agent it sends.
Currently blocked: 45.148.10.0/24, 93.123.109.0/24, 195.178.110.0/24
(Techoff SRV Limited), blocked at the user's request.
If a future range isn't octet-aligned (e.g. a /25 or /22), the
string-prefix trick stops working and the check needs real integer/bitmask
CIDR matching instead — don't force a non-aligned range into this regex.
To add a range: append |a\.b\.c\. (escaping the dots) to
blockedIpRangeRegex for a /24, or |a\.b\. for a /16, and add an IP
sample to the blockedIps fixture in function.test.js.
The S3 origin returns a 302 for a directory-style request with no trailing
slash; a search engine or client following that redirect chain sees a
temporary redirect where a permanent one is correct, and duplicate-content
crawlers may index both the slash and non-slash URL separately. This check
runs last, after every bot/security check, so a bad actor's request is
blocked (404) before it can trigger a redirect, and the check itself is pure
string inspection of uri — no need to touch uriLower or ua.
"Asset-looking" is approximated as the final path segment contains a dot —
cheaper than a file-extension allowlist and correct for every real static
asset (.jpg, .css, .js, .woff2, .pdf, ...), since browsers never
percent-encode a literal . (it's an RFC 3986 unreserved character). A dot
earlier in the path (e.g. /v1.2/about) doesn't suppress the redirect —
only a dot in the last segment does.
No query string is ever appended — this site never links to a directory-style page with one, so there's nothing to preserve.
CloudFront Functions run at every edge location with sub-millisecond startup and are ~6× cheaper than Lambda@Edge. They are the right tool for stateless, CPU-light request manipulation that requires no network I/O, no large runtimes, and no response body streaming. This filter fits that profile exactly: pure string matching, no external calls.
The trade-off is a restricted runtime — no setTimeout, no fetch, no Node.js built-ins. The function is written deliberately to stay within those constraints.
Copy the body of function.js into the CloudFront Functions editor in the AWS Console (or deploy via AWS CLI / CDK / Terraform). Associate the function with the viewer request event of your distribution.
Important: remove the
export { handler }line before deploying — CloudFront's JS 2.0 runtime does not support ES moduleexportsyntax. That line exists solely so Vitest can import the function during testing.
A Claude Code PreToolUse hook automatically validates function.js against the live cloudfront-js-2.0 runtime before every git push. It uploads the local code to the DEVELOPMENT stage and runs aws cloudfront test-function, blocking the push if any syntax or runtime error is detected.
1. Set your function name
echo "Block_Intrusions" > .cloudfront-function-name
2. Configure the test event
Edit test-event.json to match a representative viewer request for your distribution. The default covers a standard GET with a User-Agent header.
3. AWS credentials
Ensure your shell has credentials with at least these permissions:
{
"Effect": "Allow",
"Action": [
"cloudfront:DescribeFunction",
"cloudfront:UpdateFunction",
"cloudfront:TestFunction"
],
"Resource": "*"
}On each git push Claude Code will:
- Fetch the current ETag via
aws cloudfront describe-function - Upload local
function.jsto the DEVELOPMENT stage viaaws cloudfront update-function - Run
aws cloudfront test-function --stage DEVELOPMENT - Block the push and display the error if the runtime rejects the function
The hook script is at .claude/hooks/cloudfront-pre-push.sh.
Tests are written with Vitest.
Vitest was chosen over Jest for this project because:
- Native ESM support — no Babel transform needed. The function uses
export { handler }which works out of the box with"type": "module"inpackage.json. - Zero config — no
jest.config.js, no transform pipeline to maintain. - Fast — Vitest starts in milliseconds; the full suite runs in under 250 ms.
- Jest-compatible API —
describe,it,expect,it.eachare identical, so the syntax is familiar.
npm test # run once
npm run test:watch # watch mode (re-runs on file save)function.test.js covers all behaviours with 279 tests:
| Suite | What is tested |
|---|---|
| PHP / bad folder / security scan blocking | File extensions, scanner folders, sensitive/credential paths, /ip |
| Scrapper bot blocking by user-agent | 60+ bot/scraper patterns, matched case-insensitively |
| IP range blocking | Known-malicious /24 ranges blocked regardless of UA; boundary IPs just outside a range pass through |
| robots.txt disallow-all for blocked bots | Blocked bots and bad actors get a 200 disallow-all body on /robots.txt; normal browsers pass through untouched |
| sitemap.xml empty urlset for blocked bots | Blocked bots and bad actors get a 200 empty <urlset> body on /sitemap.xml; normal browsers pass through untouched |
| feed.xml empty atom feed for blocked bots | Blocked bots and bad actors get a 200 empty <feed> body on /feed.xml; normal browsers pass through untouched |
| Null / empty user-agent blocking | Missing/empty/whitespace user-agent |
| Percent-encoded URI handling | URI decoding before pattern matching |
| ads.txt and llms.txt | Follow normal UA blocking rules (no special bypass) |
| Trailing-slash redirect | 301 for directory-style paths; assets, /, and blocked bad actors (dotfile paths included) are unaffected |
| Pass-through | Normal requests forwarded unchanged |
Each test builds a minimal CloudFront event object ({ request: { uri, headers } }) and asserts on the return value — either the original request object (pass-through) or a synthetic response with statusCode, headers, and body.