Error page that traps vulnerability scanners, e-mail harvesters and botnets
# push (client certificate):
git remote set-url --push origin https://coredump.ws:8443/repos/cunty-error.git
Branches and tags
Files at HEAD (4a1686af94) · download zip
| build_fpmap.py | 1.7 KB |
| build_markov.py | 2.2 KB |
| dictionary | 69.5 KB |
| error.php | 257.3 KB |
| honeypot-selftest.sh | 7.0 KB |
| README.md | 7.0 KB |
Recent commits · all
| 4a1686af | Update honeypot: front-controller, fingerprint matching, dropper capture, webshell/CVE coverage | no-body-in-particular | 18 h ago |
| 592f38c2 | cunty-error: the coredump.ws error page that traps scanners and bots | no-body-in-particular | 19 h ago |
README.md
cunty-error
A deceptive error/404 handler for the Hiawatha web server. Instead of a plain error page, it serves convincing but entirely fabricated content to automated vulnerability scanners and exploit bots: realistic Markov-generated prose, fake SQL/PHP/framework errors, fake product logins, fake config/secret files, and more — wasting the bots' time and generating abuse-report intelligence, while never exposing anything real.
Everything it returns is fake. It touches no database and no real file
beyond its own dictionary and markov.db. Every value reflected from a
request is HTML-escaped, so it cannot become an XSS vector against a human.
What it traps
- SQL injection — error-based across MySQL/MariaDB/PostgreSQL/MSSQL/Oracle/ SQLite/DB2, framework exceptions (Laravel/Django/Rails/Spring/.NET) and Flask/Node/Tomcat/ColdFusion tracebacks; boolean-blind simulation (a TRUE condition matches the baseline byte-for-byte, FALSE returns a short "no records" page, so scanners "confirm" the fake injection); time-based blind (it sleeps, capped).
- LFI / path traversal / RCE / SSTI / XXE / NoSQL / deserialization.
- Secret/config disclosure —
.env,.gittree,.aws/.kube,auth.json,filezilla.xml,appsettings.json,application.properties,terraform.tfstate, log files, install wizards… many seeded with a honeytoken. - ~70 product login/admin banners — WordPress, Joomla, Drupal, Jenkins, GitLab, Grafana, Kibana, Tomcat, phpMyAdmin, FreePBX, Exchange/OWA, Fortinet, Palo Alto, Citrix, CouchDB, Druid, AEM, QNAP, Synology, … .
- Exposed Docker / Kubernetes APIs, cloud metadata (GCP/Azure/AWS) SSRF.
- Named CVE probes (Log4Shell with callback extraction, ProxyShell, CVE-2024-3400, CVE-2023-22527, CVE-2023-4966, ThinkPHP, Drupalgeddon, …).
- Behavioral tarpits — slow byte-trickle, infinite pagination + API cursor maze, a 20,000-URL sitemap-index, fake
robots.txt, fake directory listings, HTTP 429 jitter, scanner canaries, a fake Cloudflare WAF block page, fake fingerprint headers, a crackable JWT bait + forged-JWT detection, a login "success" maze, and webshell decoys. - Intelligence — every trapped hit is logged with the tripped trap types, the classified scanner tool, and any extracted attacker callback/OAST host.
- Scanner-fingerprint matching - each page embeds, per-URL, the exact word-matcher strings the common scanner templates look for (hidden in a comment), so content-matching detections flag the honeypot.
- Malware dropper capture - download-and-execute payloads (Mirai-style, incl. router RCEs in the POST body and the
Hostheader) have their URL extracted and logged/queued (never fetched in-request). - Webshell-hunt traps - probes for pre-existing backdoors (known shell names, obfuscation extensions, numeric/short
.php) get a fake shell UI.
Requirements
- Hiawatha with CGI enabled (
php-cgi) - PHP 8.x with the SQLite3 extension
python3(only to buildmarkov.dboffline)
Setup
- Place the page and its data in your web root (e.g.
/var/www/hiawatha):
error.php— the handler (this repo).dictionary— a newline-separated word list used for random titles/links. Any list works, e.g.cp /usr/share/dict/words dictionaryor extract one from your corpus.markov.db— the SQLite Markov model (built below).
- Build
markov.dbfrom public-domain text (Project Gutenberg works well):
mkdir corpus && cd corpus
curl -L https://www.gutenberg.org/files/1342/1342-0.txt -o pride.txt
curl -L https://www.gutenberg.org/files/2701/2701-0.txt -o moby.txt
cat *.txt > ../corpus_all.txt && cd ..
python3 build_markov.py . # writes markov.db next to corpus_all.txt
cp markov.db /var/www/hiawatha/
- Wire it into Hiawatha. Minimal (returns HTTP 500 — realistic for a SQL/PHP error, and enough to keep sqlmap engaged):
# hiawatha.conf
ErrorHandler = 401:/error.php
ErrorHandler = 403:/error.php
ErrorHandler = 404:/error.php
ErrorHandler = 405:/error.php
ErrorHandler = 500:/error.php
ErrorHandler = 501:/error.php
ErrorHandler = 503:/error.php
Recommended (front-controller): route every unknown path through the
handler so it can return realistic status codes (200 for a fake page,
500 only for a SQL error, 401/403/429 where appropriate) instead of the
pinned error code. Add to your UseToolkit chain, after your deny rules:
UrlToolkit {
ToolkitID = secure
# ... your existing deny rules (dotfiles, sensitive extensions) ...
RequestURI exists Return
# keep your real apps/proxies serving:
Match ^/(app1|app2|assets)(/|$) Return
Match ^/?$ Return
Match .* Rewrite /error.php
}
With the front-controller, error.php returns 200 by default and 500 only on
an error-based SQL page. Validate any config change with hiawatha -k before
restarting.
- Optional static bait (served as real 200 files; their URLs all 404 back into the honeypot):
robots.txtwith juicyDisallow:entries + aSitemap:line.sitemap.xmlas a sitemap-index pointing atsitemap-1..N.xml, each full of thousands of fabricated URLs.
- Logging. The page appends one tab-separated line per trapped request to
/var/log/honeypot.log(create it writable by the CGI user). Since it is a dedicated file, point your log rotation at it.honeypot-abuse-report.pyturns the log into per-source-IP abuse-report drafts (via RDAP; it never sends anything on its own).
Change these before you rely on it
error.php contains two public-by-definition values you must change for your
own deployment:
HONEYTOKEN— the canary embedded in fake secrets. If an attacker who scraped a fake.env/backup later uses it, the honeypot logsHONEYTOKEN-USED. Set a unique value.- the JWT bait secret (
'secret') — a deliberately weak HMAC key; a cracked, forged token logsJWT-FORGED-USED. Keep it weak but change it so it is not literally this repo's value.
Testing
honeypot-selftest.sh exercises every trap type and the boolean-blind SQLi
invariants against a running instance and prints PASS/FAIL. Run it after any edit:
sh honeypot-selftest.sh http://127.0.0.1
Files
| file | purpose |
|---|---|
error.php | the handler (place in web root) |
build_markov.py | builds markov.db from a text corpus |
honeypot-selftest.sh | regression test for every trap + invariants |
honeypot-abuse-report.py | turns honeypot.log into abuse-report drafts |
Caveats
- Routing every unknown path through a CGI changes how your server behaves; test with
hiawatha -kand verify your real apps still serve before relying on it. - Publishing this source reveals the mechanisms; a determined attacker can fingerprint it. That is the normal trade-off for an open-source honeypot.