# eBilanz Fabrik — robots.txt (PUBLIC-LAUNCH version) # STATUS: this file IS what ships. deploy-porkbun.sh defaults to LIVE=1 (since the # 2026-07-15 incident); the holding "Disallow: /" variant now requires an explicit # --hold flag. (The old comment here claimed the inverse — that private beta # published "Disallow: /" by default — which stopped being true on 2026-07-15 and # would mislead a reader into thinking the live site is de-indexed. Corrected # 2026-07-17.) AI / LLM crawlers are explicitly welcome (GEO/AEO visibility + # agent discovery). Agent brief: /llms.txt User-agent: * # DW-1337 — CONTENT-SIGNAL: the posture this file already states in prose, made machine-readable. # Measured 2026-09-24: ebilanzplus.de carries three Content-Signal lines; we carried zero, while # the header of this very file says "AI / LLM crawlers are explicitly welcome (GEO/AEO visibility + # agent discovery)". A permission an operator states only in a comment is a permission no crawler # can read — the same presence-vs-function gap as a price that lives only in prose. # # Why all three are YES, and none of them is a new decision: # search=yes we want to be in AI search results. That is the entire point of /llms.txt, # the JSON-LD, and the citation probe. # ai-input=yes being used as INPUT to a generated answer IS the citation we are chasing. # Saying no here would contradict the whole programme. # ai-train=yes our pages are public guides about German law that we publish to be quoted. # We already LEAD prose mentions (6, the highest of any vendor, measured # 2026-09-24) — models knowing us is the asset, not the leak. Nothing # proprietary is on these pages; the engine, the ERiC path and customer data # are all behind the funnel and none of it is crawlable. # # This grants; it does not restrict. Per the spec, a signal we omit neither grants nor denies — # so stating them is strictly more informative than silence. Content-Signal: ai-train=yes, search=yes, ai-input=yes Allow: / # DW-677 — the Werkbank is a token-gated BETA surface built with one design-partner # customer, deliberately off the live funnel. Not a product page: never linked, never # in sitemap.xml or the agent briefs. This line is the belt to that braces — a crawler # that guesses the path, or learns it from a referrer, is told no rather than left to # infer it from the absence of a link. # /start.html and /start-en.html are NOT disallowed here — deliberately, and this is a correction # (2026-08-11). Both pages carry , which is the RIGHT # tool for "do not index this form": Google must be able to CRAWL a page to READ its noindex. A # Disallow prevents exactly that, so the pair was self-defeating — the URL could still be indexed # bare (from external links) while the noindex was never seen, and `follow` could never pass equity. # It also contradicted our own agent briefs: /llms.txt and /llms-full.txt name /start.html three # times each, i.e. we advertise the page to AI agents while forbidding them to fetch it. Measured: # /start.html + /start-en.html were the entry page for 4 of 36 answer-engine arrivals, and # /start.html is Perplexity's joint-top landing page. check-werkbank-unlisted.py now asserts that # anything named in an agent brief is fetchable by the AI crawlers. Disallow: /api/ Disallow: /werkbank/ # --- AI / LLM crawlers (explicitly welcome) --- # ONE group, thirteen agent lines, and the SAME Disallow set as `*` — deliberately, and this is a # correction (M204, 2026-08-11). Previously each agent had its own `User-agent: X / Allow: /` pair. # Per RFC 9309 §2.2.1 a crawler obeys exactly one group — the most specific User-agent match — and # the `*` group applies ONLY when no named group matches. So every named AI group SILENTLY REPLACED # the `*` group, and the three Disallow lines above bound Googlebot and bingbot and nobody else: # all thirteen were free to fetch /start.html, /start-en.html and the token-gated /werkbank/. # The DW-677 line called itself "the belt to the braces"; measured with a real robots parser, the # belt was not on the crawler class it most needed to bind. Consecutive User-agent lines with no # rule between them form a single group, so the welcome is unchanged and the exclusions now apply. # Probe: robots_probe.py — every agent × path, ALLOW/block, before and after. User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-User User-agent: anthropic-ai User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Google-Extended User-agent: Applebot User-agent: Applebot-Extended User-agent: Amazonbot User-agent: meta-externalagent User-agent: CCBot User-agent: Bytespider Allow: / Disallow: /api/ Disallow: /werkbank/ Sitemap: https://ebilanzfabrik.de/sitemap.xml # Agent / LLM brief (short): https://ebilanzfabrik.de/llms.txt # Agent / LLM brief (full — every citable fact, the boundaries, the Q&A): https://ebilanzfabrik.de/llms-full.txt