# As a condition of accessing this website, you agree to abide by the following # content signals: # (a) If a Content-Signal = yes, you may collect content for the corresponding # use. # (b) If a Content-Signal = no, you may not collect content for the # corresponding use. # (c) If the website operator does not include a Content-Signal for a # corresponding use, the website operator neither grants nor restricts # permission via Content-Signal with respect to the corresponding use. # The content signals and their meanings are: # search: building a search index and providing search results (e.g., returning # hyperlinks and short excerpts from your website's contents). Search does not # include providing AI-generated search summaries. # ai-input: inputting content into one or more AI models (e.g., retrieval # augmented generation, grounding, or other real-time taking of content for # generative AI search answers). # ai-train: training or fine-tuning AI models. # use: how AI systems may consume the content (immediate, reference, or full). # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF # RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT # AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET. # BEGIN Cloudflare Managed content User-agent: * Content-Signal: search=yes,ai-train=no,use=reference Allow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CloudflareBrowserRenderingCrawler Disallow: / User-agent: Google-Extended Disallow: / User-agent: GPTBot Disallow: / User-agent: meta-externalagent Disallow: / # END Cloudflare Managed Content # robots.txt for kundasang.com # Homepage: https://kudasnan.com # Sitemap: https://kudasnan.com/sitemap.xml # Llms.txt: https://kudasnan.com/llms.txt # ============================================================ # UNIVERSAL POLICY — all bots, spiders, AI agents welcome # The wildcard below explicitly allows every crawler not # enumerated below. This covers current AND future AI agents. # ============================================================ User-agent: * Allow: / Disallow: /data/ Disallow: /counter.php Disallow: /nginx_kundasang.com.conf # Sitemap location Sitemap: https://kundasang.com/sitemap.xml # ============================================================ # Search Engine Crawlers # ============================================================ User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: /assets/ User-agent: Googlebot-News Allow: / User-agent: Googlebot-Video Allow: / User-agent: AdsBot-Google Allow: / User-agent: Mediapartners-Google Allow: / User-agent: Bingbot Allow: / User-agent: MSNBot-Media Allow: / User-agent: Slurp Allow: / User-agent: DuckDuckBot Allow: / User-agent: Baiduspider Allow: / User-agent: YandexBot Allow: / User-agent: YandexImages Allow: / User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / User-agent: facebookexternalhit Allow: / User-agent: Facebot Allow: / User-agent: Twitterbot Allow: / User-agent: LinkedInBot Allow: / User-agent: Pinterestbot Allow: / User-agent: Redditbot Allow: / User-agent: TelegramBot Allow: / User-agent: WhatsApp Allow: / User-agent: Discordbot Allow: / User-agent: Slackbot Allow: / User-agent: Slack-ImgProxy Allow: / User-agent: Embedly Allow: / User-agent: Quora Link Preview Allow: / User-agent: Showyoubot Allow: / User-agent: Outbrain Allow: / User-agent: Pinterest Allow: / User-agent: Bitlybot Allow: / User-agent: Nimbostratus-Bot Allow: / User-agent: W3C_Validator Allow: / # ============================================================ # AI / LLM Crawlers — explicitly enumerated # Below is the broadest public list of known AI agents. # Anything NOT listed falls through to the wildcard policy # at the top (also Allow: /). # ============================================================ # --- OpenAI --- User-agent: GPTBot Allow: / User-agent: GPTBot-User Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OpenAI-Image-Scraper Allow: / # --- Anthropic --- User-agent: ClaudeBot Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: ClaudeWeb Allow: / User-agent: anthropic-ai Allow: / User-agent: Claude-Indexer Allow: / # --- Perplexity --- User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Perplexity-Bot Allow: / User-agent: PerplexityAI Allow: / # --- Google AI --- User-agent: Google-Extended Allow: / User-agent: GoogleOther Allow: / User-agent: Google-CloudVertexBot Allow: / # --- Apple Intelligence --- User-agent: Applebot-Extended Allow: / # --- Microsoft Copilot / Bing AI --- User-agent: MSNBot-Media Allow: / User-agent: BingPreview Allow: / # --- Cohere --- User-agent: cohere-ai Allow: / User-agent: cohere-training-data-crawler Allow: / User-agent: cohere-search Allow: / # --- Meta --- User-agent: Meta-ExternalAgent Allow: / User-agent: Meta-ExternalFetcher Allow: / User-agent: Meta-WebIndexer Allow: / # --- Mistral --- User-agent: MistralAI-User Allow: / User-agent: MistralAI Allow: / User-agent: LeChatBot Allow: / # --- DeepSeek --- User-agent: DeepSeekBot Allow: / User-agent: DeepSeekCrawler Allow: / # --- DeepL --- User-agent: DeepLBot Allow: / # --- You.com --- User-agent: YouBot Allow: / User-agent: YouBot-Search Allow: / # --- Kagi --- User-agent: KagiBot Allow: / User-agent: Kagi-Fetcher Allow: / # --- Brave Search AI --- User-agent: Bravebot Allow: / # --- DuckDuckGo AI Assist --- User-agent: DuckAssistBot Allow: / User-agent: DuckDuckGo-AI Allow: / # --- Amazon (Alexa / Rufus) --- User-agent: Amazonbot Allow: / User-agent: RufusBot Allow: / # --- Common Crawl (used for many AI training pipelines) --- User-agent: CCBot Allow: / User-agent: CCBot-User Allow: / # --- IBM watsonx --- User-agent: IBMCommonCrawler Allow: / # --- Allen AI Institute --- User-agent: AI2Bot Allow: / User-agent: AI2Bot-Dolma Allow: / User-agent: aiHitBot Allow: / # --- DataForSEO --- User-agent: DataForSeoBot Allow: / # --- Diffbot --- User-agent: Diffbot Allow: / # --- Turnitin --- User-agent: TurnitinBot Allow: / # --- img2dataset --- User-agent: img2dataset Allow: / # --- Skroutz --- User-agent: SkroutzBot Allow: / # --- Semrush --- User-agent: SemrushBot Allow: / User-agent: SemrushBot-BA Allow: / User-agent: SemrushBot-SI Allow: / User-agent: SemrushBot-SOA Allow: / User-agent: SemrushBot-CT Allow: / User-agent: SemrushBot-OCOB Allow: / User-agent: SemrushBot-FT Allow: / User-agent: SemrushBot-Mirror Allow: / # --- Imagesift --- User-agent: ImagesiftBot Allow: / # --- iAsk --- User-agent: iasksystem Allow: / User-agent: iask Allow: / # --- Other AI-adjacent fetchers --- User-agent: magpie-crawler Allow: / User-agent: Webzio Allow: / User-agent: Webzio-Extended Allow: / User-agent: Neticle Allow: / User-agent: NetCraftSurveyAgent Allow: / User-agent: panscient Allow: / User-agent: memorybot Allow: / User-agent: mega-indexer Allow: / User-agent: marin-software Allow: / User-agent: web-fetcher Allow: / User-agent: scrible Allow: / User-agent: ponder Allow: / User-agent: cotoyogi Allow: / User-agent: timpassbot Allow: / User-agent: SeekportBot Allow: / User-agent: AspiegelBot Allow: / User-agent: Buck Allow: / User-agent: bnf.fr_bot Allow: / User-agent: femtosearch Allow: / User-agent: Neuroscape Allow: / User-agent: PetalBot Allow: / User-agent: Seekr Allow: / User-agent: YisouSpider Allow: / User-agent: Sogou Pic Spider Allow: / User-agent: 360Spider Allow: / User-agent: Soso Allow: / User-agent: Linguee Bot Allow: / User-agent: Embeditor Allow: / User-agent: TDKBot Allow: / User-agent: Crawlson Allow: / User-agent: Acoon Allow: / User-agent: ABACHOBot Allow: / # ============================================================ # AI / LLM discoverability file (Answer Engine Optimization) # ============================================================ # llms.txt spec — https://llmstxt.org # Tells LLMs: "this site allows you to read and summarize it" # ============================================================ # Host-specific directives # ============================================================ Host: https://kundasang.com