# As a condition of accessing this website, you agree to abide by the following # content signals: # (a) If a Content-Signal = yes, you may collect content for the corresponding # use. # (b) If a Content-Signal = no, you may not collect content for the # corresponding use. # (c) If the website operator does not include a Content-Signal for a # corresponding use, the website operator neither grants nor restricts # permission via Content-Signal with respect to the corresponding use. # The content signals and their meanings are: # search: building a search index and providing search results (e.g., returning # hyperlinks and short excerpts from your website's contents). Search does not # include providing AI-generated search summaries. # ai-input: inputting content into one or more AI models (e.g., retrieval # augmented generation, grounding, or other real-time taking of content for # generative AI search answers). # ai-train: training or fine-tuning AI models. # use: how AI systems may consume the content (immediate, reference, or full). # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF # RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT # AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET. # --------------------------------------------------------------------------- # AirNav Radar crawler policy # # Informational pages -- /about, /faq, /blog, the ADS-B and MLAT explainers, the # product, hardware and API pages -- are open to every crawler, AI crawlers # included. We would rather models know what AirNav Radar is than not. # # /data/ is not. It is millions of generated flight, aircraft, airport and # airline pages: the bulk of the crawl cost and the asset we sell. AI crawlers # are disallowed there. Search engines and user-directed AI fetches are not. # # ai-train is deliberately absent from every Content-Signal below. Per clause # (c) above that neither grants nor restricts training rights, so nothing here # waives the Article 4 reservation. Access, not signalling, carries the policy: # content signals cannot be scoped to a path, ordinary robots.txt rules can. # # Machine-readable site summary for LLMs: https://www.airnavradar.com/llms.txt # --------------------------------------------------------------------------- # Everything not named below. Note that a named group REPLACES this one for that # agent rather than adding to it, so the two replay endpoints have to be repeated # in every group that is allowed to crawl the site. Disallow lines are listed # before "Allow: /" on purpose: RFC 9309 requires longest-match, but parsers # that match in file order need the narrower rule first to reach it. User-agent: * Content-Signal: search=yes,ai-input=yes,use=reference Disallow: /data/getReplay Disallow: /data/get-replay Allow: / # AI crawlers that collect for training. Free of the site except /data/. # Google-Extended and Applebot-Extended are control tokens, not crawlers: # Googlebot and Applebot still crawl /data/ for Search under the group above. # Caveat on Google-Extended -- Google uses the one token for both Gemini # training and Gemini grounding, so this also keeps /data/ out of grounded # Gemini answers. Move it to its own group with "Allow: /" if being cited on # flight pages is worth more than keeping them out of training. User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: meta-externalagent User-agent: Amazonbot User-agent: Bytespider Content-Signal: search=yes,ai-input=yes,use=reference Disallow: /data/ Allow: / # Answer-time agents: search indexes for AI assistants, and fetches a user has # explicitly asked for. These are the ones that cite us and send traffic, and # they get /data/ too. Named explicitly so a future tightening of the wildcard # group above cannot catch them by accident. User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User Content-Signal: search=yes,ai-input=yes,use=reference Disallow: /data/getReplay Disallow: /data/get-replay Allow: / # Cloudflare's Browser Rendering crawler, which third parties can point at any # site through Cloudflare's API. Blocked outright, as it was under Cloudflare's # managed robots.txt. User-agent: CloudflareBrowserRenderingCrawler Disallow: / Sitemap: https://www.airnavradar.com/sitemap.xml