# LLMs Use Policy for Animal Hospital (animalhospital.it) # Place this file at https://www.animalhospital.it/llms.txt # Purpose: Declare how Large Language Models (LLMs) and AI crawlers may use this website’s content. site: https://www.animalhospital.it owner: Studio Veterinario Associato A. & G. Marcenò (Animal Hospital) contact: info@animalhospital.it last-updated: 2025-09-04 policy-version: 1.0 ## GLOBAL POLICY # Human-readable directives for AI systems and dataset operators. Allow-Index: yes # LLMs may crawl and index pages for the purpose of search/snippets/summarization with attribution. Allow-Training: no # Content may NOT be used to train, fine-tune, or improve LLMs or AI models without prior written consent. Allow-FineTune: no Allow-Embedding: yes, with attribution and excerpt limits (see below). Allow-Redistribution: no # Do not redistribute, mirror, or compile our content into datasets or archives. Attribution-Required: yes Attribution-Format: clear hyperlink to the canonical page URL with brand name “Animal Hospital” Excerpt-Limit: 200-words # If quoting content, keep excerpts short (≤ 200 words) and include a link to the source page. Respect-Robots-Txt: yes Respect-Noindex: yes Respect-Copyright: yes Jurisdiction: IT / EU (GDPR) Personal-Data-Handling: Do not store, process, or republish personal data beyond what is necessary to render the page. No profile building. Data-Deletion-Requests: Send requests to info@animalhospital.it Polite-Crawl: yes Rate-Limit: ≤ 1 request/second; max 5 requests/10 seconds; avoid bursts Crawl-Window: 02:00–06:00 Europe/Rome preferred Sitemaps: - https://www.animalhospital.it/sitemap.xml Prohibited-Paths: - /search - /go/ - /*?* (avoid parameterized duplicates where possible) ## INTENDED USE - Permitted: Non-commercial search, discovery, and short-form summarization with visible attribution and source links. - Prohibited: Any model training, fine-tuning, dataset building, embeddings for product features, or derivative works used commercially without explicit written permission. ## CONTACT FOR PERMISSIONS For written licensing or research access, email: info@animalhospital.it (Subject: “LLM Content License Request”) ## AI CRAWLER SPECIFIC OVERRIDES (Guidance for robots.txt) # The following should be mirrored in robots.txt for enforcement, as most AI crawlers respect robots.txt. # Include in robots.txt if you wish to enforce. Shown here as reference. # OpenAI User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: / # Google User-agent: Google-Extended Disallow: / User-agent: GoogleOther Allow: / User-agent: GoogleOther-Image Allow: / # Anthropic User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / # Perplexity User-agent: PerplexityBot Disallow: / # Common Crawl User-agent: CCBot Disallow: / # Apple User-agent: Applebot-Extended Disallow: / # Meta/Facebook external ML User-agent: Meta-ExternalMachine Disallow: / # Amazon User-agent: Amazonbot Disallow: / # ByteDance User-agent: Bytespider Disallow: / # Cohere (if any) User-agent: cohere-ai Disallow: / # Others of concern User-agent: AI2Bot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: Diffbot Disallow: / # Housekeeping for all bots User-agent: * Disallow: /search Crawl-delay: 1 ## NOTES - This llms.txt is a human- and machine-readable declaration of intent for LLMs. Actual enforcement depends on crawler compliance. - For robust control, keep robots.txt aligned with the “AI CRAWLER SPECIFIC OVERRIDES” above. - Our content is protected by copyright. All rights reserved unless a separate written license is granted.