[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f8bWrkChDltq2HfXlNXF1WJuLJi6m06TH9olO7T-BSBI":3,"$fndpbZerdi19iFdU7tL0ECYXpBXEwrR0suD2fO7zFuiE":125,"$fSbWH-92TGv7YL8DBle8qgXxchU8q6ws0fqRA88cYSgA":130},{"slug":4,"name":5,"version":6,"author":7,"author_profile":8,"description":9,"short_description":10,"active_installs":11,"downloaded":12,"rating":13,"num_ratings":13,"last_updated":14,"tested_up_to":15,"requires_at_least":16,"requires_php":17,"tags":18,"homepage":24,"download_link":25,"security_score":26,"vuln_count":13,"unpatched_count":13,"last_vuln_date":27,"fetched_at":28,"discovery_status":29,"vulnerabilities":30,"developer":31,"crawl_stats":27,"alternatives":37,"analysis":27,"fingerprints":27},"parseless","ParseLess","0.6.0","phalkmin","https:\u002F\u002Fprofiles.wordpress.org\u002Fphalkmin\u002F","\u003Cp>ParseLess serves your WordPress content as clean Markdown to AI crawlers (via User-Agent detection) and on manual \u003Ccode>?format=md\u003C\u002Fcode> requests. Same URL, same content, none of the theme chrome, navigation, widgets, or page-builder scaffolding that AI bots and CLI tools don’t need.\u003C\u002Fp>\n\u003Cp>It also exposes \u003Ccode>\u002Fllms.txt\u003C\u002Fcode> for the emerging AI-indexing standard, so models know where to find your content.\u003C\u002Fp>\n\u003Ch4>Who is this for?\u003C\u002Fh4>\n\u003Cp>\u003Cstrong>1. Developers using Claude Code, Cursor, Aider, or any CLI that feeds your own site content into an LLM.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>If you’ve ever piped a blog post into Claude or ChatGPT for analysis, rewriting, or summarization, you’ve watched it burn through tokens parsing nav menus, footer markup, and CSS classes that have nothing to do with your content. A typical WordPress page measured live: \u003Cstrong>~19,800 tokens of HTML for ~975 tokens of actual content\u003C\u002Fstrong> — a 20x reduction just by stripping the theme. Heavy page-builder sites (Elementor, Divi) routinely hit 100x or more.\u003C\u002Fp>\n\u003Cp>That’s the difference between fitting a handful of pages in a context window and fitting dozens.\u003C\u002Fp>\n\u003Cp>Just append \u003Ccode>?format=md\u003C\u002Fcode> to any post URL and pipe it straight into your tool of choice:\u003C\u002Fp>\n\u003Cpre>\u003Ccode>curl https:\u002F\u002Fyoursite.com\u002Fmy-post\u002F?format=md | claude \"summarize this\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cstrong>2. Site owners getting “high resource usage” warnings from their host because AI bots are hammering the site.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>GPTBot, ClaudeBot, PerplexityBot, Google-Extended, and a dozen others crawl WordPress sites constantly. Each request renders your full theme, runs widget queries, loads page-builder assets, and ships hundreds of KB of HTML per page — most of which the bot discards before extracting the actual text.\u003C\u002Fp>\n\u003Cp>ParseLess intercepts these crawls and serves a tiny Markdown payload instead. You keep the AEO\u002FSEO benefit of being indexed by AI search, without paying the server cost of rendering your full theme for every bot hit.\u003C\u002Fp>\n\u003Cp>Typical impact on AI bot traffic (measured on real WordPress pages):\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>95%+ less bandwidth\u003C\u002Fstrong> per crawled page (typically 15–30x smaller; e.g. a 79 KB page \u003Cspan aria-hidden=\"true\" class=\"wp-exclude-emoji\">→\u003C\u002Fspan> 4 KB. Heavy page-builder sites see 100x+)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>~60–80% fewer database queries\u003C\u002Fstrong> per request\u003C\u002Fli>\n\u003Cli>\u003Cstrong>~60% lower peak PHP memory\u003C\u002Fstrong> per request\u003C\u002Fli>\n\u003Cli>\u003Cstrong>~50–80% faster Time To First Byte\u003C\u002Fstrong> (measured 54% on a content-heavy page, 67% on a short one)\u003C\u002Fli>\n\u003Cli>On a site with 500 posts crawled monthly by major AI bots, that’s roughly \u003Cstrong>gigabytes \u003Cspan aria-hidden=\"true\" class=\"wp-exclude-emoji\">→\u003C\u002Fspan> tens of megabytes\u003C\u002Fstrong> of monthly bandwidth and a fraction of the cumulative PHP execution time\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Numbers vary by theme and content. Heavier setups (Elementor, Divi, Avada) see the biggest savings; lightweight themes see less dramatic but still meaningful gains. The conversion is cached as a transient, so repeated bot hits cost almost nothing.\u003C\u002Fp>\n\u003Ch4>How it works\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>AI crawlers are detected by User-Agent and served Markdown automatically — no configuration required.\u003C\u002Fli>\n\u003Cli>Humans, search engines, and unknown bots receive your normal HTML output. ParseLess never affects what real visitors see.\u003C\u002Fli>\n\u003Cli>The \u003Ccode>?format=md\u003C\u002Fcode> query parameter works on any post URL for manual preview or CLI piping.\u003C\u002Fli>\n\u003Cli>\u003Ccode>\u002Fllms.txt\u003C\u002Fcode> is published at your site root with a list of available content for AI indexers.\u003C\u002Fli>\n\u003Cli>Conversion happens once per post and is cached. Subsequent requests serve a single transient read.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Features\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>Automatic Markdown for known AI bots (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, CCBot, Google-Extended, Applebot-Extended, Bytespider, Meta-ExternalAgent, cohere-ai, and more)\u003C\u002Fli>\n\u003Cli>Manual preview and CLI access via \u003Ccode>?format=md\u003C\u002Fcode> on any post URL\u003C\u002Fli>\n\u003Cli>\u003Ccode>\u002Fllms.txt\u003C\u002Fcode> endpoint for AI-indexing standards\u003C\u002Fli>\n\u003Cli>Respects noindex flags from Yoast SEO, Rank Math, and Genesis\u003C\u002Fli>\n\u003Cli>Skips private, draft, password-protected, and trashed posts\u003C\u002Fli>\n\u003Cli>Works with all public post types (configurable)\u003C\u002Fli>\n\u003Cli>Optional YAML frontmatter (title, URL, author, date, categories, tags, excerpt)\u003C\u002Fli>\n\u003Cli>Transient-based caching with configurable TTL\u003C\u002Fli>\n\u003Cli>Settings page at \u003Cstrong>Tools \u003Cspan aria-hidden=\"true\" class=\"wp-exclude-emoji\">→\u003C\u002Fspan> ParseLess\u003C\u002Fstrong> for detection mode, bot list, cache TTL, post types, and llms.txt control\u003C\u002Fli>\n\u003Cli>Per-post meta box with Markdown preview and copy-to-clipboard\u003C\u002Fli>\n\u003Cli>Extensible via filters: \u003Ccode>md4ai_bot_list\u003C\u002Fcode>, \u003Ccode>md4ai_supported_post_types\u003C\u002Fcode>, \u003Ccode>md4ai_markdown_output\u003C\u002Fcode>, \u003Ccode>md4ai_cache_ttl\u003C\u002Fcode>, \u003Ccode>md4ai_should_serve_markdown\u003C\u002Fcode>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>Privacy and data collection\u003C\u002Fh3>\n\u003Cp>When logging is enabled (off by default), ParseLess records the following for\u003Cbr \u002F>\neach Markdown request:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Post ID, URL, and request timestamp\u003C\u002Fli>\n\u003Cli>The full User-Agent header\u003C\u002Fli>\n\u003Cli>A salted SHA-256 hash of the requester’s IP address (the raw IP is never stored)\u003C\u002Fli>\n\u003Cli>Matched bot identifier, bytes served, and whether the response came from cache\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>IP hashes cannot be reversed from an email address, so the WordPress privacy\u003Cbr \u002F>\nexporter\u002Feraser tools will report no personal data on demand.\u003C\u002Fp>\n\u003Cp>Logs are pruned daily according to the configured retention window\u003Cbr \u002F>\n(7\u002F30\u002F90\u002F365 days, default 30). Site owners can disable logging or click\u003Cbr \u002F>\n“Delete all logged requests” at any time from Tools \u003Cspan aria-hidden=\"true\" class=\"wp-exclude-emoji\">→\u003C\u002Fspan> ParseLess.\u003C\u002Fp>\n","Serve your posts as Markdown to AI crawlers and CLI tools — typically 95%+ less bandwidth, ~60–80% fewer DB queries, no theme bloat.",50,417,0,"2026-06-16T17:58:00.000Z","7.0.2","6.8","8.2",[19,20,21,22,23],"ai","bots","crawlers","llms","markdown","https:\u002F\u002Fgithub.com\u002Fphalkmin\u002Fparseless","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fparseless.0.6.0.zip",100,null,"2026-07-22T17:31:50.256Z","no_bundle",[],{"slug":7,"display_name":7,"profile_url":8,"plugin_count":32,"total_installs":33,"avg_security_score":26,"avg_patch_time_days":34,"trust_score":35,"computed_at":36},2,80,30,94,"2026-08-28T19:49:53.933Z",[38,63,77,91,108],{"slug":39,"name":40,"version":41,"author":42,"author_profile":43,"description":44,"short_description":45,"active_installs":46,"downloaded":47,"rating":48,"num_ratings":49,"last_updated":50,"tested_up_to":15,"requires_at_least":51,"requires_php":52,"tags":53,"homepage":59,"download_link":60,"security_score":61,"vuln_count":32,"unpatched_count":13,"last_vuln_date":62,"fetched_at":28},"better-robots-txt","Better Robots.txt – AI-Ready Crawl Control & Bot Governance","3.1.2","Pagup","https:\u002F\u002Fprofiles.wordpress.org\u002Fpagup\u002F","\u003Cp>Better Robots.txt replaces the default WordPress robots.txt workflow with a smarter, structured version you can configure and preview before publishing.\u003C\u002Fp>\n\u003Cp>Instead of a blank textarea, you get a guided wizard with presets, plain-language explanations, and a final Review & Save step so you can inspect the generated robots.txt before it goes live.\u003C\u002Fp>\n\u003Cp>Built for beginners and advanced users alike, Better Robots.txt helps you control how search engines, AI crawlers, SEO tools, archive bots, bad bots, social preview bots, and other automated agents interact with your site.\u003C\u002Fp>\n\u003Cp>Trusted by thousands of WordPress sites, Better Robots.txt is designed for the AI era without resorting to hype, vague promises, or hidden rules.\u003C\u002Fp>\n\u003Cp>Better Robots.txt is available in Free, Pro, and Premium editions. The free plugin covers the guided workflow and essential crawl control features, while Pro and Premium unlock additional governance, protection, and AI-ready modules. Some screenshots on the plugin page show features from all three editions.\u003C\u002Fp>\n\u003Ch3>A quick overview\u003C\u002Fh3>\n\u003Cp>\u003Ciframe loading=\"lazy\" title=\"Better robots.txt Video — AI-Ready Crawl Control for WordPress\" src=\"https:\u002F\u002Fplayer.vimeo.com\u002Fvideo\u002F1169756981?dnt=1&app_id=122963\" width=\"750\" height=\"372\" frameborder=\"0\" allow=\"autoplay; fullscreen; picture-in-picture; clipboard-write; encrypted-media; web-share\" referrerpolicy=\"strict-origin-when-cross-origin\">\u003C\u002Fiframe>\u003C\u002Fp>\n\u003Ch3>Why Better Robots.txt is different\u003C\u002Fh3>\n\u003Cp>Most robots.txt plugins fall into one of three categories:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Simple text editor\u003C\u002Fli>\n\u003Cli>Virtual robots.txt manager\u003C\u002Fli>\n\u003Cli>Single-purpose AI or policy add-on\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Better Robots.txt goes further.\u003C\u002Fp>\n\u003Cp>It gives you a complete, guided crawl control workflow so you can:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Choose a preset that matches your goals\u003C\u002Fli>\n\u003Cli>Control major crawler categories without writing everything by hand\u003C\u002Fli>\n\u003Cli>Keep core WordPress protection rules visible and editable\u003C\u002Fli>\n\u003Cli>Clean up low-value crawl paths that waste crawl budget\u003C\u002Fli>\n\u003Cli>Generate a cleaner robots.txt output\u003C\u002Fli>\n\u003Cli>Preview the final result before saving\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>What you can control\u003C\u002Fh3>\n\u003Cp>Better Robots.txt helps you manage:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Search engine visibility\u003C\u002Fli>\n\u003Cli>AI and LLM crawler behavior\u003C\u002Fli>\n\u003Cli>AI usage signals such as search, ai-input, and ai-train preferences\u003C\u002Fli>\n\u003Cli>SEO tool crawlers\u003C\u002Fli>\n\u003Cli>Bad bots and abusive crawlers\u003C\u002Fli>\n\u003Cli>Archive and Wayback access\u003C\u002Fli>\n\u003Cli>Feed crawlers and crawl traps\u003C\u002Fli>\n\u003Cli>WooCommerce crawl cleanup\u003C\u002Fli>\n\u003Cli>CSS, JavaScript, and image crawling rules\u003C\u002Fli>\n\u003Cli>Social media preview crawlers\u003C\u002Fli>\n\u003Cli>ads.txt and app-ads.txt allowance\u003C\u002Fli>\n\u003Cli>llms.txt generation\u003C\u002Fli>\n\u003Cli>Advanced directives such as crawl-delay and custom rules\u003C\u002Fli>\n\u003Cli>Final review before publishing\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>Editions\u003C\u002Fh3>\n\u003Cp>Better Robots.txt is available in three editions:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Free – Includes the guided setup, the Essential preset, core crawl control features, and the final Review & Save workflow.\u003C\u002Fli>\n\u003Cli>Pro – Adds more advanced governance and protection modules, including additional AI, crawler, and cleanup controls.\u003C\u002Fli>\n\u003Cli>Premium – Unlocks the most restrictive and advanced protection options, including the Fortress preset and additional high-control modules.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Some options shown in the interface are marked Free, Pro, or Premium so users can immediately understand which modules belong to each edition.\u003C\u002Fp>\n\u003Ch3>AI crawler controls by edition\u003C\u002Fh3>\n\u003Cp>The Free edition lets site owners choose a basic global policy for AI search and discovery crawlers: allow or block. AI training crawlers remain separated from AI search crawlers, so a site can block training use while still allowing AI-powered search and answer discovery.\u003C\u002Fp>\n\u003Cp>The Pro and Premium editions add advanced controls, including per-bot AI crawler rules, custom user agents, AI usage signals, llms.txt, and additional machine-readable governance files.\u003C\u002Fp>\n\u003Cp>Better Robots.txt separates AI-related agents into three practical families:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>AI training crawlers, such as GPTBot, Google-Extended, and ClaudeBot.\u003C\u002Fli>\n\u003Cli>AI search and discovery crawlers, such as OAI-SearchBot, PerplexityBot, and Claude-SearchBot.\u003C\u002Fli>\n\u003Cli>User-action fetchers, such as ChatGPT-User, Perplexity-User, and Claude-User.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>User-action fetchers are triggered by user requests. Depending on the provider, robots.txt may not apply to those fetches in the same way as it applies to automated crawlers. For stronger enforcement, combine robots.txt with server logs, WAF rules, provider verification, or IP-based controls where available.\u003C\u002Fp>\n\u003Ch3>Presets\u003C\u002Fh3>\n\u003Cp>Setup starts with four modes:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Essential – A clean, practical configuration for most websites that want a better robots.txt without complexity.\u003C\u002Fli>\n\u003Cli>AI-First – For publishers and content sites that want AI-ready governance without shutting down discovery.\u003C\u002Fli>\n\u003Cli>Fortress – For websites that want stronger protection against scraping, archive capture, and unnecessary crawl activity.\u003C\u002Fli>\n\u003Cli>Custom – For users who prefer to configure each module manually.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>For many sites, one preset plus a quick review is enough.\u003C\u002Fp>\n\u003Ch3>Built for beginners and experts\u003C\u002Fh3>\n\u003Cp>Beginners get:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>A guided setup instead of a raw robots.txt box\u003C\u002Fli>\n\u003Cli>Preset-based configuration\u003C\u002Fli>\n\u003Cli>Plain-language explanations for important choices\u003C\u002Fli>\n\u003Cli>A safer workflow with a final preview step\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Advanced users get:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Editable core WordPress protection rules\u003C\u002Fli>\n\u003Cli>Fine-grained crawler controls by category\u003C\u002Fli>\n\u003Cli>WooCommerce-oriented cleanup options\u003C\u002Fli>\n\u003Cli>Consolidated output options\u003C\u002Fli>\n\u003Cli>Advanced directives and custom rules\u003C\u002Fli>\n\u003Cli>A final output they can inspect before publishing\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>AI-ready, without hype\u003C\u002Fh3>\n\u003Cp>Better Robots.txt includes features for modern AI-related crawl governance, including:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>AI crawler handling\u003C\u002Fli>\n\u003Cli>Optional llms.txt support\u003C\u002Fli>\n\u003Cli>AI usage signals for compliant systems\u003C\u002Fli>\n\u003Cli>Optional machine-readable governance signals for advanced use cases\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>These features help you express how you want automated systems to use your content.\u003C\u002Fp>\n\u003Cp>However, Better Robots.txt does not claim to control AI by force. Like robots.txt itself, these signals are most useful with compliant systems and good-faith crawlers.\u003C\u002Fp>\n\u003Ch3>What Better Robots.txt is\u003C\u002Fh3>\n\u003Cp>Better Robots.txt is:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>A robots.txt governance plugin for WordPress\u003C\u002Fli>\n\u003Cli>A guided configuration workflow instead of a raw text editor\u003C\u002Fli>\n\u003Cli>A crawl control layer to reduce wasteful crawling\u003C\u002Fli>\n\u003Cli>A practical bridge between SEO, crawl hygiene, and AI-era policy signaling\u003C\u002Fli>\n\u003Cli>A way to keep your crawl policy clearer for humans and machines\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Technical reference for advanced users: Better Robots.txt also maintains a public \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FGautierDorval\u002Fbetter-robots-txt\" rel=\"nofollow noopener noreferrer ugc\">GitHub repository\u003C\u002Fa> with product definition, governance notes, and machine-readable artefacts.\u003C\u002Fp>\n\u003Ch3>What Better Robots.txt is not\u003C\u002Fh3>\n\u003Cp>Better Robots.txt is not:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>A firewall or Web Application Firewall (WAF)\u003C\u002Fli>\n\u003Cli>An anti-scraping enforcement engine\u003C\u002Fli>\n\u003Cli>A legal compliance engine\u003C\u002Fli>\n\u003Cli>A guarantee that every bot will obey your rules\u003C\u002Fli>\n\u003Cli>A replacement for server-level security or access control\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>It helps you publish a clearer crawl policy.\u003C\u002Fp>\n\u003Cp>It does not replace infrastructure-level protection.\u003C\u002Fp>\n\u003Ch3>Typical use cases\u003C\u002Fh3>\n\u003Cp>Use Better Robots.txt if you want to:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Clean up a weak or noisy default robots.txt\u003C\u002Fli>\n\u003Cli>Reduce crawl waste on WordPress or WooCommerce\u003C\u002Fli>\n\u003Cli>Keep major search engines allowed while restricting other bots\u003C\u002Fli>\n\u003Cli>Control whether archive bots can snapshot your site\u003C\u002Fli>\n\u003Cli>Publish AI usage preferences more clearly\u003C\u002Fli>\n\u003Cli>Keep social preview bots allowed while limiting scrapers\u003C\u002Fli>\n\u003Cli>Review the final file before making it live\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3>Key Features\u003C\u002Fh3>\n\u003Cul>\n\u003Cli>Guided step-by-step wizard\u003C\u002Fli>\n\u003Cli>Preset-based setup: Essential, AI-First, Fortress, Custom\u003C\u002Fli>\n\u003Cli>Search engine visibility controls\u003C\u002Fli>\n\u003Cli>AI and LLM crawler governance\u003C\u002Fli>\n\u003Cli>AI usage signals support\u003C\u002Fli>\n\u003Cli>SEO tool crawler controls\u003C\u002Fli>\n\u003Cli>Bad bot and abusive crawler options\u003C\u002Fli>\n\u003Cli>Archive and Wayback access controls\u003C\u002Fli>\n\u003Cli>Spam, feed, and crawl trap cleanup\u003C\u002Fli>\n\u003Cli>WooCommerce crawl cleanup options\u003C\u002Fli>\n\u003Cli>CSS, JavaScript, and image crawling rules\u003C\u002Fli>\n\u003Cli>Social media preview crawler controls\u003C\u002Fli>\n\u003Cli>ads.txt and app-ads.txt allowance\u003C\u002Fli>\n\u003Cli>Optional llms.txt generation\u003C\u002Fli>\n\u003Cli>Consolidated output option\u003C\u002Fli>\n\u003Cli>Core WordPress protection rules remain visible and editable\u003C\u002Fli>\n\u003Cli>Final Review & Save preview screen\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>About the publisher\u003C\u002Fh4>\n\u003Cp>Better Robots.txt is developed and maintained by \u003Ca href=\"https:\u002F\u002Fpagup.com\u002F\" rel=\"nofollow ugc\">Pagup\u003C\u002Fa>, a digital readability firm based in Quebec, Canada. Pagup helps organizations become correctly understood by search engines, generative AI systems, and autonomous agents.\u003C\u002Fp>\n\u003Cp>The robots.txt file is the first surface that AI crawlers read when they discover a site. A well-structured robots.txt that references governance files such as llms.txt, ai-manifest.json, and interpretation policies helps AI systems understand your site faster and more accurately.\u003C\u002Fp>\n\u003Cp>Better Robots.txt is one component of a broader digital readability practice that includes \u003Ca href=\"https:\u002F\u002Fpagup.com\u002Fen\u002Fservices\u002Fsemantic-content-architecture\u002F\" rel=\"nofollow ugc\">semantic content architecture\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fpagup.com\u002Fen\u002Fservices\u002Fai-governance-and-machine-readability\u002F\" rel=\"nofollow ugc\">AI governance and machine readability\u003C\u002Fa>, and \u003Ca href=\"https:\u002F\u002Fpagup.com\u002Fen\u002Fglossary\u002Finterpretive-seo\u002F\" rel=\"nofollow ugc\">interpretive SEO\u003C\u002Fa>.\u003C\u002Fp>\n\u003Ch4>Part of the Pagup ecosystem\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fpagup.com\u002F\" rel=\"nofollow ugc\">pagup.com\u003C\u002Fa> — Digital readability firm. Diagnostic, semantic architecture, AI governance.\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fgautierdorval.com\u002F\" rel=\"nofollow ugc\">gautierdorval.com\u003C\u002Fa> — Doctrine, canonical definitions, interpretive governance research.\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Finterpretive-governance.org\u002F\" rel=\"nofollow ugc\">interpretive-governance.org\u003C\u002Fa> — Formal versioned standard for interpretive governance.\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fbetter-robots.com\u002F\" rel=\"nofollow ugc\">better-robots.com\u003C\u002Fa> — Documentation and resources for Better Robots.txt.\u003C\u002Fli>\n\u003C\u002Ful>\n","Replace the default WordPress robots.txt workflow with a smarter, structured version you can preview before publishing, with Free, Pro, and Premium ed &hellip;",5000,321566,90,102,"2026-07-04T14:07:00.000Z","5.0","7.4",[54,55,56,57,58],"ai-crawlers","bot-blocker","llms-txt","robots-txt","seo","","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fbetter-robots-txt.3.1.2.zip",99,"2023-02-14 00:00:00",{"slug":64,"name":65,"version":66,"author":67,"author_profile":68,"description":69,"short_description":70,"active_installs":13,"downloaded":71,"rating":13,"num_ratings":13,"last_updated":72,"tested_up_to":15,"requires_at_least":51,"requires_php":52,"tags":73,"homepage":75,"download_link":76,"security_score":26,"vuln_count":13,"unpatched_count":13,"last_vuln_date":27,"fetched_at":28},"miniorange-ai-visibility","AI Crawlers Visibility, Analytics & Control","1.2.0","miniOrange","https:\u002F\u002Fprofiles.wordpress.org\u002Fcyberlord92\u002F","\u003Cp>\u003Cstrong>AI Crawlers Visibility, Analytics & Control\u003C\u002Fstrong> helps WordPress site owners understand and manage how AI systems access their content. Monitor visits from \u003Cstrong>GPTBot, ClaudeBot, Google-Extended, PerplexityBot\u003C\u002Fstrong>, and 73+ known AI crawlers, then choose which bots to monitor, discourage through robots.txt, or block with a PHP 403 response.\u003C\u002Fp>\n\u003Cp>The plugin combines \u003Cstrong>AI crawler analytics\u003C\u002Fstrong>, \u003Cstrong>AI bot control\u003C\u002Fstrong>, and \u003Cstrong>AI discoverability tools\u003C\u002Fstrong> in one WordPress dashboard. It includes an AI Visibility Score, \u003Ccode>llms.txt\u003C\u002Fcode> and \u003Ccode>llms-full.txt\u003C\u002Fcode>, Markdown for Agents, JSON-LD, downloadable reports, email alerts, and an authenticated REST API.\u003C\u002Fp>\n\u003Cp>It works alongside SEO and security plugins. It does not replace traditional SEO, a firewall, or server-level bot protection.\u003C\u002Fp>\n\u003Ch4>AI SEO, GEO, and AEO for WordPress\u003C\u002Fh4>\n\u003Cp>Traditional SEO helps search engines index a site. \u003Cstrong>AI SEO\u003C\u002Fstrong>, \u003Cstrong>Generative Engine Optimization (GEO)\u003C\u002Fstrong>, and \u003Cstrong>Answer Engine Optimization (AEO)\u003C\u002Fstrong> focus on making content easier for AI crawlers and answer engines to discover and understand.\u003C\u002Fp>\n\u003Cp>This plugin supports the technical foundations of AI visibility by helping you:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Measure which AI crawlers access your content and which pages they request.\u003C\u002Fli>\n\u003Cli>Check discovery, structured-data, and content-readiness signals.\u003C\u002Fli>\n\u003Cli>Publish AI-readable discovery files and Markdown versions of public content.\u003C\u002Fli>\n\u003Cli>Provide machine-readable site identity and article information with JSON-LD.\u003C\u002Fli>\n\u003Cli>Control crawler access without changing the content human visitors see.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>No plugin can guarantee inclusion, citations, rankings, or traffic from ChatGPT, Claude, Gemini, Perplexity, or other AI services.\u003C\u002Fp>\n\u003Ch4>Why monitor AI crawlers on WordPress?\u003C\u002Fh4>\n\u003Cp>AI companies use automated \u003Cstrong>AI crawlers\u003C\u002Fstrong> to index, train on, or retrieve website content. Common examples include \u003Cstrong>GPTBot\u003C\u002Fstrong> (OpenAI), \u003Cstrong>ClaudeBot\u003C\u002Fstrong> (Anthropic), \u003Cstrong>Google-Extended\u003C\u002Fstrong> (Google), \u003Cstrong>Bytespider\u003C\u002Fstrong> (ByteDance), \u003Cstrong>CCBot\u003C\u002Fstrong> (Common Crawl), \u003Cstrong>PerplexityBot\u003C\u002Fstrong>, and many others.\u003C\u002Fp>\n\u003Cp>With this AI crawlers plugin you can:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>See which AI crawlers visit your site and which pages they request.\u003C\u002Fli>\n\u003Cli>Compare AI bot traffic across date ranges and spot trends over time.\u003C\u002Fli>\n\u003Cli>Block AI crawlers with robots.txt rules or PHP-level 403 responses.\u003C\u002Fli>\n\u003Cli>Keep logs for compliance, debugging, and content licensing decisions.\u003C\u002Fli>\n\u003Cli>See a 7-day crawler summary on the WordPress admin dashboard.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>How it works\u003C\u002Fh4>\n\u003Col>\n\u003Cli>Enable crawler tracking and choose how long activity logs should be retained.\u003C\u002Fli>\n\u003Cli>Review crawler trends, top bots, most-crawled pages, and recent requests.\u003C\u002Fli>\n\u003Cli>Use robots.txt or optional PHP blocking for crawlers you do not want accessing your content.\u003C\u002Fli>\n\u003Cli>Run an AI Visibility Score scan and enable the discoverability features that fit your site.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch4>AI crawler analytics dashboard\u003C\u002Fh4>\n\u003Cp>Understand which AI systems are interested in your site instead of relying on guesswork.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Dashboard with total requests, unique AI crawlers, pages crawled, and blocked attempts.\u003C\u002Fli>\n\u003Cli>Charts for AI crawler requests over time and traffic by bot category.\u003C\u002Fli>\n\u003Cli>Top AI crawlers and most-crawled pages tables with period comparison.\u003C\u002Fli>\n\u003Cli>Recent AI crawler activity log with filters for bot, HTTP status, and content type.\u003C\u002Fli>\n\u003Cli>Download CSV analytics reports for preset or custom date ranges.\u003C\u002Fli>\n\u003Cli>Blocked-attempt logging when PHP blocking is enabled.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>AI Visibility Score\u003C\u002Fh4>\n\u003Cp>Find technical signals that can limit AI discovery and get prioritized recommendations for improvement.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Analyze any page or your entire site for AI readiness.\u003C\u002Fli>\n\u003Cli>26 automated checks across discovery, structured data, content structure, and technical signals.\u003C\u002Fli>\n\u003Cli>0–100 score with letter grade and prioritized recommendations.\u003C\u002Fli>\n\u003Cli>Site-wide batch scan, scan history, comparison diff, and PDF export.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>AI discoverability: llms.txt, Markdown, and JSON-LD\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>Auto-generate \u003Ccode>\u002Fllms.txt\u003C\u002Fcode> and optional \u003Ccode>\u002Fllms-full.txt\u003C\u002Fcode> for AI systems.\u003C\u002Fli>\n\u003Cli>Choose public post types, set URL limits, and include or exclude individual content.\u003C\u002Fli>\n\u003Cli>Exclude content by URL pattern and automatically exclude noindex content.\u003C\u002Fli>\n\u003Cli>Respect noindex metadata from Yoast SEO, Rank Math, All in One SEO, SEOPress, and Genesis.\u003C\u002Fli>\n\u003Cli>Regenerate discovery files when content is published or on a schedule.\u003C\u002Fli>\n\u003Cli>Optionally add \u003Ccode>LLMs:\u003C\u002Fcode> and \u003Ccode>LLMs-full:\u003C\u002Fcode> references to robots.txt.\u003C\u002Fli>\n\u003Cli>Serve \u003Ccode>.md\u003C\u002Fcode> URL endpoints for public content and taxonomy archives.\u003C\u002Fli>\n\u003Cli>Respond to \u003Ccode>Accept: text\u002Fmarkdown\u003C\u002Fcode> requests from compatible AI agents.\u003C\u002Fli>\n\u003Cli>Output Organization, WebSite, and optional Article JSON-LD with AI discovery pointers.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>73+ built-in AI crawlers\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>Pre-configured detection for major AI, search, and scraper bots.\u003C\u002Fli>\n\u003Cli>Category labels (training, search, assistant, scraper, and more).\u003C\u002Fli>\n\u003Cli>Per-crawler info with user-agent pattern details.\u003C\u002Fli>\n\u003Cli>Enable or disable monitoring for individual AI crawlers.\u003C\u002Fli>\n\u003Cli>Includes crawlers associated with OpenAI, Anthropic, Google, Perplexity, Meta, Amazon, Apple, ByteDance, Common Crawl, and others.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Custom AI crawler detection\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>Add your own user-agent patterns for bots not in the built-in list.\u003C\u002Fli>\n\u003Cli>Assign company name and category for consistent reporting.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>robots.txt control for AI crawlers\u003C\u002Fh4>\n\u003Cp>Publish clear crawler preferences and identify bots that continue visiting after being disallowed.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Add Disallow rules for selected AI crawlers.\u003C\u002Fli>\n\u003Cli>Bulk templates to disallow or allow entire categories (training, scraper, and more).\u003C\u002Fli>\n\u003Cli>Compliance monitor shows top non-compliant crawlers that still visit.\u003C\u002Fli>\n\u003Cli>Rules are appended to your WordPress robots.txt output.\u003C\u002Fli>\n\u003Cli>Live preview before you publish changes.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Block AI crawlers with PHP (403)\u003C\u002Fh4>\n\u003Cp>robots.txt communicates a preference; PHP blocking enforces access control at the WordPress application level.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Hard-block selected AI crawlers before WordPress serves a response.\u003C\u002Fli>\n\u003Cli>Block specific IP addresses with the same 403 PHP response.\u003C\u002Fli>\n\u003Cli>Customize the plain-text message returned to blocked bots (since 1.1.0).\u003C\u002Fli>\n\u003Cli>Blocked requests are logged in the AI crawler analytics dashboard.\u003C\u002Fli>\n\u003Cli>Works alongside robots.txt for stronger enforcement when needed.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Data management\u003C\u002Fh4>\n\u003Cp>Crawler activity is stored in your WordPress database, giving you control over retention and deletion.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Configurable log retention (7 days to 1 year, or keep all data).\u003C\u002Fli>\n\u003Cli>Optional delete-all-data action from the admin.\u003C\u002Fli>\n\u003Cli>Delete plugin data on uninstall (optional setting).\u003C\u002Fli>\n\u003Cli>Multi-select crawler, category, HTTP status, and content-type activity filters.\u003C\u002Fli>\n\u003Cli>Download reports for preset or custom date ranges.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Reports and integrations\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>Daily, weekly, or monthly HTML email reports with configurable recipients and alert types.\u003C\u002Fli>\n\u003Cli>Email summaries cover crawler activity, traffic changes, blocked attempts, new crawlers, and robots.txt non-compliance.\u003C\u002Fli>\n\u003Cli>Test email delivery and view the next scheduled report from the Reports page.\u003C\u002Fli>\n\u003Cli>Authenticated REST API for analytics summaries, recent activity, CSV exports, visibility scans, and blocked attempts.\u003C\u002Fli>\n\u003Cli>Professional report emails include key metrics, the top crawler, and most-crawled pages.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Clean admin experience\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>WordPress admin menu: \u003Cstrong>AI Crawlers\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>Redesigned UI with grouped sidebar navigation and dark mode.\u003C\u002Fli>\n\u003Cli>Dashboard, Crawlers, Control, Visibility Score, Reports, and Settings in one place.\u003C\u002Fli>\n\u003Cli>Built for site administrators — no code required.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Who is this plugin for?\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>\u003Cstrong>Bloggers and publishers\u003C\u002Fstrong> who want to know when AI crawlers access their articles.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Business and marketing sites\u003C\u002Fstrong> tracking AI bot traffic to key landing pages.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Developers and agencies\u003C\u002Fstrong> managing AI crawler policy across client WordPress sites.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Site owners\u003C\u002Fstrong> exploring robots.txt and PHP options for AI crawler control.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch4>Privacy and data ownership\u003C\u002Fh4>\n\u003Cul>\n\u003Cli>AI crawler activity logs are stored locally in your WordPress database.\u003C\u002Fli>\n\u003Cli>The analytics dashboard tracks detected crawler requests, not general human visitor analytics.\u003C\u002Fli>\n\u003Cli>Log retention and deletion are controlled from the plugin settings.\u003C\u002Fli>\n\u003Cli>Scheduled reports are sent through the standard WordPress \u003Ccode>wp_mail()\u003C\u002Fcode> function.\u003C\u002Fli>\n\u003Cli>Core crawler monitoring does not require a separate analytics account.\u003C\u002Fli>\n\u003C\u002Ful>\n","Monitor 73+ AI crawlers, analyze bot traffic, manage robots.txt, block unwanted bots, and improve AI visibility with llms.txt and readiness checks.",246,"2026-07-21T08:34:00.000Z",[54,74,55,56,57],"ai-seo","https:\u002F\u002Fplugins.miniorange.com\u002F","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fminiorange-ai-visibility.zip",{"slug":78,"name":79,"version":80,"author":81,"author_profile":82,"description":83,"short_description":84,"active_installs":13,"downloaded":85,"rating":13,"num_ratings":13,"last_updated":86,"tested_up_to":15,"requires_at_least":87,"requires_php":52,"tags":88,"homepage":59,"download_link":90,"security_score":26,"vuln_count":13,"unpatched_count":13,"last_vuln_date":27,"fetched_at":28},"puffersights-ai-crawler-insights","PufferSights – AI Crawler Insights","0.1.0","senols","https:\u002F\u002Fprofiles.wordpress.org\u002Fsenols\u002F","\u003Cp>PufferSights monitors 100+ known AI crawler and AI agent user agents, hashes IP addresses, groups traffic by bot, provider, crawl purpose, content type, and response status, tracks human referrals from AI surfaces, and can publish a dynamic llms.txt content map for public site content.\u003C\u002Fp>\n\u003Cp>The dashboard summarizes:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>HTTP traffic by bot.\u003C\u002Fli>\n\u003Cli>Crawl purpose.\u003C\u002Fli>\n\u003Cli>Content type.\u003C\u002Fli>\n\u003Cli>Response status.\u003C\u002Fli>\n\u003Cli>AI referrals and crawl-to-refer ratio.\u003C\u002Fli>\n\u003Cli>Top crawled content.\u003C\u002Fli>\n\u003Cli>Tracked agent count.\u003C\u002Fli>\n\u003Cli>Dynamic llms.txt content map.\u003C\u002Fli>\n\u003Cli>robots.txt audit and policy snippets.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The crawler registry is based on current public operator documentation and industry references for OpenAI, Anthropic, Perplexity, Google, Apple, Common Crawl, Meta, ByteDance, Microsoft, Amazon, and related AI crawler operators.\u003C\u002Fp>\n\u003Cp>The plugin does not contact any external service. All analytics data is stored in your own WordPress database.\u003C\u002Fp>\n\u003Cp>robots.txt publishing is off by default. The plugin can generate and optionally publish policies for:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Monitor only.\u003C\u002Fli>\n\u003Cli>Block training crawlers.\u003C\u002Fli>\n\u003Cli>Allow AI search\u002Fuser-action bots while blocking training crawlers.\u003C\u002Fli>\n\u003Cli>Block all known AI bots.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>llms.txt publishing is on by default and can be disabled in the PufferSights settings. The generated \u003Ccode>\u002Fllms.txt\u003C\u002Fcode> file lists selected published public pages and posts in Markdown so AI assistants can find the site’s main public content more easily.\u003C\u002Fp>\n\u003Ch3>Important Notes\u003C\u002Fh3>\n\u003Cp>User-agent detection is not bot verification. User agents can be spoofed. Raw IP addresses are not stored; the plugin stores a salted hash for rough uniqueness.\u003C\u002Fp>\n\u003Cp>robots.txt is voluntary. Use a WAF, CDN, or server-level controls when technical enforcement is required.\u003C\u002Fp>\n\u003Cp>Google-Extended and Applebot-Extended are robots.txt control tokens rather than normal request user agents, so they appear in robots.txt audits and policy snippets but usually do not appear in request logs.\u003C\u002Fp>\n\u003Cp>llms.txt is a content map, not an access-control policy. It does not replace robots.txt and does not force AI systems to use or cite your content.\u003C\u002Fp>\n\u003Ch3>Privacy\u003C\u002Fh3>\n\u003Cp>PufferSights stores local analytics for public, logged-out requests only. It does not track wp-admin pages, logged-in users, AJAX requests, or WP-Cron requests.\u003C\u002Fp>\n\u003Cp>The plugin stores:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Request time and date.\u003C\u002Fli>\n\u003Cli>Event type, such as AI crawler request or AI referral.\u003C\u002Fli>\n\u003Cli>HTTP method.\u003C\u002Fli>\n\u003Cli>Request path without query string.\u003C\u002Fli>\n\u003Cli>HTTP response status.\u003C\u002Fli>\n\u003Cli>MIME\u002Fcontent group.\u003C\u002Fli>\n\u003Cli>Matched crawler or AI referral provider.\u003C\u002Fli>\n\u003Cli>User-agent string and user-agent hash.\u003C\u002Fli>\n\u003Cli>Salted one-way hash of the request IP address.\u003C\u002Fli>\n\u003Cli>Referrer origin only, such as \u003Ccode>https:\u002F\u002Fchatgpt.com\u003C\u002Fcode>, without referrer path or query string.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The plugin does not store raw IP addresses, cookies, browser local storage, or complete referrer URLs. It does not send analytics, telemetry, crawler records, or site data to third-party services.\u003C\u002Fp>\n\u003Cp>If llms.txt publishing is enabled, the plugin serves a Markdown overview of selected published public posts and pages at \u003Ccode>\u002Fllms.txt\u003C\u002Fcode>. Drafts, private posts, and password-protected posts are not included.\u003C\u002Fp>\n\u003Cp>Administrators can disable tracking, clear captured events, and configure retention from the PufferSights admin page. The default retention period is 90 days. On uninstall, the plugin removes its custom analytics table, saved options, and scheduled cleanup hook.\u003C\u002Fp>\n\u003Cp>The plugin also adds suggested disclosure text to WordPress’ Privacy Policy Guide.\u003C\u002Fp>\n","Monitor 100+ known AI crawler and AI agent user agents with local analytics, AI referrals, llms.txt, and opt-in robots.txt policy tools.",126,"2026-06-16T08:04:00.000Z","6.5",[19,89,20,21,56],"analytics","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fpuffersights-ai-crawler-insights.0.1.0.zip",{"slug":92,"name":93,"version":94,"author":95,"author_profile":96,"description":97,"short_description":98,"active_installs":99,"downloaded":100,"rating":101,"num_ratings":102,"last_updated":103,"tested_up_to":15,"requires_at_least":104,"requires_php":17,"tags":105,"homepage":59,"download_link":107,"security_score":26,"vuln_count":13,"unpatched_count":13,"last_vuln_date":27,"fetched_at":28},"block-ai-crawlers","Block AI Crawlers","1.5.8","lastsplash (a11n)","https:\u002F\u002Fprofiles.wordpress.org\u002Flastsplash\u002F","\u003Cp>Protect Your Content from AI Scraping\u003C\u002Fp>\n\u003Cp>This plugin helps you prevent AI crawlers from using your content as training data for their products. By updating your site’s \u003Ccode>robots.txt\u003C\u002Fcode>, it blocks common AI crawlers and scrapers, aiming to protect your content from being used in the training of Large Language Models (LLMs).\u003C\u002Fp>\n\u003Ch3>Features\u003C\u002Fh3>\n\u003Ch3>Blocks AI Crawlers\u003C\u002Fh3>\n\u003Cp>Includes:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>OpenAI\u003C\u002Fstrong> – Blocks crawlers used for ChatGPT\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Google\u003C\u002Fstrong> – Blocks crawlers used by Google’s Gemini AI products\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Meta\u003C\u002Fstrong> – Blocks FacebookBot and Meta training crawlers\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Anthropic\u003C\u002Fstrong> – Blocks crawlers used by Claude\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Perplexity\u003C\u002Fstrong> – Blocks crawlers used by Perplexity\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Apple\u003C\u002Fstrong> – Blocks Applebot-Extended\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Amazon\u003C\u002Fstrong> – Blocks Amazonbot\u003C\u002Fli>\n\u003Cli>…and 150+ more via ai.robots.txt\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The blocked crawler list is generated at build time from \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fai-robots-txt\u002Fai.robots.txt\" rel=\"nofollow ugc\">ai.robots.txt\u003C\u002Fa> (\u003Ccode>robots.json\u003C\u002Fcode>, MIT). See the plugin’s \u003Ccode>THIRD-PARTY.md\u003C\u002Fcode> for license attribution.\u003C\u002Fp>\n\u003Ch3>Experimental Meta Tags\u003C\u002Fh3>\n\u003Cp>The plugin adds the “noai, noimageai” directive to your site’s meta tags, instructing AI bots not to use your content in their datasets. Please note that these tags are experimental and have not been standardized.\u003C\u002Fp>\n\u003Ch3>Custom robots.txt Rules\u003C\u002Fh3>\n\u003Cp>Have custom entries for your robots.txt file? You can now add them directly through the plugin!\u003C\u002Fp>\n\u003Ch3>Usage\u003C\u002Fh3>\n\u003Cp>After activation, the plugin will automatically update your \u003Ccode>robots.txt\u003C\u002Fcode> and add the necessary meta tags. No further configuration is required, but you can check the settings page for a full list of blocked crawlers.\u003C\u002Fp>\n\u003Ch3>Limitations\u003C\u002Fh3>\n\u003Cp>While this plugin aims to block specified crawlers, it cannot guarantee complete protection against all forms of scraping, as some bots may disregard \u003Ccode>robots.txt\u003C\u002Fcode> directives.\u003C\u002Fp>\n\u003Ch3>Support\u003C\u002Fh3>\n\u003Cp>For questions or support, \u003Ca href=\"https:\u002F\u002Fwordpress.org\u002Fsupport\u002Fplugin\u002Fblock-ai-crawlers\u002F\" rel=\"ugc\">please post on the forums\u003C\u002Fa>.\u003C\u002Fp>\n","Tell AI (Artificial Intelligence) companies not to scrape your site for their AI products.",1000,17841,92,7,"2026-07-18T12:43:00.000Z","6.9",[19,106,21,57],"chatgpt","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fblock-ai-crawlers.1.5.8.zip",{"slug":109,"name":110,"version":111,"author":112,"author_profile":113,"description":114,"short_description":115,"active_installs":116,"downloaded":117,"rating":26,"num_ratings":118,"last_updated":119,"tested_up_to":15,"requires_at_least":120,"requires_php":52,"tags":121,"homepage":123,"download_link":124,"security_score":26,"vuln_count":13,"unpatched_count":13,"last_vuln_date":27,"fetched_at":28},"ai-content-signals","AI Content Signals","1.3.0","Fernando Tellado","https:\u002F\u002Fprofiles.wordpress.org\u002Ffernandot\u002F","\u003Cp>AI Content Signals lets you declare, in a machine-readable way, how AI systems may use your content: for search indexing, real-time AI answers (RAG), or model training. It started with Cloudflare’s Content Signals in robots.txt and now expresses the same preferences across several surfaces, so more crawlers and tools can read them:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>robots.txt\u003C\u002Fstrong> — Cloudflare Content Signals (search \u002F ai-input \u002F ai-train)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>HTTP header\u003C\u002Fstrong> — \u003Ccode>X-Robots-Tag: noai, noimageai\u003C\u002Fcode>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>HTML meta\u003C\u002Fstrong> — robots \u003Ccode>noai, noimageai\u003C\u002Fcode>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>\u002F.well-known\u002Ftdmrep.json\u003C\u002Fstrong> — W3C Text and Data Mining Reservation Protocol\u003C\u002Fli>\n\u003Cli>\u003Cstrong>EU Directive 2019\u002F790\u003C\u002Fstrong> rights reservation\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Everything is opt-in and does not change your robots.txt output unless you enable it.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The signals you control\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>You set three preferences, and AI Content Signals expresses each one in the right format on every surface you enable:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>search\u003C\u002Fstrong> — allow or deny search indexing and traditional search results (links and short snippets)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>ai-input\u003C\u002Fstrong> — allow or deny using your content for real-time AI answers (RAG, grounding, AI Overviews)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>ai-train\u003C\u002Fstrong> — allow or deny using your content to train or fine-tune AI models\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>These three come from Cloudflare’s Content Signals vocabulary, written to your robots.txt. When you deny AI training, the same opt-out is also emitted on any other surface you enable: \u003Ccode>noai, noimageai\u003C\u002Fcode> in an HTML meta tag and an \u003Ccode>X-Robots-Tag\u003C\u002Fcode> header, plus a \u003Ccode>tdm-reservation\u003C\u002Fcode> in your TDMRep manifest. Declaring the same preference in several places means a crawler that ignores one signal may still honor another.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Key Features\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Easy-to-use settings page in WordPress admin\u003C\u002Fli>\n\u003Cli>Set global defaults for all crawlers\u003C\u002Fli>\n\u003Cli>Configure specific settings for individual AI bots (GPTBot, ClaudeBot, PerplexityBot, etc.)\u003C\u002Fli>\n\u003Cli>Add custom bot User-Agents\u003C\u002Fli>\n\u003Cli>Supports both physical and virtual robots.txt files\u003C\u002Fli>\n\u003Cli>Option to create physical robots.txt with basic WordPress rules\u003C\u002Fli>\n\u003Cli>Preview generated Content Signals before applying\u003C\u002Fli>\n\u003Cli>Export and import settings as JSON for easy migration between sites\u003C\u002Fli>\n\u003Cli>Optional legal text with EU Directive reference\u003C\u002Fli>\n\u003Cli>Developer-friendly: filter hook to extend the predefined bots list\u003C\u002Fli>\n\u003Cli>Works with existing robots.txt from SEO plugins\u003C\u002Fli>\n\u003Cli>Automatic sitemap detection and inclusion\u003C\u002Fli>\n\u003Cli>Optional extra output surfaces: HTML robots meta tag (noai, noimageai), X-Robots-Tag header, and a W3C TDMRep file at \u002F.well-known\u002Ftdmrep.json\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cstrong>Supported Bots\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>The plugin includes predefined settings for 28 major AI crawlers:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>OpenAI GPTBot, OAI-SearchBot, and ChatGPT-User\u003C\u002Fli>\n\u003Cli>Anthropic ClaudeBot, Claude-Web, and anthropic-ai\u003C\u002Fli>\n\u003Cli>Perplexity Bot and Perplexity-User\u003C\u002Fli>\n\u003Cli>Google Extended (Gemini) and GoogleOther\u003C\u002Fli>\n\u003Cli>Amazon Bot\u003C\u002Fli>\n\u003Cli>Apple Extended\u003C\u002Fli>\n\u003Cli>Meta\u002FFacebook Bot and meta-externalagent\u003C\u002Fli>\n\u003Cli>DuckDuckGo DuckAssistBot\u003C\u002Fli>\n\u003Cli>Allen Institute AI2Bot\u003C\u002Fli>\n\u003Cli>Mistral AI\u003C\u002Fli>\n\u003Cli>ByteDance Bytespider\u003C\u002Fli>\n\u003Cli>DeepSeek AI\u003C\u002Fli>\n\u003Cli>xAI Grok\u003C\u002Fli>\n\u003Cli>Huawei Pangu\u003C\u002Fli>\n\u003Cli>Common Crawl, Cohere AI, Diffbot, You.com Bot, and more\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cstrong>Important Notice\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>Content Signals is a declarative standard – it expresses your preferences but does not technically enforce them. AI companies are not legally required to respect these signals, though the plugin includes legal text referencing EU copyright directives.\u003C\u002Fp>\n\u003Cp>The IETF AI Preferences (AIPREF) Working Group is currently developing a formal standard based on similar concepts. This plugin implements the current Cloudflare Content Signals specification and will be updated as standards evolve.\u003C\u002Fp>\n\u003Cp>This plugin works best when combined with other protection measures like traditional robots.txt rules and server-level bot management.\u003C\u002Fp>\n\u003Ch3>Support\u003C\u002Fh3>\n\u003Cp>Need private support or custom development?\u003C\u002Fp>\n\u003Cp>Do you need one-on-one help, priority troubleshooting, or a custom feature, integration, or tweak built specifically for your site? I offer private support and custom development. Just \u003Ca href=\"mailto:ai-content-signals@ayudawp.com\" rel=\"nofollow ugc\">contact me\u003C\u002Fa> and tell me what you need.\u003C\u002Fp>\n\u003Cp>Need help or have suggestions?\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fservicios.ayudawp.com\" rel=\"nofollow ugc\">Official website\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fwordpress.org\u002Fsupport\u002Fplugin\u002Fai-content-signals\u002F\" rel=\"ugc\">WordPress support forum\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fwww.youtube.com\u002FAyudaWordPressES\" rel=\"nofollow ugc\">YouTube channel\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fayudawp.com\" rel=\"nofollow ugc\">Documentation and tutorials\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Love the plugin? Please leave us a \u003Ca href=\"https:\u002F\u002Fwordpress.org\u002Fsupport\u002Fplugin\u002Fai-content-signals\u002Freviews\u002F#new-post\" rel=\"ugc\">5-star review\u003C\u002Fa> and help spread the word!\u003C\u002Fp>\n\u003Ch3>About AyudaWP.com\u003C\u002Fh3>\n\u003Cp>We are specialists in WordPress security, SEO, and performance optimization plugins. We create tools that solve real problems for WordPress site owners while maintaining the highest coding standards and accessibility requirements.\u003C\u002Fp>\n","Control how AI uses your content — search, AI answers or training — via robots.txt Content Signals, HTTP headers, meta tags and TDMRep.",500,3221,3,"2026-06-16T17:57:00.000Z","5.6",[19,122,21,57,58],"cloudflare","https:\u002F\u002Fservicios.ayudawp.com","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fai-content-signals.1.3.0.zip",{"error":126,"url":127,"statusCode":128,"statusMessage":129,"message":129},true,"http:\u002F\u002Flocalhost\u002Fapi\u002Fplugins\u002Fparseless\u002Fbundle",404,"no bundle for this plugin yet",{"slug":4,"current_version":6,"total_versions":131,"versions":132},5,[133,139,146,153,160],{"version":6,"download_url":25,"svn_tag_url":134,"released_at":27,"has_diff":135,"diff_files_changed":136,"diff_lines":27,"trac_diff_url":137,"vulnerabilities":138,"is_current":126},"https:\u002F\u002Fplugins.svn.wordpress.org\u002Fparseless\u002Ftags\u002F0.6.0\u002F",false,[],"https:\u002F\u002Fplugins.trac.wordpress.org\u002Fchangeset?old_path=%2Fparseless%2Ftags%2F0.5.0&new_path=%2Fparseless%2Ftags%2F0.6.0",[],{"version":140,"download_url":141,"svn_tag_url":142,"released_at":27,"has_diff":135,"diff_files_changed":143,"diff_lines":27,"trac_diff_url":144,"vulnerabilities":145,"is_current":135},"0.5.0","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fparseless.0.5.0.zip","https:\u002F\u002Fplugins.svn.wordpress.org\u002Fparseless\u002Ftags\u002F0.5.0\u002F",[],"https:\u002F\u002Fplugins.trac.wordpress.org\u002Fchangeset?old_path=%2Fparseless%2Ftags%2F0.4.0&new_path=%2Fparseless%2Ftags%2F0.5.0",[],{"version":147,"download_url":148,"svn_tag_url":149,"released_at":27,"has_diff":135,"diff_files_changed":150,"diff_lines":27,"trac_diff_url":151,"vulnerabilities":152,"is_current":135},"0.4.0","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fparseless.0.4.0.zip","https:\u002F\u002Fplugins.svn.wordpress.org\u002Fparseless\u002Ftags\u002F0.4.0\u002F",[],"https:\u002F\u002Fplugins.trac.wordpress.org\u002Fchangeset?old_path=%2Fparseless%2Ftags%2F0.3&new_path=%2Fparseless%2Ftags%2F0.4.0",[],{"version":154,"download_url":155,"svn_tag_url":156,"released_at":27,"has_diff":135,"diff_files_changed":157,"diff_lines":27,"trac_diff_url":158,"vulnerabilities":159,"is_current":135},"0.3","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fparseless.0.3.zip","https:\u002F\u002Fplugins.svn.wordpress.org\u002Fparseless\u002Ftags\u002F0.3\u002F",[],"https:\u002F\u002Fplugins.trac.wordpress.org\u002Fchangeset?old_path=%2Fparseless%2Ftags%2F0.3.0&new_path=%2Fparseless%2Ftags%2F0.3",[],{"version":161,"download_url":162,"svn_tag_url":163,"released_at":27,"has_diff":135,"diff_files_changed":164,"diff_lines":27,"trac_diff_url":27,"vulnerabilities":165,"is_current":135},"0.3.0","https:\u002F\u002Fdownloads.wordpress.org\u002Fplugin\u002Fparseless.0.3.0.zip","https:\u002F\u002Fplugins.svn.wordpress.org\u002Fparseless\u002Ftags\u002F0.3.0\u002F",[],[]]