Technical GEO
72 practices
73% of Websites Have Technical Barriers Blocking AI Crawlers
OtterlyAI study of 1M+ citations found 73% of sites block AI crawlers via robots.txt, CDN rules, or JS rendering issues. Fix crawler access before anything else.
Sources
74% of Websites Are Invisible to AI Search โ Structured Data Is the Fixable Gap
SearchScore SAVI (850k sites, Q2 2026): 74.2% Invisible/Low visibility; AI visibility averages 34.1 vs technical health 70.1, with structured data scoring 23.1.
Sources
Agentic AI Search: Claude and GPT-5 Agents Now Execute Multi-Step Purchasing (Jun 2026)
Autonomous AI agents from Claude and GPT-5 now execute multi-step purchasing and comparison tasks. Agent-first content structure becomes critical for visibility.
Sources
Agentic search requires machine-readable pricing, availability, and checkout data โ protocols reshaping SEO
OpenAI Agentic Commerce Protocol (ACP, with Stripe), Web Capabilities Protocol (WebMCP, Google+Microsoft), Universal Commerce Protocol (UCP, Google+Shopify) โ sites without machine-readable
Sources
AI Agent Traffic Reaches 88% of Human Organic Search Volume; Predicted to Surpass by End 2026
BrightEdge data shows AI agent requests now rival human search traffic, with OpenAI agents dominating. Differentiated robots.txt strategy is essential.
Sources
AI Crawler Explosion โ GPTBot +305% YoY, AI Bots Reach 22% of All Bot Traffic
Presenc AI 2022-2026 trend: GPTBot grew 305% YoY. AI bots now 22% of all bot traffic. Meta-ExternalAgent surged to #2 at 16.7%. Bots = 31.2% of all HTTP requests; trajectory crosses human traffic
Sources
The #1 GEO robots.txt Mistake: Blocking OAI-SearchBot Instead of GPTBot Kills ChatGPT Citations
OpenAI runs 3 bots: GPTBot (training), OAI-SearchBot (search index), ChatGPT-User (fetches). Block OAI-SearchBot = zero ChatGPT citations. Most teams conflate them.
Sources
AI Crawler Traffic Now 3.4% of All Web Traffic; 12% of Top Sites Block OAI-SearchBot
Known Agents data shows AI search crawlers are a measurable and growing portion of web traffic. Differentiated robots.txt management is a basic GEO hygiene requirement.
Sources
AI Crawlers Now Account for 40-50% of Bot-Level Activity โ 65-70% Are Live Queries
JetOctopus server log analysis (Feb 2026): AI bots ~40-50% of Googlebot-level activity. 65-70% of AI bot traffic is user search, not training. 14+ AI user-agents need explicit Allow rules in
Sources
AI Crawlers Visit Once and Skip Homepages โ Blog Is the New Front Door
Trakkr (575,788 AI crawler visits): GPTBot reaches homepages only ~3% of the time; 88.5% of pages get exactly one visit; 21% of ChatGPT Search sessions start on blog pages.
Sources
AI Crawls Product Pages but Cites Blog Posts โ 337K Citation Mismatch
Trakkr cross-referenced 337K AI citations with 11.4M crawler visits: what AI crawls (product pages) is not what it cites (blog content). Measure both, not one.
Sources
AI Labs Don't Use llms.txt on Their Own Consumer Front Doors โ Docs Only
HTTP Archive: chatgpt.com, claude.ai, gemini.google.com lack llms.txt. But docs.anthropic.com, docs.perplexity.ai have one. Value is for agent-readiness, not marketing.
Sources
AI Mode + AI Overviews merged at Google I/O 2026 โ unified AI Search on Gemini 3.5 Flash
Google merged AI Overviews and AI Mode into one unified AI Search experience at I/O 2026. Runs on Gemini 3.5 Flash. Single citation pool eliminates need for separate strategies.
Sources
AI Mode + AI Overviews merged at Google I/O 2026 โ unified AI Search on Gemini 3.5 Flash
Google merged AI Overviews and AI Mode into one unified AI Search experience at I/O 2026. Runs on Gemini 3.5 Flash. Single citation pool eliminates need for separate strategies.
Sources
Anthropic Runs 3 Separate Crawler User-Agents With Distinct Functions
ClaudeBot (training), Claude-Web (real-time browsing), and anthropic-ai (catch-all) can be independently controlled in robots.txt for granular AI visibility management.
Sources
Bing Webmaster Tools + OAI-SearchBot Is the ChatGPT Citation Pipeline
ChatGPT Search retrieves from Bing's index: 87% of ChatGPT cited pages correspond to Bing top results (Mersel AI 2026). Submit sitemap to BWT, allow OAI-SearchBot, enable IndexNow.
Sources
ChatGPT Search cites only 15% of retrieved pages โ four-gate funnel model
AirOps study of 548K pages: ChatGPT Search retrieves via Bing index, checks OAI-SearchBot crawl access, then cites only 15% based on structure, freshness, and authority.
Sources
ChatGPT Sends 28.8% of Referrals to Internal Site Search Instead of the Answer Page (Previsible Jul 2026)
Previsible study: ChatGPT sends 28.8% of its referrals to internal search results pages. The model trusts the domain but defaults to site search when it can't identify the right page.
Sources
ChatGPT Citation Pipeline Is Two-Stage โ Bing Retrieval Then Fine-Tuned Re-Rank
ChatGPT uses Bing to retrieve candidates, then a fine-tuned model re-ranks by answer fit, domain authority, source consensus; zero JS execution.
Sources
Claude uses Brave Search (not Google/Bing) โ 86.7% overlap with Brave top 10, Brave SEO is prerequisite
Profound 2025 + multiple 2026 studies: Claude's web retrieval runs on Brave Search's independent index. Google rankings have limited transferability. Brave does not license from Google or Bing.
Sources
ClaudeBot Worst Crawl-to-Refer Ratio at 20,583 Pages Per Referral
Presenc AI study of top 1,000 sites: ClaudeBot crawls 20,583 pages per referral. PerplexityBot best at ~210:1. 25% block GPTBot. OAI-SearchBot at 85:1.
Sources
Cloudflare "Block AI Bots" Silently Kills ChatGPT Visibility
Cloudflare CDN-level bot blocking overrides robots.txt allow rules, an invisible cause of ChatGPT invisibility even when Bing-ranked.
Sources
Cloudflare Content Signals Policy โ Cite But Don't Train
New robots.txt extension declares post-fetch usage (search/ai-input/ai-train). Set ai-train=no while ai-input=yes to stay citable but not trained on.
Sources
Cloudflare Three-Tier AI Crawler Classification โ Search, Agent, Training
From Sep 15, 2026, Cloudflare blocks multi-purpose crawlers by default on ad-supported pages. Publishers can granularly allow/block each tier.
Sources
Crawl-to-Citation Efficiency Varies 10x Across AI Engines โ Perplexity Most Efficient
Presenc AI (April 2026) joins crawl events to citation outcomes. PerplexityBot 5-10x more efficient than GPTBot. Optimize differently per engine โ Perplexity for fetchable content, Claude for
Sources
GEO attacks can promote flawed products into AI recommendations by up to 83.2%
Seller-controlled GEO rewrites promote flawed products into LLM recommendation sets by up to 83.2%; structured evidence checks cut the harm by up to 39.2% (SafeGEO, 600 cases).
Sources
Google AI Mode Information Agents launched as push-based referral surface
Google launched always-on AI Mode information agents on June 12 2026 for Ultra subscribers. Unlike zero-click AIO, they push source-linked updates proactively to users.
Sources
Google AI Overviews Introduces Hover Pop-Up Link Cards (February 2026)
Google rolled out hover pop-up link cards in AI Overviews and AI Mode: hovering over highlighted text shows link cards for direct source access, creating a new path for user clicks from AI summaries.
Sources
Google AI Overviews now generates AI images directly in search results
Google launched AI image generation directly within AI Overviews (July 2026) using the Nano Banana model. Rolling out in English for AI Mode-supported countries.
Sources
Google I/O 2026: AI Mode becomes default, Search agents, Gemini 3.5, Personal Intelligence
Google announced AI Mode as default replacement for traditional search, Gemini 3.5 Flash as default model, information agents, and Personal Intelligence across 98 languages. AI Mode crossed 1B MAU.
Sources
Google I/O 2026: AI Overviews and AI Mode Merged Into Unified Gemini 3.5 Flash Surface
Google merged AI Overviews and AI Mode into one unified AI Search layer on Gemini 3.5 Flash, reaching 1B+ monthly users, eliminating separate citation strategies.
Sources
Google Lighthouse now audits llms.txt and agentic browsing readiness
Chrome Lighthouse 13.3 (May 2026) added an Agentic Browsing audit category testing llms.txt presence, WebMCP support, accessibility-tree quality and CLS โ an agent-era readiness signal.
Sources
Google States No Special AI Schema Required for AI Overviews or AI Mode
Google confirmed in 2026 that no AI-specific markup or llms.txt is required for AIO or AI Mode: 'GEO is still SEO at the core.'
Sources
Google S-CTS and S-BERT Systems Detect AI-Generated Content Networks Within Days of New Model Launch
Google's S-CTS terminates entire AI content networks (not individual pages), while S-BERT detects the mathematical 'fingerprint' of AI text. New model outputs detectable within days via LoRA updates.
Sources
Google SAGE research: AI agents pull from top 3 ranked pages โ fundamentals matter more
Google's SAGE paper (Jan 2026) reveals AI agents perform multi-step searches. Top-3 rankings are the gateway to agentic discovery. Four 'shortcut' patterns identified.
Sources
ChatGPT Business reads Bing for citations; Plus reads Google via Bright Data scraper
TUM thesis study of 370K+ results found ChatGPT Business citations are 95% Bing-sourced, Plus citations 94% Bright Data (Google scraper).
Sources
Grok DeepSearch Uses IndexNow: Fresh Content Can Appear in Citations Within 2-4 Weeks
FuelOnline analysis: Grok DeepSearch draws on Bing's index. Submitting via IndexNow immediately after publishing is the fastest path to Grok citation visibility.
Sources
Grok xAI Crawler Silently Blocked by Default Security Rules Matching 'bot' Pattern
ProAISearch found Grok's crawler (user agents 'xAI' and 'Grok') lacks 'bot' in its name, so most default security setups and WAF rules silently block it. Content never gets indexed.
Sources
Grounding queries reveal that most AI influence is invisible (Otterly Copilot data)
Otterly's 3-month Copilot data: 647 grounding queries, 30,398 grounding events; 5 pages carried 74.6% of citations, 99.6% of AI influence invisible. Fix high-grounding/low-citation pages.
Sources
HasData: 56.4% of News Publishers Block AI Crawlers, 39.5% of Blocks Fail in Practice
HasData analysis of 10,894 domains (July 2026): 56.4% of news publishers block at least one AI crawler. GPTBot banned by 50.5% of publishers. 39.5% of GPTBot blocks fail to actually block.
Sources
Indirect Prompt Injection Is a Live Web Threat Against AI Crawlers and Agents
Google's April 2026 web sweep found attackers seeding prompt injections on sites to corrupt browsing AI; data-layer governance now required.
JavaScript-rendered content fails AI parsing 77% of the time
Erlin's 2026 data: AI parse success rates โ static HTML with schema 94%, plain HTML 68%, JS-rendered 23%, PDF 7%. JS-rendered pricing/feature pages often read as empty templates.
Sources
June 2026 Spam Update โ SpamBrain Extends to AI Overviews/AI Mode Enforcement
Google's June 24-26 2026 Spam Update extended SpamBrain enforcement to AI Overviews and AI Mode. Tactics to game AI answers are now treated as spam equivalent to paid links.
Sources
llms.txt Has 784+ Implementations But No Major Provider Confirms Using It โ MCP May Supersede
Presenc AI State of llms.txt 2026: 784+ implementations but only Perplexity and Anthropic (Claude Desktop) confirmed use. OpenAI, Google, Meta silent. MCP emerging as alternative.
Sources
llms.txt State of Adoption 2026: Support Confirmed, Sector Gaps Remain
Presenc AI April 2026 report: Anthropic and Perplexity confirmed llms.txt support. Adoption <10% in finance/healthcare, high in developer SaaS. Top-100 brands lead.
Sources
llms.txt adoption reached 4-5% of mid-market sites โ moderate impact on citations, strong for agent discovery
llms.txt (proposed by Jeremy Howard Sep 2024) adopted by 4-5% of mid-market sites by mid-2026. Current evidence shows no direct citation boost but helps AI agents discover pages faster.
Sources
llms.txt and AI Discovery File Suite Becomes Competitive Differentiator
One-van firm scored ChatGPT 99/100 against national brands using 9 AI Discovery Files. Backlinks didn't determine AI assessment โ structured facts did.
llms.txt has negligible short-term impact on AI search visibility โ OtterlyAI 90-day experiment
OtterlyAI's 90-day controlled experiment: out of 62.1K total AI bot hits, only 84 went to /llms.txt. Ahrefs: no major LLM provider currently supports llms.txt. Adoption is ~2% across 1M sites.
Sources
llms.txt implementation documented 5x AI traffic increase โ single highest-leverage technical GEO tactic
The Concurate case study documented a 5x increase in AI-referred traffic after implementing llms.txt. An llms.txt file at the domain root signals content priorities to LLM crawlers.
Sources
llms.txt shows zero measurable citation lift across 300K+ domain study
SE Ranking analyzed 300K domains: llms.txt has no effect on AI citations. Removing the variable improved model accuracy. 7.4% adoption in top 10K sites. Niche value for developer docs only.
Sources
llms.txt: 800K+ Sites Published, 97% Get Zero AI Requests, Google Explicitly Ignores It
Ahrefs studied 137,000 domains: ~28% publish llms.txt but 97% of those files got zero fetches in May 2026. Google confirms it ignores the file; coding agents fetch it ~5x more than AI search bots.
Sources
llms.txt adoption reached 4-5% of mid-market sites โ moderate impact on citations, strong for agent discovery
llms.txt (proposed by Jeremy Howard Sep 2024) adopted by 4-5% of mid-market sites by mid-2026. Current evidence shows no direct citation boost but helps AI agents discover pages faster.
Sources
llms.txt shows zero measurable citation lift across 300K+ domain study
SE Ranking analyzed 300K domains: llms.txt has no effect on AI citations. Removing the variable improved model accuracy. 7.4% adoption in top 10K sites. Niche value for developer docs only.
Sources
Markdown Is Not an AI-SEO Shortcut โ Bots Find It Four Ways, Use Varies
Botify's live-server experiment: bots discover Markdown via links, head tags, llms.txt, or content negotiation. GPTBot took 1,273 reads via links/head; purpose-built AI tools request Markdown ~100%
Sources
Most AI Crawler Traffic Is for Training, Not Indexing: 53.7% of AI Bots Fetch for Training
Foglift classified 65,527 AI crawler requests: 53.7% training, 32% indexing, 14.3% user-triggered fetch. Blocking training bots won't stop citation crawls.
Sources
AI Crawlers Operate on ~2s Hard Timeout โ FCP Under 0.4s Correlates with 6.7 AI Citations
Two converging findings: AI systems fetch pages with ~2s timeout (HTTP 499), and SE Ranking finds pages with FCP under 0.4s average 6.7 citations vs 2.1 for slower pages โ a 3.2x gap.
Sources
Perplexity Penalizes Gated Content: Gartner Receives 0 Citations Despite 130 Total Across Six Engines
Machine Relations Index measurement shows Perplexity's retrieval architecture systematically deprioritizes paywalled sources โ affecting how B2B brands with gated content appear.
Sources
Perplexity post-trains models for cross-source evidence synthesis and accuracy (Aug 2026)
Perplexity's first technical explainer (Aug 6, 2026) details post-training models to connect evidence across sources and verify facts - shaping how content is synthesized and cited.
Sources
Perplexity Runs 6-Stage RAG Pipeline โ Only 5-10 Pages Retrieved Become 3-4 Cited
Perplexity's pipeline (BM25 + dense retriever + XGBoost reranker, 0.7 threshold) retrieves 5-10 pages and cites 3-4. Crawler access is a binary gate: robots.txt/Cloudflare blocks remove pages
Sources
Perplexity Uses 3-Layer Reranking Pipeline With 780M+ Monthly Queries
Perplexity processes 780M+ monthly queries via 3-layer reranking: relevance 30%, visual placement 20%, domain authority 15%, freshness 15%, source diversity 10%, schema 10%.
Sources
ChatGPT Has 'Labrador' VIP Lane for Licensed Publishers Bypassing Normal Retrieval
Researchers discovered ChatGPT's VIP tier via network traffic analysis. Licensed publishers (Reuters, WSJ, Wikipedia) get pre-summarized full-article extracts.
Sources
Rail Europe doubled ChatGPT traffic by prioritizing AI crawler access and value-led SEO
Botify case study: Rail Europe shifted from volume-led to value-led SEO, reorganized sitemaps by page type, and prioritized AI crawler governance โ resulting in 2x ChatGPT referral traffic.
Sources
Search and Training AI Crawlers Are Separate โ Allow Citation Bots, Block Training Only If Deliberate
OAI-SearchBot/PerplexityBot/ClaudeBot power citations; GPTBot/Google-Extended affect training only. Blocking the wrong one invisibilizes you.
Sources
Short URLs do NOT get cited more โ /guide/ pages average 42% above baseline, 'blog/' next
Otterly AI analyzed 1,028,959 URLs cited across 6 AI platforms. URL length correlates near-zero with citations (r = -0.025). /guide/ pages average 2.7 citations (42% above 1.9 avg). Clean URLs beat qu
Sources
Split robots.txt Strategy: Allow Retrieval Bots, Block Training Bots
89% of AI crawler traffic is training/mixed, only 8% search-related. Allow OAI-SearchBot, PerplexityBot, Claude-Web. Block GPTBot, CCBot, Bytespider.
Sources
AI Crawlers Mostly Do Not Execute JavaScript โ Server-Render Citation Content
Vercel analysis of 1.3B fetches found near-zero JS execution for GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot; only Google-Extended renders JS.
Sources
Stripe uses llms.txt to steer AI agents away from deprecated APIs โ agent behavior shaping
Stripe's llms.txt includes 'Instructions for Large Language Model Agents' section. Explicitly steers coding assistants away from legacy Card Element. Real value of llms.txt is agentic-web control
Sources
Stripe uses llms.txt to steer AI agents away from deprecated APIs โ agent behavior shaping
Stripe's llms.txt includes 'Instructions for Large Language Model Agents' section. Explicitly steers coding assistants away from legacy Card Element. Real value of llms.txt is agentic-web control, not
Sources
Structured data + llms.txt deliver cheapest AI visibility ROI โ 200% monthly traffic growth in auto parts case study
Hedges Company case study: Product schema markup + llms.txt drove 200% monthly AI referral traffic growth for 3 months. ChatGPT/Perplexity citation rates tripled month-over-month.
Sources
Structured data + llms.txt deliver cheapest AI visibility ROI โ 200% monthly traffic growth in auto parts case study
Hedges Company case study: Product schema markup + llms.txt drove 200% monthly AI referral traffic growth for 3 months. ChatGPT/Perplexity citation rates tripled month-over-month.