Getting found
Analytics, SEO and growth
Once a site is live, two questions follow: how many people visit and what do they do there, and how do new people find it in the first place. Analytics tools count visits and actions, ad tracking reports sales back to ad platforms, and SEO is the work of being found through search engines and, increasingly, through AI assistants.
The pieces stack up. A script on each page collects visits and events; small files such as sitemap.xml and robots.txt tell crawlers what exists and where they may go; structured data and social tags describe each page; and speed, measured by Core Web Vitals, affects both rankings and whether visitors stay.
20 terms · 4 comparisons · prices checked September 2026
20 terms, click any to open
Analytics
Web analytics
ConceptMeasuring how many people visit a website or app, where they come from and what they do there, so decisions rest on numbers rather than guesses.
Most tools work through a small script on every page. Each visit records a page view with details such as the referring site, campaign tags, country, device and browser, plus custom events such as sign-ups, button clicks and purchases. Dashboards turn that into visitors, sessions, top pages, traffic sources, conversion rates and funnels, which show how many people get through each step of, say, a checkout.
The main families are Google Analytics 4 (free, detailed and tied to Google Ads), cookieless tools such as Plausible, Fathom and Umami (simple and privacy-first), and product analytics such as PostHog, Mixpanel and Amplitude, which follow individual users through an app. Server logs and a CDN's own statistics count requests without any script at all.
No tool sees everyone. Ad blockers, declined cookie banners and bots all skew the numbers, so two tools rarely agree exactly. Trends over weeks and months say more than any single figure.
Related
Also called: site analytics, traffic stats, page views, product analytics
Google Analytics 4
ServiceFreeGoogle's free analytics service, which shows how many people use a website or app, where they came from and what they did, and links the results to Google Ads.
A Google tag (gtag.js, or a tag in Google Tag Manager) on each page sends events to a GA4 property. Page views, scrolls, outbound clicks, site searches and file downloads are collected automatically, and you add your own, such as sign_up, or purchase with a value. The events that matter most are marked as key events (formerly called conversions). Standard reports cover traffic, engagement and revenue; Explorations build funnels, paths and segments; and a free BigQuery export hands over the raw events.
GA4 replaced Universal Analytics, which stopped collecting data in 2023, and its event-based model takes some learning. The free tier samples very large Explorations, hides small numbers when some privacy features are on, and keeps user-level data for Explorations for 14 months at most. Analytics 360, the enterprise tier, raises those limits and adds service agreements.
It stores first-party cookies (_ga) and sends data to Google, so in the EU, the UK and similar places it normally needs consent through a cookie banner. Consent Mode tells Google tags what each visitor agreed to, and Google models some of the missing data. Ad blockers stop it for part of the audience.
Pros
- Free, with generous limits for almost any site
- Tight links with Google Ads, Search Console and BigQuery
- Covers websites and mobile apps in one property
- Funnels, audiences and explorations for detailed questions
- Widely known, so agencies, plugins and tutorials are easy to find
Cons
- Steep learning curve, and the interface changes often
- Needs cookie consent in the EU and UK, which leaves gaps in the data
- Blocked by many ad blockers and privacy-focused browsers
- Visitor data goes to Google, which some organisations must avoid
Pick it when
- You run Google Ads and want conversions to flow back into campaigns
- You need funnels, audiences or raw event data in BigQuery
- A marketing team or agency already works in it
Skip it when
- You only need visitors, sources and top pages (a cookieless tool is simpler)
- You want to avoid a consent banner or sending visitor data to Google
What it costs · Free
Free for nearly all sites. Analytics 360, with higher limits and service agreements, is sold on custom contracts through Google and its partners; there is no public price.
Approximate, checked September 2026.
Related
Also called: GA4, Google Analytics, GA, Analytics 360, gtag.js
Open Google Analytics 4 as a pageOfficial site (opens in a new tab)
Google Tag Manager
ServiceFreeA free Google tool for adding and managing tracking scripts, such as analytics, ad pixels and heatmaps, from a web dashboard instead of changing the site's code each time.
You install one container snippet on the site. In the GTM dashboard you then define tags (the scripts to run, such as GA4, Google Ads, the Meta Pixel or LinkedIn's Insight Tag), triggers (when to run them: page views, clicks, form submissions, custom events) and variables. The site's own code can push events and details, such as an order total, into a data layer that tags read. Changes are published as versions, with a preview mode for testing first.
It frees marketers from waiting on developers, but every tag adds third-party JavaScript that can slow pages, leak data or break things, so access and reviews matter. It works with Consent Mode and cookie banners to hold tags back until a visitor agrees. Server-side tagging runs a container on a server you control, which receives events from the site and forwards them to each vendor, so fewer scripts load in the browser and you decide what is shared.
What it costs · Free
Free. Tag Manager 360 is sold to large companies through Google. Server-side containers need hosting, about $45 a month per server on Google Cloud Run.
Approximate, checked September 2026.
Related
Also called: GTM, Tag Manager, server-side GTM, sGTM, data layer
Open Google Tag Manager as a pageOfficial site (opens in a new tab)
UTM parameters
ConceptLabels added to the end of a link, such as ?utm_source=newsletter, that tell your analytics exactly which campaign, email or post sent a visitor.
There are five classic tags: utm_source (who sent the visit, such as google or newsletter), utm_medium (the channel, such as cpc, email or social), utm_campaign (the campaign name), and the optional utm_term and utm_content for keywords and ad variants. Analytics tools read them from the landing page's address and credit the visit, and any later conversion, to that campaign. The name comes from Urchin, the analytics company Google bought and turned into Google Analytics.
Consistent naming matters: GA4 treats Email and email as different sources, so teams keep a shared naming sheet or a link builder. Tagging internal links on your own site is a mistake, because it overwrites where the visitor originally came from. Ad platforms add their own click ids automatically (gclid for Google Ads, fbclid for Meta), and canonical tags stop tagged URLs from showing up as duplicate pages in search.
Related
Also called: UTM tags, utm_source, utm_campaign, campaign tracking, tracking links
Ad tracking
Meta Pixel
ServiceFreeA snippet of JavaScript from Meta that reports what visitors do on your site, such as viewing a product or buying it, back to Facebook and Instagram ads.
The base code loads Meta's script and sends a PageView on every page; you add standard events such as ViewContent, AddToCart, Lead and Purchase (with value and currency), or custom events of your own. It sets a first-party _fbp cookie and keeps the fbclid id from ad clicks, so Meta can match events to its users, credit ads with the conversions they drove, build audiences such as people who viewed a product, and steer ad delivery towards people likely to act.
Because it runs in the browser, ad blockers, Safari and Firefox tracking protection and visitors who decline cookies all cut into the data, and under GDPR and similar laws it should only load after consent. Meta recommends pairing it with the Conversions API and deduplicating the two with a shared event id. Events Manager shows what arrives, and the Meta Pixel Helper browser extension helps debug a setup.
Pros
- Free, and quick to add by snippet, Tag Manager or a platform plugin
- Lets Meta optimise ads towards real purchases or leads
- Builds retargeting and lookalike audiences from site visitors
- Ready-made integrations for Shopify, WordPress and most site builders
Cons
- Blocked by many ad blockers and privacy-focused browsers
- Needs cookie consent in the EU, UK and similar regions
- Adds third-party JavaScript that can slow pages
- Shares visitor behaviour with Meta, which your privacy policy must disclose
Pick it when
- You run, or plan to run, Facebook or Instagram ads
- You want retargeting audiences built from site visitors
- Alongside the Conversions API, for more complete data
Skip it when
- You do not advertise on Meta's platforms
- Sensitive sites, such as health services, where sharing visit data with an ad network is risky
What it costs · Free
Free to use; you pay only for the ads it measures.
Approximate, checked September 2026.
Related
Also called: Facebook Pixel, fbq, Meta tracking pixel, Facebook conversion tracking
Conversions API (CAPI)
APIFreeMeta's server-to-server way of reporting conversions: your server tells Meta about a sign-up or purchase directly, so it is not lost to ad blockers or browser limits.
When something that matters happens, such as an order being paid, your backend sends an HTTPS request to Meta's Graph API with an access token, the event name and time, the page URL and customer details. Email addresses and phone numbers are normalised and hashed with SHA-256 before sending; the _fbp and _fbc cookie values, IP address and browser user agent go along unhashed to help matching. Events Manager grades the result with an Event Match Quality score.
Because it runs on your server, it also reports events a browser never sees: a payment confirmed later by a webhook, a phone order, an in-store sale or a lead that becomes a customer in your CRM. When the Pixel and the API both send the same event, giving both the same event name and event id lets Meta keep one copy, as long as they arrive within 48 hours of each other.
Setup can be direct code, a platform integration (Shopify, WooCommerce), server-side Google Tag Manager or Meta's Conversions API Gateway. TikTok, Google Ads, LinkedIn and Snap offer similar server-side APIs.
Pros
- Not affected by ad blockers or browser cookie limits
- Reports events after the page has closed, such as webhook-confirmed payments
- Better matching when hashed email addresses or phone numbers are included
- You control exactly which data leaves your server
Cons
- Needs backend work, an access token and careful hashing
- Without deduplication, conversions get counted twice
- Consent rules still apply, so server-side is not a loophole
Pick it when
- Meta ads are a meaningful part of your marketing spend
- Purchases or leads are confirmed on the server, for example by a payment webhook
- Pixel data is visibly missing conversions
Skip it when
- You do not advertise on Meta's platforms
- A small test campaign where the Pixel alone is enough
What it costs · Free
Free to use. Costs come only from the server, gateway or partner tool that sends the events.
Approximate, checked September 2026.
Related
Also called: CAPI, Meta Conversions API, Facebook Conversions API, server-side tracking
Open Conversions API (CAPI) as a pageOfficial site (opens in a new tab)
Getting found
SEO
ConceptSearch engine optimisation: making pages easy for Google, Bing and other search engines to find, understand and trust, so they appear when people search.
Search engines work in three steps. They crawl the web by following links and sitemaps, index what they find (store it and work out what each page is about), and rank indexed pages for each search using many signals: relevance to the query, content quality, links from other sites, freshness, location and page experience. SEO is work on each of those steps.
Technical SEO keeps a site crawlable and fast: clean URLs, HTML that arrives already rendered, a sitemap and robots.txt, canonical tags, structured data, mobile-friendly layouts and good Core Web Vitals. On-page SEO is the content itself: pages that answer what people actually search for, with clear titles, meta descriptions, headings and internal links. Off-page SEO is reputation: links and mentions from other sites, reviews and, for local businesses, a Google Business Profile.
Results take weeks to months, nobody can guarantee a ranking, and shortcuts such as bought links or mass-produced pages break Google's spam policies. AI answers draw on the same search index, so the same groundwork also supports answer engine optimisation.
Pros
- Visitors keep arriving without paying for each click
- People arrive already looking for what you offer
- Good pages can bring visitors for a long time after publishing
- Technical fixes also make the site faster and easier to use
- Also the groundwork for being cited in AI answers
Cons
- Slow: results usually take months
- No guarantees, and algorithm updates can shift rankings quickly
- Competitive topics need steady content and reputation work
- Hard to tie revenue to specific pages or efforts
Pick it when
- People already search for what you sell or explain
- You can publish useful content steadily over months
- A local business that wants to show up in map and 'near me' results
Skip it when
- You need customers this week (ads work faster)
- Nobody searches for the product yet, such as a brand-new category
Related
Also called: search engine optimisation, search engine optimization, organic search, technical SEO
Answer engine optimisation
ConceptShaping content so that AI assistants such as ChatGPT, Perplexity and Google's AI answers cite it, or mention your brand, when people ask them questions.
AI answer engines usually run a web search, read the top results and write a summary with links to the pages they used. Being one of those sources rests mostly on the same things as SEO: pages that can be crawled and indexed, and content that is clear, specific and trustworthy. Google states that optimising for its AI Overviews and AI Mode is still SEO, and that no special files or markup are needed.
Habits that help: answer the question directly near the top of the page, back claims with specific facts and sources, keep pages up to date, earn mentions on other reputable sites, and serve content as HTML rather than only through JavaScript, since many AI crawlers do not run scripts. Search crawlers such as OAI-SearchBot (for ChatGPT) and PerplexityBot need to be allowed in robots.txt for a site to appear in those tools.
Measuring it is still rough. Search Console shows impressions in Google's AI features, and analytics tools show referrals from sites such as chatgpt.com and perplexity.ai. Promises of guaranteed placement in AI answers deserve caution.
Related
Also called: AEO, GEO, generative engine optimisation, AI SEO, LLM SEO
Open Answer engine optimisation as a pageOfficial site (opens in a new tab)
Google Search Console
ServiceFreeGoogle's free dashboard for site owners, showing which searches bring people to the site, which pages are indexed and what problems Google has found.
After you prove you own a site (with a DNS record, an HTML file, a meta tag, or through Analytics or Tag Manager), the Performance report lists search queries with clicks, impressions, click-through rate and average position for the last 16 months. The Page indexing report explains why pages are or are not in Google's index, URL Inspection shows how Google sees a single page and lets you request indexing, and sitemaps are submitted here too.
Other reports cover Core Web Vitals from real Chrome users, structured data errors, manual penalties, security problems and links. Newer reports show impressions in generative AI features such as AI Overviews and AI Mode. Bing Webmaster Tools is the equivalent for Bing and can import a Search Console setup.
What it costs · Free
Free for any site you can verify.
Approximate, checked September 2026.
Also called: GSC, Search Console, Google Webmaster Tools
Open Google Search Console as a pageOfficial site (opens in a new tab)
Structured data (JSON-LD)
FormatMachine-readable labels in a page's code that state exactly what the page describes, such as a product's price or an event's date, in a shared vocabulary.
The vocabulary is schema.org, backed by Google, Microsoft, Yahoo and Yandex, with types such as Organization, Product, Article, Event, Recipe, LocalBusiness and BreadcrumbList. JSON-LD places it in a script tag of type application/ld+json as a block of JSON, separate from the visible HTML, and is the format Google recommends; Microdata and RDFa weave the same information into HTML attributes instead.
Valid markup makes a page eligible for rich results such as star ratings, prices, event dates and breadcrumbs, and helps search engines and AI systems read facts reliably. It is not a ranking boost on its own, it must match what visitors can see, and Google has retired several rich result types over the years (How-to, then FAQ), so check the current list. Google's Rich Results Test and the Schema Markup Validator check the code.
Related
Also called: JSON-LD, schema.org, schema markup, rich results, rich snippets
Open Structured data (JSON-LD) as a pageOfficial site (opens in a new tab)
Open Graph and social cards
FormatTags in a page's head that set the title, description and image shown when someone shares the link on WhatsApp, LinkedIn, Facebook, Slack or X.
Facebook introduced the Open Graph protocol in 2010, and nearly every app that builds link previews now reads it. The core tags are og:title, og:description, og:image, og:url and og:type, written as meta tags in the page's head. X reads its own twitter:card tags and falls back to Open Graph for the rest. An image of 1200 by 630 pixels displays well on most platforms.
Previews are fetched once and cached, so an updated image may not appear until the cache is refreshed with tools such as Facebook's Sharing Debugger or LinkedIn's Post Inspector. Image URLs must be absolute and publicly reachable. Frameworks such as Next.js can generate a card image for every page automatically.
Also called: og:image, OG tags, Twitter cards, link previews, social share image
Open Open Graph and social cards as a pageOfficial site (opens in a new tab)
Canonical URL
ConceptA tag that tells search engines which address is the main one when the same page can be reached at several URLs.
The same content often lives at several addresses: with and without www, over http and https, with tracking parameters such as ?utm_source=, with sort or filter options, or as a copy republished on another site. A link element with rel='canonical' in the page's head (or a Link HTTP header, for files such as PDFs) names the preferred version, so ranking signals gather on that URL and it is the one shown in results.
Google treats it as a strong hint rather than a command, weighing it alongside redirects, sitemaps and internal links, and URL Inspection in Search Console shows which version it chose. Use absolute URLs, point an original page at itself, and use a permanent (301) redirect instead when the duplicate does not need to exist at all.
Related
Also called: rel=canonical, canonical tag, duplicate content
Open Canonical URL as a pageOfficial site (opens in a new tab)
Telling crawlers
XML sitemap
FormatA file listing the pages of a site, with when each last changed, so search engines can find them without relying only on links.
A sitemap is an XML file, usually at /sitemap.xml, holding a url entry for each page with its address (loc) and, ideally, its last modified date (lastmod). The format, published at sitemaps.org, is shared by Google, Bing and most other engines. One file can hold up to 50,000 URLs or 50 MB uncompressed; bigger sites split into several files listed in a sitemap index, and extensions cover images, video and news.
Submit it in Search Console and Bing Webmaster Tools, or point to it with a Sitemap line in robots.txt. It helps pages get discovered but does not force them into the index. Google ignores the priority and changefreq fields and trusts lastmod only when it is consistently accurate. Most frameworks and CMSs generate one automatically.
Also called: sitemap.xml, sitemap index, sitemaps
Open XML sitemap as a pageOfficial site (opens in a new tab)
robots.txt
FormatA small text file at the root of a site that tells search engine and AI crawlers which parts they may visit and which to leave alone.
It lives at /robots.txt and groups rules by crawler: a User-agent line (Googlebot, Bingbot, GPTBot, or * for everyone) followed by Allow and Disallow lines with paths, plus an optional Sitemap line. The format was standardised as RFC 9309 in 2022. Reputable crawlers obey it, but it is a request, not a lock: it offers no security, and anyone can read the file.
It controls crawling, not indexing. A blocked page can still appear in results, without a description, if other sites link to it; to keep a page out of search, let it be crawled and add a noindex robots meta tag or X-Robots-Tag header. AI companies often use separate crawlers for training and for search (OpenAI has GPTBot and OAI-SearchBot), so each can be allowed or blocked on its own.
Also called: robots exclusion protocol, crawl rules, Disallow, User-agent
IndexNow
ProtocolFreeA free way to tell search engines the moment a page is added, changed or deleted, instead of waiting for them to crawl again. Bing and several others support it; Google does not.
You generate a key, host it as a text file on your site to prove ownership, and then send each changed URL, or up to 10,000 of them in one JSON request, to an IndexNow endpoint. Participating engines share submissions with each other, so one ping reaches Bing, Yandex, Naver, Seznam.cz, Yep and the rest. A success response only means the URL was received, not that it will be indexed.
Microsoft Bing and Yandex launched it in 2021. Google does not take part, so sitemaps and Search Console remain the way to reach it. Many CMS plugins, hosting platforms and CDNs (Cloudflare's Crawler Hints, for example) send pings automatically, and it suits sites whose pages change often, such as shops, listings and news.
What it costs · Free
Free to use; there is nothing to buy.
Approximate, checked September 2026.
Also called: Index Now, IndexNow API, instant indexing
llms.txt
FormatA proposed Markdown file at /llms.txt that gives AI tools a short, curated guide to a site, with links to its most useful pages. It is a community proposal, not an official standard.
Jeremy Howard of Answer.AI proposed it in 2024. The file opens with the site's name as a heading and a one-line summary, followed by sections of links with short notes, ideally pointing to clean Markdown versions of the pages. It targets AI agents and coding assistants that need concise context rather than pages full of navigation, adverts and scripts. Some sites also publish an llms-full.txt with the complete text.
Support is uneven. It is most common on software documentation, including the developer docs of several AI companies, and Chrome's Lighthouse checks for it in its Agentic Browsing audits, but Google says its Search, AI Overviews included, ignores the file. It does not control access either; that remains the job of robots.txt.
Related
Also called: llms-full.txt, LLMs.txt, AI text file
Speed
Core Web Vitals
ConceptGoogle's three measures of how a page feels to real visitors: how quickly the main content appears, how fast it responds to taps and clicks, and how much it jumps around.
Largest Contentful Paint (LCP) is the time until the biggest image or block of text appears; good is 2.5 seconds or less. Interaction to Next Paint (INP), which replaced First Input Delay in 2024, measures how long the page takes to respond visibly to taps, clicks and key presses; good is 200 milliseconds or less. Cumulative Layout Shift (CLS) scores how much content moves unexpectedly as the page loads; good is 0.1 or less.
A page passes when at least 75 per cent of real visits meet all three targets, measured from Chrome users over the previous 28 days in the Chrome UX Report. That field data appears in Search Console and PageSpeed Insights, while lab tools such as Lighthouse simulate one load to help diagnose problems. The scores feed Google's page experience signals, which weigh less than relevance, but slow, jumpy pages lose visitors either way.
Related
Also called: CWV, LCP, INP, CLS, page experience
Open Core Web Vitals as a pageOfficial site (opens in a new tab)
Lighthouse
ToolOpen sourceGoogle's free, open-source tool that loads a page and scores it for performance, accessibility, best practices and SEO, with a list of what to fix.
Lighthouse runs in Chrome DevTools, from the command line, as a Node module, in CI pipelines through Lighthouse CI, and behind PageSpeed Insights. It loads the page under controlled lab conditions, by default a simulated mid-range phone on a slow mobile connection, scores each category from 0 to 100, and lists audits such as oversized images, render-blocking scripts or missing alt text, each with an explanation.
Lab scores vary between runs and machines, and they are not what Google uses in search; that is field data from real Chrome users (Core Web Vitals). Treat the score as a diagnostic rather than a goal. Recent versions add an Agentic Browsing category that checks how easily AI agents can read and use a page, including a check for llms.txt.
What it costs · Open source
Free (Apache 2.0).
Approximate, checked September 2026.
Related
Also called: Google Lighthouse, PageSpeed Insights, Lighthouse CI, Lighthouse score
Image optimisation
ConceptServing every image as a small, modern file at the right size for the visitor's screen, which is one of the simplest ways to make a site load faster.
Images are usually the heaviest part of a page. The main techniques are modern formats (WebP and AVIF files are typically much smaller than JPEG or PNG at similar quality), resizing to the displayed size with srcset and sizes so phones never download desktop-sized images, sensible compression, lazy loading for images below the fold with loading='lazy', and SVG for logos and icons.
Setting width and height, or an aspect ratio, reserves space and prevents layout shift. The main hero image is often the page's LCP element, so it should load early rather than lazily, often with fetchpriority='high'. Much of this can be automated: next/image in Next.js, Astro's image component, image CDNs such as Cloudinary, imgix or Cloudflare Images, and libraries such as sharp in a build step.
Related
Also called: WebP, AVIF, responsive images, srcset, lazy loading, next/image
Side by side
Differences
How the options in this area compare on the questions that usually decide the choice.
Meta Pixel vs Conversions API
Open as a page: Meta Pixel vs Conversions APIThe Pixel reports from the visitor's browser; the Conversions API reports from your server. Meta recommends running both and deduplicating, because each catches events the other misses.
| Compare | Meta Pixel | Conversions API |
|---|---|---|
| Runs in | The visitor's browser | Your server |
| Setup | A snippet, Tag Manager or a platform plugin | Backend code, a platform integration or a gateway |
| Ad blockers | Often blocked | Not affected |
| Browser privacy limits | Reduced by browser tracking protection | Largely avoided, since events come from your server |
| Events it can send | Only what happens on the page | Any event, including offline and webhook-confirmed sales |
| Matching data | Cookies, plus optional hashed details | Hashed email or phone, IP, user agent, cookie ids |
| Consent | Needed where the law requires it | Same rules; server-side is not a loophole |
| Cost | Free | Free; you pay for the server or tool |
How to choose
- Use the Pixel for a quick start and for browsing events such as page views and add-to-cart.
- Add the Conversions API when sales or leads are confirmed on your server, or when Pixel data is missing conversions.
- Running both with a shared event id for deduplication is the setup Meta recommends.
SEO vs answer engine optimisation
Open as a page: SEO vs answer engine optimisationBoth are about being found. SEO aims for a place in search results; answer engine optimisation aims to be the source an AI assistant quotes. Most of the groundwork is shared.
| Compare | SEO | Answer engine optimisation |
|---|---|---|
| Goal | Rank in search results and earn clicks | Be cited or mentioned in AI answers |
| Where it shows | Google, Bing and other search results | ChatGPT, Perplexity, Google AI Overviews and AI Mode |
| What counts | Relevance, quality, links, page experience | The same, plus direct answers and mentions elsewhere |
| Success looks like | Rankings, clicks and organic traffic | Citations, brand mentions and referral visits |
| Measuring it | Search Console and analytics | Still rough: AI reports, referrals, manual checks |
| Maturity | Long-established and well documented | New, with many untested claims |
| Special files | Sitemap, robots.txt, structured data | None required; llms.txt is optional |
How to choose
- Start with SEO, because AI answer engines mostly draw on pages that search engines already index and trust.
- Add answer-focused habits, such as direct answers, specific facts and up-to-date pages, once the basics are in place.
- Treat promises of guaranteed placement in AI answers with caution.
sitemap.xml vs robots.txt vs IndexNow vs llms.txt
Open as a page: sitemap.xml vs robots.txt vs IndexNow vs llms.txtFour ways of talking to crawlers that are easy to confuse. Each does a different job, and none of them forces a page into or out of a search index.
| Compare | sitemap.xml | robots.txt | IndexNow | llms.txt |
|---|---|---|---|---|
| What it does | Lists the pages that exist | Sets where crawlers may go | Announces changed pages | Summarises a site for AI tools |
| Format | XML | Plain-text rules | An HTTP request plus a key file | Markdown |
| Lives at | /sitemap.xml, or any submitted path | /robots.txt only | Key file on your site; pings go to an API | /llms.txt |
| Who reads it | Google, Bing and most engines | Nearly all reputable crawlers | Bing, Yandex and partners; not Google | Some AI tools and agents; not Google Search |
| Status | Long-established standard | Standard (RFC 9309) | Open protocol from Bing and Yandex | Community proposal |
| Controls indexing | No, it only helps discovery | No, it controls crawling | No, it only notifies | No |
| Effort | Usually generated by the framework | A few lines, rarely changed | A plugin or a small script | Written by hand or generated |
How to choose
- Every public site benefits from a sitemap and a robots.txt; both take minutes, and most frameworks generate them.
- Add IndexNow when quick updates in Bing and its partners matter, as for shops, listings or news.
- Add llms.txt if your audience uses AI coding tools or agents, knowing that Google Search ignores it.
Crafted in the dark. Shipped to the world.
Tell us what you are building. You get a private project space with a proposal and a line-by-line quote within a day.