Google Says Do Not Chunk Your Website for AI. Microsoft Says AI Selects Chunks. They Are Both Right.
One of the articles that brings us the most readers argues that the opening of a page, roughly the first 30%, does the heavy lifting in whether AI systems cite it. Then Google published guidance saying websites do not need to break content into special AI-friendly pieces at all. Microsoft, meanwhile, tells publishers that AI assistants select individual sections of a page and that structure improves selection. That sounds like a contradiction, and it also sounds like our own advice is caught in the middle. It is not a contradiction. The two companies are answering different questions, and seeing that changes how you should structure a website, including how we would restate our own most-read claim.
Key Takeaways
- Google’s guidance for its generative AI features says there is no need to break content into small pieces, no ideal page length, and no special AI-specific markup required. It also says files such as llms.txt do not help or harm visibility in Google Search.
- Microsoft’s guidance for inclusion in AI answers, published in October 2025, says systems parse pages into smaller pieces and recommends direct concise answers, descriptive headings, Q&A blocks, lists and tables, and claims anchored in specific, measurable facts.
- These answer different questions. Google is answering an optimisation question: you do not need to artificially restructure content. Microsoft is answering a retrieval question: systems select usable passages from within what they find.
- The resolution is what we call evidence architecture: structure pages into coherent evidence units, not artificial AI chunks. A page should be large enough to prove its claim and modular enough to retrieve.
- Our own earlier advice deserves the same treatment: putting the direct answer early helps people and retrieval alike, but there is no evidence for a magic percentage boundary, and we should not have implied one.
What Does Google Actually Say?
That you should stop performing surgery on your content for machines. Google’s guide to optimising for its generative AI features is unusually blunt for a Google document: there is no need to break content into small pieces for AI systems, no ideal page length, and no AI-specific markup required to appear in its AI experiences, and files such as llms.txt do not help or harm visibility in Google Search.
What Google says matters instead reads like a list from before the AI era, because it is. Pages need to be crawlable and indexable, content should be written for people, and the material that wins is unique and non-commodity: original information, real experience, and a perspective that is not interchangeable with ten other pages.
The warning underneath addresses an obvious temptation: sites shredding themselves into hundreds of fragments, one per imagined AI query, or maintaining separate AI-focused content. Google says creating separate content for every possible search variation is unnecessary, and doing so primarily to manipulate rankings or generative AI responses can violate its scaled content abuse policy. Its systems, it says, can understand multiple topics on a page and surface the relevant part.
What Does Microsoft Actually Say?
That selection happens inside the page, and structure affects it. Microsoft’s guidance on optimising content for inclusion in AI search answers, published in October 2025, describes AI assistants parsing pages into smaller pieces that can be evaluated for authority and relevance. Its February 2026 AI Performance reporting then gave publishers a way to see which pages were actually being cited and the grounding queries associated with them.
Its recommendations follow from that mechanic: give direct, concise answers to real questions, use descriptive headings that say what each section contains, keep titles, headings, and descriptions aligned, use lists, tables, and question-and-answer blocks where the information genuinely fits those shapes, anchor claims in specific, measurable facts, and keep your entity information consistent. None of it is exotic. All of it assumes a machine will be choosing a passage, not reading your page like a novel.
So one company says do not chunk, and the other describes chunk selection as how the system works. Read carelessly, that is a contradiction. Read carefully, it is two views of the same machine from different ends.
Why Are They Both Right?
Because they are answering different questions. Google is answering an optimisation question: do publishers need to artificially restructure content for generative AI? No. Fragmentation produces thin pages, duplicated effort, and a worse site for people, and the retrieval systems do not need it.
Microsoft is answering a retrieval question: how does an AI system select usable information from content it has found? Through structure, clarity, and self-contained sections, because selection operates on passages.
Put together, the instruction is not “chunk” or “do not chunk.” It is: write whole, coherent pages, and make the pieces inside them findable. We call the result evidence architecture, and the working rule is one sentence. A page should be large enough to prove its claim and modular enough to retrieve. Large enough to prove, because trust needs context, evidence, and depth on one URL. Modular enough to retrieve, because the systems citing you may take a section rather than the whole page.
The failure modes sit at the two extremes. Content chopping produces fragments that retrieve but prove nothing. Wall-of-text pages prove their claim but give retrieval nothing clean to take. The architecture lives between them.
What About Our Own First 30% Claim?
It gets the same fact-check we give everyone else. The observation behind it holds up: putting the direct answer early in a page and early in each section helps people, helps snippets, and gives passage-level retrieval a clean unit to select. Answer-first structure remains our standing recommendation, and both companies’ guidance is consistent with it.
The percentage does not hold up, and we should say so plainly. No platform has published evidence that a specific share of a page, 30% or any other number, functions as a citation boundary. Treat the figure as a memorable way of saying front-load your answers, which is what we intend it to mean, and not as a measured threshold, which nobody has demonstrated. Advice that survives its own scrutiny is the only kind worth publishing.
What Does Evidence Architecture Look Like in Practice?
At the site level, it is about giving retrieval one strong candidate instead of many weak ones. Give each important task, entity, or decision one clear URL that owns it. Connect related pages with descriptive internal links and genuinely useful hub pages. Resist creating a thin page for every possible fan-out query, which is the trap we covered in our query fan-out article and runs directly against Google’s guidance that you do not need separate content for every variation of how someone might search. Keep the pages that matter crawlable, indexed, and present in visible HTML, and treat schema as corroboration of what the page visibly says, never a substitute for it.
At the page level, it is about making each evidence unit complete. Align the title, main heading, and description around one clear promise. Give the direct answer early. Organise the supporting questions under descriptive headings, and place the evidence beside the claim it supports: examples, dates, measurements, named sources, first-hand experience. Use tables and lists when the information is genuinely comparative or procedural, not as decoration. End with the reader’s next logical step and a contextual link to it.
And at the interaction level, there is now a third reader to serve. Agents that act on pages depend on the same things assistive technology always has: real buttons and links with descriptive accessible names, properly associated form labels, and stable layouts. Google’s agent-friendly website guidance frames this directly: the semantic structure and accessibility work that serves people is what agents operate on. Nothing about serving the third reader requires abandoning the first two.
How Do You Know If Any of This Is Working?
You stop guessing and read the citation data, because for the first time the platforms are showing it. Microsoft’s AI Performance report in Bing Webmaster Tools, in public preview since February 2026 and expanded in June with Intents, Topics, Citation Share and Compare, shows which URLs are cited, the grounding queries associated with them, and how citation activity changes over time. Google’s generative AI performance reporting in Search Console began rolling out to a subset of sites in June 2026. It shows impressions and which URLs appeared in AI Overviews and AI Mode, but does not currently provide Microsoft’s level of citation or grounding-query detail.
That data answers the question this whole article raises better than any formatting theory can: which of your pages, structured which way, actually get selected. Our benchmark article covers how to turn those reports into a weekly measurement you can keep. The honest loop is structure, publish, measure, adjust, and no step in it requires believing anyone’s magic number, including ours.
Frequently Asked Questions
Q: Should I break my website content into chunks for AI?
A: No. Google’s guidance explicitly says there is no need to break content into small pieces for its AI features, and warns against creating separate content for every possible query variation. Write coherent pages, and use clear sections and headings so systems can select the relevant part.
Q: Does page structure affect whether AI cites my content?
A: Microsoft’s guidance says AI systems parse pages into smaller pieces and recommends direct answers, descriptive headings, Q&A blocks, lists, tables, and claims anchored in specific, measurable facts to improve selection. Google recommends clear structure for usability without promising it causes inclusion, so treat structure as sound practice rather than a guaranteed lever.
Q: Is there an ideal page length for AI search?
A: No. Google states there is no ideal page length for its generative AI features. The practical rule is that a page should be long enough to prove its claim with evidence and structured enough that each section can stand alone when retrieved.
Q: Does the first part of a page matter more for AI citation?
A: Placing direct answers early helps people, snippets, and passage-level retrieval, and both Google’s and Microsoft’s guidance is consistent with answer-first structure. No platform has published evidence of a specific percentage boundary, so front-load answers without treating any number as a measured threshold.
Q: How can I measure whether AI systems cite my pages?
A: Microsoft’s AI Performance report in Bing Webmaster Tools, in public preview since February 2026, shows cited URLs, grounding queries, and citation trends for Copilot and Bing AI answers. Google has begun rolling out generative AI performance reporting in Search Console. Microsoft Clarity’s AI Visibility reporting also provides citation data, cited pages, grounding queries, share of authority and AI referral traffic, with branded and non-branded grounding-query segmentation added in August 2026.
Two companies looked at the same machine and gave advice from opposite ends of it, and the industry heard a contradiction because a contradiction is more shareable than a synthesis. The synthesis is better: build pages that prove and sections that travel. That was true before anyone called it AI optimisation, and it will still be true when the acronyms change again.
AI Visibility Studio helps websites structure content so AI systems can find it, understand it, cite it, and actually use it when generating answers. aivisibilitystudio.com
References
- Google Search Central, “Optimizing Your Website for Generative AI Features on Google Search”: developers.google.com/search/docs/fundamentals/ai-optimization-guide
- Google Search Central, “AI Features and Your Website”: developers.google.com/search/docs/appearance/ai-features
- Google web.dev, “Build agent-friendly websites” (updated April 2026): web.dev/articles/ai-agent-site-ux
- Microsoft Bing Blog, “Introducing AI Performance in Bing Webmaster Tools Public Preview” (February 2026): blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview
- Microsoft Bing, “Optimizing Your Content for Inclusion in AI Search Answers” (October 8, 2025): about.ads.microsoft.com/en/blog/post/october-2025/optimizing-your-content-for-inclusion-in-ai-search-answers
- AI Visibility Studio, “You Answered the Question. The AI Asked Five More.” and “How to Build an AI Citation Benchmark” (internal): aivisibilitystudio.com/blog
Originally published on Medium ↗