How AI search decides what to cite (a worked example)
I built a Wikipedia-style page with 156 dated, sourced facts to test what AI answer engines cite. Here's the structure, and what it means for local sites.
Last week I published a page with 156 facts on it. Each one has a source URL and the date I checked it. It took two research passes, one in Claude Code and one in ChatGPT, cross-checked against each other, before a single sentence got written.
That's a stupid amount of work for a page about SEO courses. I did it because I wanted to know, with my own site, what AI answer engines actually reward. Not what the conference talks say. What happens when you build the page and watch.
The page is aiseocourse.wiki. It's a reference article on AI SEO courses, styled like Wikipedia, with 43 footnotes. It's also my entry in a 30-day ranking contest, which is a separate story. This post is about the structure, because the structure is the part that transfers to a plumber's website in Klang.
Why Wikipedia's shape, not a blog post
When ChatGPT, Perplexity or Google's AI Mode answer a question, they pull from pages they can extract a clean fact from. Wikipedia gets cited more than any other site not because it's Wikipedia, but because every Wikipedia page is a machine for extraction: a definition in the first sentence, an infobox with label-value pairs, numbered sources, neutral voice, no sales pitch to filter out.
So I copied the shape. The model was charlesfloate.wiki, a page Charles Floate built for himself in August with the same skeleton. Infobox. Contents box. Numbered references. "Last edited on" line at the bottom. It costs nothing to use that layout on a static site and it removes every reason for a retrieval system to skip you.
There's research on this now. A March 2026 study led by the University of Tokyo with Tsukuba, Hiroshima and the National Institute of Informatics, published as GEO-SFE, found that structural changes alone, with no change to the content itself, lifted citation rates 17.3% across six generative engines. Same words, better shape, more citations.
A .wiki domain, by the way, does nothing. Google gives no credit to the extension. The shape is what matters. You could do this on a .com or on a subfolder of your existing site.
The four things I'd copy onto any local site
1. A machine summary at the top
Two or three sentences right under the title, bold lead, every fact footnoted. Name, what it is, where, since when, the one number that matters. This is the block an AI engine lifts when it's in a hurry, which is always. A competitor site launched the same day as mine with this exact block, and it was the smartest thing on their page.
2. Dated facts
Every number on my page says when it was checked. "$289 a month on 17 September 2026." Not "$289 a month." Engines are increasingly biased toward freshness, and a date on a fact is a freshness signal a scraper can read without understanding anything else. For a local business that means: opening hours with a "checked" date, prices with a date, "serving Subang Jaya since 2014" rather than "over a decade of experience".
3. A sources page and a data file
My site has a /sources page listing every citation with a method note, and a /data.json file with all 156 facts as structured records. Also an llms.txt at the root pointing to both. I don't think any of this helps rankings today. I think it will within a year, and it costs an hour. LocusPilot does the machine-readable half of this automatically on the sites it builds: llms.txt and llms-full.txt, a clean markdown version of every page, and the .well-known files agents look for. The sources page and the data file I still put together by hand.
4. Neutral voice, criticism included
The Rainmakers article on my site has a criticism section. Four subsections, sourced. My first draft had nine, and it read like an investigation, so I cut it. But I kept it, because a page with zero downside is a sales page, and retrieval systems have gotten good at sorting sales pages from reference pages. On a local site this is the "what we don't do" paragraph. The aircon company that says "we don't service VRF systems, call these guys instead" is more citable than the one that claims everything.
What I'd skip
Schema stuffing. I put Article, Organization and BreadcrumbList markup on the pages, matching what's visible, and nothing else. I removed an AggregateRating block after launch because the rating came from Skool, not from my own users, and Google's rules are clear on that. Fake review markup is the fastest way to lose the little trust a new site has.
Interlinking a pile of new domains. I bought seven domain variations and built two. Two fresh domains linking to each other pass roughly nothing. A single link from an old, indexed page on my main site did more than any of that.
Waiting to be found. New domains take one to four weeks to index on their own. I submitted everything to Search Console and Bing on day one, pinged IndexNow, and ran a paid speed indexer on every URL. Your local site is probably already indexed, so you can skip this, but if you're launching a new location page, don't just publish and hope.
The honest part
I don't know yet whether the page will get cited. It's been live eleven days. The contest scoring window is 14 to 16 October and I'll publish whatever the numbers say, good or bad. What I do know is that building it this way forced every claim through a source, and that alone made it a better page than anything I'd have written from memory.
If you want to see the structure, the page is aiseocourse.wiki. The companion, aiseocourse.directory, does the same thing as comparison cards with check dates, and a free eight-lesson version runs the same structure for people starting out. The same shape carries to a hosted page too, 45 programs, every check sourced. Copy the shape. Put your own facts in it.
Ready to build a rank-ready site?
Get 30+ pages, auto-applied images, and full SEO setup in minutes.
Get started

