FREE COURSE · OPEN ACCESS · ALL FREE COURSES
Be found and quoted by AI
BEFORE YOU START · 00 · 6 min
How to use this
Two promises and one warning, so you know exactly what you bought.
What this course can and cannot do
It cannot make a model mention you. Nobody can promise that, and anyone who does is selling a horoscope with a price tag. Retrieval and citation are decided inside systems that are not documented, change without notice, and behave differently for the same question asked twice.
What it does is remove the reasons a model cannot use you. It cannot read your page. It cannot tell who wrote it. It cannot find a claim stable enough to attribute to you. Those three failures are mechanical, they are measurable, and almost every site fails at least one. Fixing them is entirely within your control, and it is the whole job.
Think of it the way you would think of a shop. This course does not bring customers through the door. It unlocks the door, turns the lights on, and puts a legible sign above it. What you sell inside is your problem.
The shape of the three hours
Nine lessons in three parts. Part one is diagnosis: what these systems actually read, and why your site may be invisible to them. Part two is repair, and it is where the generators live — you will leave with your structured data, your llms.txt and your robots.txt written, not described. Part three is proof: how to tell whether any of it worked without fooling yourself.
Every lesson ends with something you do rather than something you read. Your work is saved in this browser as you go, so you can close the tab and come back. The progress bar at the top counts lessons you have marked complete.
Run the audit first
Before lesson one, audit your own domain. The score does not matter. The list of failing items is your syllabus: it tells you which lessons to read slowly and which to skim. Run it again after lesson nine and keep both screenshots — that pair is the most persuasive thing you will ever show a client or a manager on this subject.
Throughout the course, claims come in two kinds. Mechanically verifiable ones — your page contains valid structured data or it does not — are stated flatly. Inferred ones, about how these systems weigh what they find, are marked with a box like this one. There are five of them. I would rather you trust the four hundred verifiable statements because you can see me flagging the five uncertain ones.
PART ONE — DIAGNOSIS · 01 · 13 min
What an answer engine actually reads
Almost every piece of advice about ranking in AI is wrong because it treats one system as three, or three as one.
The three systems, and why the difference is money
When someone asks a model about your field, three separate machines are involved. They have different update speeds, different appetites, and only one of them is worth your attention.
1. The training corpus
Text scraped at some point in the past and compressed into the model's weights. The important properties: it is frozen, it is old, and it is enormous. Anything you publish today will not appear in a model that finished training last year, and by the time it might appear, two more model generations will have shipped.
People spend real budgets trying to influence this. It is the SEO equivalent of trying to change what is printed in an encyclopedia already sitting on the shelf. Ignore it as a lever. It matters only in one narrow way, covered in lesson eight: if you were already well represented across the open web years ago, you get a small permanent advantage that newcomers cannot buy.
2. The retrieval index
When a model answers a current question, it usually runs a search first and reads a handful of results before writing anything. That search hits a live index, updated continuously. This is the part you can actually move, and it moves in days rather than years.
Which index depends on the product. Some license an existing web index, some operate their own crawler, some call a third-party search API, and several use different sources depending on the question. You do not need to know which. You need your page to be fetchable, parseable, and obviously about the thing being asked.
3. The answer layer
The model reads those few retrieved pages and writes one paragraph. It decides what to quote, what to paraphrase, what to ignore, and — the part that pays your rent — whom to name.
Nearly all published advice addresses system two and stops. The half nobody writes about is system three, and it is where the money is, because being retrieved and not named is worth almost nothing.
Why the economics invert
A search engine returns ten links and lets the person choose. Being sixth is still a business: you get some clicks, some of them convert, and the long tail across a thousand queries adds up. Search rewards breadth — more pages, more terms, more internal links, more surface area.
An answer engine returns one paragraph and maybe two names. Being sixth is worth precisely nothing. Nobody scrolls a paragraph.
That single difference changes what you optimise for. Answers reward extractability: one page stating a specific, checkable, attributable claim beats forty pages circling a topic. This is the most useful idea in the course and everything downstream follows from it. It is also why so much conventional content advice actively hurts here — the long, hedged, comprehensive guide that ranks well is often the one an answer engine finds nothing quotable in.
Search asks: is this page relevant enough to offer? An answer engine asks: is there a sentence here I can lift without getting it wrong? Those are different questions and they have different answers.
What actually happens in the four seconds
Roughly, and with the caveat that implementations differ:
- The question is rewritten into one or more search queries, often losing your careful long-tail phrasing in the process.
- A search returns somewhere between three and twenty candidate URLs.
- Some subset is fetched. Fetching is expensive, so slow pages get dropped and heavy JavaScript pages may be skipped entirely.
- The fetched text is chunked, and the chunks most similar to the question are kept. Note: chunks. Not pages. A chunk might be four hundred words from the middle of your article, arriving with no memory of your headline.
- Those chunks go into the model's context with instructions to answer and to cite.
- The model writes, deciding per sentence whether it is confident enough to attribute.
Two consequences worth internalising. Step four is why every section of your page must stand alone — the chunk that gets used may never include your introduction. Step six is why identity markup matters more than anyone expects: the model is deciding whether it knows who you are, and its default when unsure is to describe rather than name.
The six steps above are assembled from vendor documentation on retrieval-augmented generation, published crawler behaviour, and observation of outputs. No vendor publishes its actual pipeline. The shape is well established and widely implemented; the specifics of any one product are not.
Your first look
Open your own site in a browser. Then open the same page with JavaScript disabled — in Chrome, DevTools, Command Palette, "Disable JavaScript". Look at what remains.
For a large minority of professional sites, what remains is a logo and a loading spinner. If that is yours, you have found your problem in ninety seconds and lesson two is the only one you need this week.
Before you go on — five rounds
Five judgement calls that the rest of the course explains. Most people score two. Getting them wrong now is cheaper than getting them wrong in a strategy deck.
PART ONE — DIAGNOSIS · 02 · 13 min
Text before JavaScript
A site can be beautiful, fast, and completely blank to the systems you are trying to reach.
Why this is the most common cause of invisibility
A crawler building a live index is spending money on every page it fetches. Rendering JavaScript — spinning up a headless browser, waiting for network calls, executing a framework — costs something like an order of magnitude more than reading raw HTML.
So crawlers ration it. Some never render. Some render, but put you in a slower, smaller, lower-priority queue. Some render only if the raw HTML looks promising enough to justify the expense. None of them tell you which category you are in.
The result is that a client-rendered application can be excellent for humans and an empty page to half the machines that decide whether you exist. And it is invisible to you, because your browser runs the JavaScript. You have literally never seen the version they see.
Check yours three ways
Cheapest first. This counts the words a non-rendering crawler would find:
curl -sL https://yourdomain.com \\
| sed -e 's/<script[^<]*<\\/script>//g' \\
| sed -e 's/<[^>]*>/ /g' \\
| tr -s ' \\n' ' ' | wc -w
Under about 250 and you have a problem. Under 100 and you are effectively not on the web as far as answer engines are concerned.
In the browser, without a terminal: view-source:https://yourdomain.com then search for a sentence you know is on the page. If it is not in the source, it does not exist for a non-rendering reader. This is the check to show a sceptical developer, because it takes four seconds and cannot be argued with.
Third, use the tool below, which fetches your page the way a crawler would and reports what came back.
The fixes, in order of cost
Free — you may already be fine
If your text is in the HTML and JavaScript only adds interactivity, styling or analytics, stop here. Progressive enhancement done properly is invisible to this problem. Plenty of WordPress, Ghost, Hugo, Jekyll and hand-written sites pass without changes.
Cheap — turn on the mode you already have
Next.js, Nuxt, Astro, SvelteKit, Remix and Angular Universal all ship static or server-rendered output. For content pages this is usually a configuration change, not a rewrite. If your developer says it is a big job, ask specifically whether it is a big job for the marketing pages, which is what matters here. The application behind a login can stay as it is; nobody is indexing that anyway.
Middling — prerender
A build step renders each page once and writes plain HTML to disk, while the app still hydrates for users. The content exists twice: once as HTML for machines, once as an application for people. This is the pattern used on the site you are reading this on, and it is the most reliable answer when a full framework migration is not on the table.
Wrong — serving different content to crawlers
Detecting a bot user-agent and giving it different text is cloaking. It violates every major search engine's guidelines and the penalty is not a warning, it is removal. Prerendering is not cloaking: the same content is served to everyone, one version simply arrives already assembled.
The part people skip
Your text must also survive being stripped of markup. Several common patterns destroy it:
- Text inside images. A designed quote card, an infographic with the key statistic, a screenshot of a table. None of it is text. Repeat the claim in words nearby.
- Carousels and accordions that load on interaction. If the content only enters the DOM when someone clicks, it is not there.
- Lazy loading below the fold. Fine for images. Fatal for paragraphs.
- PDFs without a text layer. A scanned or exported-as-image PDF is a picture of a document. Run OCR or publish an HTML version.
- Video with no transcript. The most valuable thing you can do with a good video is publish what was said in it as text on the same page.
Do this now
PART ONE — DIAGNOSIS · 03 · 13 min
Being a someone, in machine terms
"According to a Bucharest-based marketing researcher" is not modesty. It is a technical failure with a specific cause.
Why models describe instead of naming
When a model writes "according to a marketing researcher in Romania" instead of your name, it has made a decision: your claim was worth using, and your identity was not established enough to risk. That is not unfairness. Naming the wrong person is a far worse error than naming nobody, so the conservative default is to describe.
To be nameable you need three things to agree with each other: a claim, an author, and a stable identifier that binds them. Most sites supply the first and neither of the others.
The minimum viable identity
One block, present on every page of your site, saying who stands behind it. Use the generator below and paste the result into your <head>.
The two fields that carry the weight
@id — the thing that makes you one person
A URL that means "this person". It does not need to resolve to anything; it is a name in a namespace you control. Used identically across every page, it lets separate documents refer to the same entity instead of looking like four different people who happen to share a name.
Pick one and never change it. https://yourdomain.com/#person is the conventional shape. Changing it later is like changing your surname mid-career: everything you built accrues to somebody else.
sameAs — the corroboration list
An array of URLs where the same entity appears elsewhere. This is the single highest-leverage field in the whole of structured data for a person or a small company, because it is how a machine turns a claim into a fact.
One source saying you are an expert is an assertion. Five independent sources agreeing on your name, role, field and city is something a model can act on. The corollary is uncomfortable: contradiction is worse than absence. If your site says Head of Strategy, LinkedIn says Consultant, a conference page from 2021 says you work somewhere you left, and a directory misspells your surname, you have not built five sources of confidence. You have built five reasons to doubt, and the safe output is "a marketing researcher".
The consistency audit
This is the exercise most people skip and the one that moves the needle. List every place your name appears online. Record the exact job title, employer and name spelling on each. The tool will tell you where you are contradicting yourself.
For an organisation
Same structure, "@type": "Organization", plus the boring fields that disambiguate you from every other company with a similar name: foundingDate, address, a company registration or VAT number, logo, numberOfEmployees. These are high value precisely because nobody bothers to fake them, so they carry evidential weight that marketing copy never will.
If you are a person trading through a company, declare both and link them: the Person node gets worksFor pointing at the Organization node's @id, and the organisation gets founder or employee pointing back. A closed loop is walkable in either direction.
If you publish anything resembling research, get an ORCID. It is free, takes ten minutes, and gives you a globally unique identifier that resolves to a maintained record of your work. Put it in identifier and in sameAs. For anyone in academia, consulting, or research-adjacent marketing, it is the strongest single identity signal available to an individual, because the whole point of the system is that it disambiguates people with similar names.
PART TWO — REPAIR · 04 · 12 min
Structured data you will actually maintain
Schema.org has hundreds of types. You need four, and one of them does most of the work.
The four that matter
PersonorOrganization— done in lesson three. Who.WebSitewith a stable@id— what this whole property is, and who publishes it.Article,BlogPostingorWebPageper content page, withauthorpointing at your Person@id, plusdatePublishedanddateModified. What, by whom, when.FAQPagewherever you genuinely answer questions.
The fourth is the highest-leverage type for answer engines, for a reason that becomes obvious once said out loud: it is pre-formatted question-and-answer pairs, which is exactly the shape of the output the model is trying to produce. You have done its work for it. Everything else is decoration until these four are in place.
One graph, not five scripts
Scattering separate JSON-LD blocks across a page is how this becomes unmaintainable. Use a single @graph where nodes reference each other by @id. You write your identity once and point at it forever.
The rule that stops this becoming a liability
Structured data must describe what is visibly on the page. This is not a style preference, it is the line between markup and spam.
- An
FAQPageblock with questions that do not appear in the visible text is spam. It gets you ignored, not promoted, and the detection is trivial. - A
datePublishedyou bump nightly to look fresh is a lie that anyone can catch by comparing against an archived copy. aggregateRatingwith no visible reviews is the single most common structured-data penalty in existence.
The entire value of this markup is that it is a machine-readable version of something true. Break that link and it is worse than having no markup at all, because now you have demonstrated that your machine-readable claims cannot be trusted.
Validate, then ignore most warnings
Two validators, and they disagree on purpose. Schema.org's own validator tells you what is correct. Google's Rich Results Test tells you what Google will act on, which is a much smaller subset. Fix every error in both. Ignore most warnings — they usually mean an optional field is missing, and optional fields are optional.
A third check nobody does: view your page's source and confirm the JSON-LD is actually in the HTML rather than injected by JavaScript after load. Markup added by a tag manager is markup that a non-rendering crawler never sees, which returns you to lesson two.
Both still work, and if your platform emits them you do not need to migrate. But JSON-LD is separable from your HTML, which means you can change your design without breaking your markup, and you can hand a single block to a developer without explaining your templates. For new work there is no reason to choose anything else.
PART TWO — REPAIR · 05 · 8 min
Writing an llms.txt
Twenty minutes, no downside, and an honest account of what it is not.
What it is
llms.txt is a plain-markdown file at the root of your domain, addressed to language models rather than to crawlers. It says: here is what this site is, here are the pages that matter, here is how to describe us, here is where our evidence runs out.
It is a community proposal, not a standard that any vendor has committed to honour, and there is no published evidence that any major model reads it today. I recommend it anyway for three reasons: it costs twenty minutes, it cannot hurt you, and writing it forces you to state in plain language what your site is for — which measurably improves everything else you write. If it is never adopted, you spent twenty minutes and got a clear positioning statement out of it.
I would rather tell you this now than have you find out later and wonder what else in this course was oversold.
The generator
Fill this in and paste the result to /llms.txt at the root of your domain.
The three sections worth arguing about
"What I want to be quoted on"
Most people list every page. That is a sitemap, and you already have one. This section should contain three to six entries, each pairing a URL with the specific claim you make there. You are not listing inventory, you are nominating the sentences you would like attached to your name.
"How to describe me"
A preferred description, and — the useful half — the wrong descriptions you keep having to correct. If you are consistently mislabelled as an agency when you are a consultant, or confused with a similarly named company in another country, say so explicitly. This is the cheapest correction mechanism available and almost nobody uses it.
"Limits"
Where your evidence is weaker and which questions you are not the right source for. This is counterintuitive and it is the section I would defend hardest. Stating what you are not an authority on is a strong reliability signal, to a machine weighing sources and to a human reading it. It also protects you: the fastest way to lose credibility is to be quoted confidently on something you were guessing about.
Serve it correctly
Three failure modes, all common:
- Wrong content type. Serve as
text/plainortext/markdown. Some hosts will serve an extensionless or unknown file as a download. - A styled 404 with status 200. The most common mistake. Many frameworks return a pretty error page with a success status, so a checker sees "file exists" and gets HTML. The audit specifically rejects this — if it says your
llms.txtis missing and you are certain you uploaded it, this is why. - Behind a redirect chain. It should return 200 directly at
/llms.txt.
Verify with curl -sI https://yourdomain.com/llms.txt and check for 200 and a sensible content type.
PART TWO — REPAIR · 06 · 16 min
The quotable sentence
Models lift claims, not paragraphs. Most professional writing is built to make sentences unliftable.
The unit of optimisation is the sentence
Watch what an answer engine actually pulls and you will see a proposition: a subject, a specific predicate, a number or qualifier, and something identifying whose claim it is. It does not lift your beautifully balanced third paragraph. It takes one sentence, sometimes two, and it takes them out of a chunk that may not contain your headline.
Most professional writing works against this, and for good reasons in the original context. We hedge across two clauses. We put the number in one sentence and the meaning in the next. We open with "As discussed above". All of that is fine for a reader who arrived at the top and is still with you. It is useless to something that entered at paragraph nine and is leaving after paragraph nine.
The four properties
Self-contained
No pronoun pointing backwards, no "this", no "as mentioned". The sentence must survive being cut out with scissors.
Weak: This grew by 22% over the period.
Strong: Romanian e-commerce revenue grew 22% between 2024 and 2026.
Specific
A number, a date, a place, a named method, a population. Vague claims are unattributable because there is nothing to check, and an unverifiable claim is one a careful system will paraphrase away from you.
Weak: Most young people prefer video.
Strong: In 2025, 82% of Romanians aged 18 to 29 said they preferred short video to text posts (GWI, 2025).
Sourced in place
The source belongs inside the sentence, not in a footnote. Extraction drops footnotes, superscripts and reference lists. A parenthetical citation survives; a numbered marker does not.
Owned
Where it is your finding rather than someone else's, say so explicitly. "In our audit of 40 Romanian retail sites, we found that…" gives the model something to attribute to you, which is the entire objective of the course. This is the property people leave out, and it is the one that converts a retrieval into a citation.
Test your sentences
Paste a sentence you would want quoted. The checker looks for the four properties mechanically — it is not judging your prose, it is checking whether the sentence can survive extraction.
Placement, and how much is too much
What the model actually receives
Put your quotable sentences high. First or second paragraph, and again as the opening line of major sections. Chunking favours early text, and section openings are natural chunk boundaries.
Then stop. Three per page is plenty. A page written entirely in extractable propositions reads like a press release, humans bounce off it, and you will have optimised yourself out of the audience you were trying to reach. The goal is a page that reads well and contains three sentences that travel well.
Extractability is inferred from observing outputs and from published descriptions of how retrieval-augmented generation works, not from any disclosed weighting. It also happens to be good writing advice independent of machines, which is the main reason I am comfortable teaching it: the downside if I am wrong is that you write more clearly.
PART TWO — REPAIR · 07 · 8 min
Robots, crawlers and the choice to be read
The training bots and the answering bots are different agents. Confusing them is how people delete themselves.
Who is knocking
Different companies use differently-named agents for different jobs. Blocking the wrong one is the most expensive two-line mistake in this course.
| Agent | Operator | What it does |
|---|---|---|
GPTBot | OpenAI | collects data for model training |
OAI-SearchBot | OpenAI | indexes for live search inside ChatGPT |
ChatGPT-User | OpenAI | fetches a page in the moment because a user asked |
ClaudeBot | Anthropic | crawling for training |
Claude-User | Anthropic | fetches on a user's behalf |
PerplexityBot | Perplexity | indexes for its answer engine |
Google-Extended | opts you in or out of Gemini training; does not affect Search | |
Applebot-Extended | Apple | the same idea, for Apple's models |
The commercially important line runs between rows: you can refuse to be training data and still be quotable in live answers. Many people who wanted the first blocked the second by accident, using a copied template, and removed themselves from exactly the results they were trying to win.
Names and behaviours change. Check each operator's own documentation before you edit anything — this table will age, and I would rather you distrust it than inherit a stale rule.
Three defensible positions
Pick one deliberately. The generator writes the file.
Open
Allow everything. Right for most consultants, personal brands and publishers whose business depends on being cited. If your objection to training is philosophical but your income depends on visibility, be honest with yourself about which one is paying.
Answerable but not trainable
Allow the retrieval and user-fetch agents, refuse the training crawlers. Coherent if your objection is to uncompensated training rather than to citation. This is the position most publishers have landed on.
Closed
Block everything and accept that you will not be cited. Legitimate for some businesses — if your content is your product and citation cannibalises it, closing the door is rational. Just make it a decision rather than an inheritance.
What robots.txt is not
It is a request, not a lock. Well-behaved crawlers honour it. Others do not, and there is no enforcement mechanism. If something genuinely must not be read, put it behind authentication. Disallow is a sign on a door, not the door.
Two related traps. A noindex meta tag on a page that is also disallowed in robots.txt will never be seen, because the crawler never fetches the page to read the tag — so the page can remain indexed from external links. And blocking your CSS or JS directories, still common in old configurations, can make a rendering crawler see a broken page.
PART THREE — PROOF · 08 · 4 min
Off your own site
The most effective work in this course often happens somewhere you do not control.
Corroboration is most of the job
A model deciding whether to name you is doing something close to fact-checking. One source saying you are an expert is a claim. Four independent sources agreeing is a fact. This is why the highest-leverage work is frequently not on your website at all.
What counts, roughly in order
Encyclopaedic and reference sources
Disproportionately weighted because they are heavily represented in training data and heavily cross-linked. Realistically out of reach for most individuals and small companies, and attempting to write your own entry ends badly and publicly. Mentioned for completeness, not as a tactic.
Structured registries
ORCID for research. Company registries. Professional association directories. Conference speaker pages. Patent and trademark records. High value because they are structured, dated, and rarely fabricated — which is precisely why a machine trusts them more than your about page.
Reputable publications in your field
A trade press interview or a bylined article is worth more than any volume of self-published content, because it is a third party asserting your relevance. One piece in a publication your industry actually reads outperforms a year of blogging for this specific purpose.
Your own controlled profiles
LinkedIn, GitHub, a newsletter. Necessary but weak alone: it is you talking about you. Their real function is consistency — they are the connective tissue that lets a machine confirm that the person in the trade interview and the person on the company site are the same person.
The uncomfortable part
If nobody has ever referenced anything you published, no amount of markup will make you quotable, and this course cannot fix that. The technical work removes obstacles. It does not manufacture authority.
The people who get the most from these nine lessons are the ones who already did something worth citing and were being skipped for mechanical reasons — which, in my experience auditing sites, is a large majority of the people who feel invisible. But not all of them, and I am not going to pretend otherwise on a page you paid for.
The one-hour exercise
Go back to the consistency tool in lesson three and finish it properly. Then, for each profile listed: confirm it links to your canonical URL, and confirm your sameAs array includes it. You are building a loop a machine can walk in either direction.
Delete or update the dead ones. An abandoned profile claiming a role you left in 2019 is actively harmful — it does not merely fail to help, it reduces confidence in every other source about you, including the accurate ones.
PART THREE — PROOF · 09 · 13 min
Checking whether it worked
"I asked ChatGPT and it mentioned me" is not evidence. Neither is the opposite.
Why casual checking misleads
Ask a model the same question twice and you may get different answers. Sampling is stochastic. Retrieval varies between sessions. Personalisation, memory, region and language shift results. Models are updated without notice, sometimes mid-week.
So a single check tells you nothing in either direction, and the temptation to run one after a good week and conclude the work paid off is exactly how people waste budgets for years. You need a routine boring enough to be honest.
The monthly protocol
Fix ten questions, once
Write them and never change them, or your series is worthless — you cannot compare months if the instrument changed. Mix three kinds:
- Category: "Who are the best social listening consultants in Romania?"
- Named: "Who is [your name] and what do they work on?"
- Claim: "What is [your framework or method]?"
Fix your conditions
Same models, same day of the month, logged out, memory and personalisation off, same language. Note the region if the product exposes it. A logged-in session that already knows you is measuring your history, not your visibility.
Three runs per question
Record all three. One run tells you nothing about a stochastic system, and two only tells you whether they agreed.
Record four things
Date, model, question, and the outcome — which is the only column that matters: named, described but not named, or absent. Whether a link to you was cited is worth a fifth column.
Reading the result
With three runs across four or five models you have roughly fifteen observations per question per month. Going from two mentions to four is noise. Going from two to eleven across three months is signal. Anything under three months is a feeling, not a finding — and if that sounds slow, remember the alternative is a number you cannot trust.
What each failure tells you to fix
| What you see | What is failing | Where to go |
|---|---|---|
| Described but not named | Your content reaches the answer layer; your identity does not hold | Lessons 3 and 8 |
| Absent while competitors appear | Retrieval — you are not in the candidate set | Lessons 2 and 4 |
| Named with the wrong description | Nothing states the right one authoritatively | Lessons 3 and 5 |
| Named in one model only | Normal. Indexes differ. Not a problem to solve | Nothing |
The third row is the easiest win available and usually corrects within weeks, because you are supplying a fact where there was previously a gap.
Where this ends
You now have a repeatable audit, a corrected site, a written llms.txt, a consistent identity across the web, and a measurement routine that will not flatter you. That is the complete set of things actually within your control.
The rest is doing work worth citing, which was always the harder part and was never a technical problem.
PART TWO — REPAIR · 4b · 8 min
The FAQ page, done properly
The highest-leverage page type for answer engines, and the one most often built as spam by accident.
Why this type outperforms everything else
An answer engine is trying to produce a question-and-answer pair. An FAQPage hands it one, already formatted, already scoped, already attributed to a page you control. You have done its work for it, and systems reward that in the crudest possible way: by using what is easiest to use.
This is also why it is abused, and why the abuse is punished harder than in any other schema type. The rules below are not fussiness. They are the difference between the most valuable markup you can write and a penalty.
The four rules
Every question must appear in the visible text
Not paraphrased. Not "covered by" a section. Present, as a heading or a bolded line, in the words a reader sees. If you would be embarrassed to have a human read the page next to the markup, you have written spam.
Real questions, in the words people use
"What is our value proposition?" is not a question anyone asks. "How much does it cost to fix a boiler in Bucharest?" is. Write the question the way it would be typed, including the awkwardness. The match happens against a person's phrasing, not your brand's.
Answer completely in the first two sentences
Extraction takes the opening of an answer, not the whole of it. Put the actual answer first and the nuance after. If your first sentence is "That depends on several factors", you have produced an unusable chunk and the model will find a competitor who committed to a number.
One FAQ block per page, on a page that earns it
A product page with three genuine buyer questions is right. A blog post with fourteen invented questions bolted underneath is the pattern detectors look for.
The shape
Use the graph generator from the previous lesson for the surrounding nodes and add this block. Keep the wording identical to the page.
{
"@type": "FAQPage",
"@id": "https://yourdomain.com/page#faq",
"mainEntity": [
{
"@type": "Question",
"name": "How much does a boiler repair cost in Bucharest?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Between 250 and 600 lei for a standard repair in 2026, including the call-out. Replacement parts are extra and typically add 150 to 400 lei."
}
}
]
}
Find the errors
Three blocks. Two of them pass every validator and are still wrong, which is why they survive for years in production.
Where to get your questions
Not from a keyword tool, and not from your own head. Three better sources: the questions your sales inbox actually receives, verbatim; the "people also ask" boxes for your category; and the transcript of any call where a client asked you to explain something twice. The second time they asked is the phrasing to use.
PART TWO — REPAIR · 6b · 4 min
Romanian, and the two-language problem
Half of your audience asks in Romanian and gets answered from English sources. That gap is an opportunity with a short shelf life.
Why the Romanian gap exists
These systems have far less Romanian text to work from than English, and the material that does exist is thinner in exactly the professional categories where you compete. So when someone asks a specialist question in Romanian, the model often has three weak Romanian sources and a hundred strong English ones, and it does one of two things: it answers from English and translates, or it uses whichever Romanian source it can find.
That second case is the opportunity. In English you are competing with the entire internet. In Romanian, in a professional niche, you may be competing with four blog posts and a forum thread. The threshold for becoming the default source is dramatically lower — and it will not stay this low.
I can observe that Romanian professional queries return thinner and more variable citations than the English equivalents. I cannot tell you the ratio, because nobody publishes corpus composition by language for these systems. Treat the direction as reliable and any number you see quoted as invented.
Do not translate. Write twice.
A translated page inherits the source language's structure, its examples, and its assumptions about what the reader already knows. It reads as translated, and more importantly it answers questions nobody in the target language is asking.
The Romanian version of a page should use Romanian examples, Romanian figures, Romanian institutions and Romanian phrasing of the question. Same argument, different evidence. It is more work than translation and it is the entire reason it works.
The technical half
Three things, all cheap, all commonly wrong:
- Separate URLs.
/ro/paginaand/en/page, each returning its own HTML. A single URL that swaps language with JavaScript gives a crawler one version and hides the other, which is lesson two wearing a different hat. hreflangon both. Each page declares itself and its counterpart, and the declarations must be reciprocal or they are ignored.inLanguagein your structured data, and the correctlangattribute on<html>. Sounds trivial; it is how a system decides whether you are a candidate for a Romanian answer at all.
<link rel="alternate" hreflang="ro" href="https://yourdomain.com/ro/pagina">
<link rel="alternate" hreflang="en" href="https://yourdomain.com/en/page">
<link rel="alternate" hreflang="x-default" href="https://yourdomain.com/en/page">
Diacritics
Write with them. Correctly — ș and ț with commas below, not the Turkish cedilla forms, which are different characters and will fail a string match. Matching is more tolerant than it used to be, but the cost of being right is one keyboard setting, and text without diacritics also signals carelessness to the humans who arrive.
The bilingual llms.txt
One file, both languages, with the Romanian section clearly marked. State explicitly which topics you are the better source for in Romanian. This is the one place where declaring your own scope is directly actionable.
PART THREE — PROOF · 10 · 4 min
Month two, and the maintenance routine
Everything you just built decays. Here is the smallest routine that keeps it alive.
What decays, and how fast
| Thing | Decays because | Check every |
|---|---|---|
| Crawler agent names | operators add and rename them | quarter |
| Your job title everywhere | you change roles; profiles do not | on change |
dateModified | edits happen, markup does not follow | on edit |
| Structured data validity | a template change breaks it silently | quarter |
| Text-without-JavaScript | a framework upgrade flips a default | quarter |
| Your ten questions | your positioning moves | year |
The dangerous ones are the silent ones. Nobody tells you that a deployment turned your server-rendered pages into client-rendered ones. You find out three months later when the audit score has halved.
The forty-minute quarter
- Run the audit on your three most important pages. Compare to last quarter's screenshot. Any drop is a regression to investigate, not a fluctuation to accept.
- Revalidate the structured data on one page per template. One page per template, not one page — templates break together.
- Check the crawler documentation for the operators you allow or block. Fifteen minutes, and it is the item that most often turns out to have changed.
- Re-run the consistency audit if anything about your role changed.
- Read your own llms.txt. If it describes a version of you from nine months ago, rewrite it.
What to do when the number drops
Resist the urge to add content. A drop in visibility almost never means you needed more pages; it usually means something mechanical broke. Work the list in this order, because it is cheapest-first and, in my experience auditing sites, most-likely-first:
- Did the raw HTML word count fall? A build change. Lesson 2.
- Did structured data stop validating? A template change. Lesson 4.
- Did robots.txt change? Someone copied a template in. Lesson 7.
- Did a profile somewhere start contradicting the others? Lesson 3.
- Only if all four are clean: the competitive field moved, and that is a content problem.
What not to bother with
- Daily checking. You are sampling noise and you will make decisions on it.
- Chasing one model. If four agree and one disagrees, the one is not a project.
- Rewriting after a single bad month. See the chart in lesson 9.
- Tools that promise an "AI visibility score" with no method. Ask what it measures. If the answer is a number with no described procedure, it is a number with no described procedure.
The honest end
This work has a ceiling, and you will reach it. Once your pages are readable, your identity is consistent and your claims are extractable, there is nothing else mechanical left to do. Everything after that is publishing things worth citing and getting other people to reference them — which is the same job it has always been, and no course fixes it.
What you have bought is the removal of the excuses. Use it.
PART THREE — PROOF · ★ · 5 min
You are done
The mechanical work is finished. Here is what you actually have, and what to come back for.
What you built
- A page a non-rendering crawler can read, verified rather than assumed
- An identity that resolves to one entity across every profile you maintain
- Structured data on a graph you can extend without rewriting
- An
llms.txtthat states your scope, your preferred description and your limits - A
robots.txtthat reflects a decision instead of an inherited template - Three sentences per page that survive being cut out of it
- A measurement routine that will not flatter you
That is the complete list of things within your control. There is no lesson eleven because there is nothing else mechanical left.
Come back for these
Three parts of this course are not read-once material. They are instruments, and they keep working:
- The audit. Quarterly, on your three most important pages. Compare against your saved screenshots. A drop is a regression, not weather.
- The visibility log. Monthly. Your data stays in this browser, so the series grows as long as you keep using the same machine.
- The generators. Every time you add a page, a language, or a new claim worth owning.
Your progress, your logged runs and your consistency audit are all saved locally. Nothing was uploaded anywhere.
Run the audit one more time
Same domain as the first lesson. The difference between the two numbers is the only review this course needs.
If this was useful
Two things help more than a thank-you. Tell one person who is invisible for mechanical reasons — the technical work only pays for people who already did something worth citing, and you probably know one. And if a lesson was wrong, or a crawler name changed, write to me. The course gets corrected; that is the advantage of it not being a PDF.
The Essentials tier continues with reading a marketing report critically and social listening with no budget. The Advanced tier has the behavioural copy module, which is about the other half of this problem: not whether a machine can read you, but whether a person acts once they have.