From PDF to AI-friendly page: what happens to a monthly letter along the way

A monthly letter in PDF reaches a model as one block, with no declared hierarchy and no machine-readable authorship. This article follows the conversion end to end: what the old format loses, what the page gives back and how much work is left for the asset manager's team.

What exactly does an AI lose when it reads a PDF?

A PDF stores the position of text on the page rather than the level of a heading, so hierarchy disappears. Authorship and date carry no machine-readable markup. And because a model retrieves short passages, a twelve-page document reaches it as one indivisible unit.

What the letter hasWhat the PDF delivers to a modelWhat the page gives back
Section headingtext in a larger size, with no semantic levelan <h2> in question form, answered right below
Authorshipa name printed on the cover, unmarkedPerson in JSON-LD plus a visible authorship box
Revision datea date in the header, with no declared fielda declared, verifiable dateModified
Retrievable unitthe whole fileeach section, retrieved on its own
Table of figuresa drawn grid that extraction tends to scramblea <table> with headers and a declared source
Link to the rest of the archivenoneinternal links to neighbouring pages

Source: ReBo AI Visibility method, September 2026.

The most expensive loss is the last line of the middle column. Retrieval systems work through query fan-out: one prompt becomes dozens of parallel searches, and what each returns is a passage, with its heading travelling alongside. A PDF enters that contest as a closed block. Either the whole file is judged relevant to the sub-search, or it is absent from it.

There is also the case where the loss is total. A letter exported as an image, or scanned from paper, has no text layer at all. The mechanism is covered in why AIs cannot read PDFs.

How does a monthly letter become an AI-friendly page, step by step?

In five steps: an inventory of what is durable, decomposition into self-contained sections, the layer a PDF never carries, preservation of the original figures and caveats, and publication at a permanent URL with valid JSON-LD.

It helps to follow a real fragment. Take the "Macro outlook" section of an August letter, four paragraphs on rates, currency and the effect on the fund's positions. In the PDF, that fragment is a rectangle of text on page 3.

In the inventory, that section is classified as commentary with a short shelf life. The thesis on duration at the end of it is classified as durable and gets a destination of its own. Much of what an asset manager writes every month is commentary; the durable part is what holds a page up for years.

In the decomposition, the durable fragment becomes a section with a question-form heading ("Why does short duration protect in a hiking cycle?") and an anchor answer of roughly 45 words right below it. The rule behind the cut: the passage has to make sense alone, because alone is how it will be read.

In the addition step, in comes everything the PDF never carried. Structured data with author, publisher and date. A visible FAQ with exact parity against the FAQPage. An authorship box with verifiable credentials. A table with a declared source in place of a chart image.

In the preservation step, figures, sources and caveats stay intact, including the wording compliance approved. In publication, the page ships as static HTML at a canonical URL with a trailing slash, linked to its neighbours in the archive. The operational walkthrough, with the common mistakes at each stage, is in how to turn a monthly letter into AI-friendly content.

Does publishing in HTML mean dropping the PDF?

No. The PDF remains the relationship piece for investors, with its own design and signature. The page exists to be read by machines and cited. Each one points to the other, and the canonical URL for that edition sits on the page.

The two formats solve different problems and coexist well. The PDF is the document an investor files, prints and attaches to an email. The page is the asset a retrieval system can read, cut and cite. The practice that works: the page links the PDF for that edition, and the PDF carries the page's permanent URL.

One warning about attribution. Publishing the same analysis at two indexable URLs splits the signal between them. So the canonical for each edition points at the page, and the PDF sits as its attachment.

How much of the work stays with the asset manager's team?

The letter the team already writes each month. Repackaging, technical markup and publication sit with whoever runs AI Visibility. What comes back to the firm is a compliance review of the converted text, with no new production.

The principle behind this is one source, many assets. The raw material already exists inside the firm, produced by the people whose credentials matter. The AI Visibility work is repackaging that material into the formats each reader consumes, including the reader that is a machine.

In the monthly operation, what reaches the in-house team is a review: the converted text goes back for compliance sign-off and for the author to confirm nothing was distorted. Everything else, from markup to the distribution cycle, stays off the analyst's desk.

Keep reading

To understand why the old format fails, start with why AIs cannot read PDFs. For the conversion itself, go to AI-friendly investor letters. To see how a model decides whom to cite once the page exists, read how ChatGPT chooses the sources it cites. To measure the result, start with the diagnostic.

Frequently asked questions

Is it worth converting old letters?

It is worth it when the analysis still holds. A macro commentary from three years ago has aged; a structural thesis about a sector, an in-house glossary and a risk methodology have not. The test is durability, and an unlocked archive usually yields more pages than the current month's production.

Does the conversion change the analysis itself?

It does not. The analytical text is preserved, including figures, sources and caveats. What gets added sits around it: question-form headings, an anchor answer, an FAQ, an authorship box and structured data. Changing the letter's conclusion would mean redoing the analyst's work.

How long until a converted page gets cited?

There is no guaranteed timeline, and ReBo does not sell one. What can be observed is the sequence: the crawler fetches the page, the page enters the search index that feeds the engine, and only then can it enter source selection. Each stage has its own measurement.

Does every letter need its own page?

It depends on what the letter carries. An edition covering four independent subjects goes further as four thematic pages with their own URLs, because each answers a different question. Short single-subject editions work well as one page.