Skip to content

Why AI does not read websites, it reads knowledge structures

Object of the analysis

This article discusses what “reading” means when talking about AI systems and why that notion is not equivalent to human reading of a website.

The goal is not to describe interfaces, crawlers or publication formats, but to dismantle an inherited assumption: that AIs “read pages” in the same way that people read documents.

The analysis is placed after having understood:

  • what it means for an AI to “cite” a source
  • how an AI decides what information to reuse

A deeper level is addressed here: what kind of object the system actually operates on.

1.1 What does “read” mean in AI systems?

In this context, “reading” does not describe an experience or a sequential act.

Describes a technical operation on information representations.

When it is stated that an AI “reads”, it is actually alluding to processes such as:

  • identification of semantic units
  • evaluation of relationships between concepts
  • reuse of compatible fragments
  • integration in a generated synthesis

It does not exist:

  • layout perception
  • linear path of the text
  • narrative comprehension
  • reading experience as such

The system does not “walk” a web site.

It operates on structures that have already been abstracted from the document.

1.2 What is NOT analyzed when we talk about reading

This article does not analyze:

  • how pages are indexed
  • how HTML is crawled
  • how results are presented to users
  • how human reading experiences are designed

Nor does it address:

  • publishing tactics
  • formats “best read by AI”.
  • technical adjustments to facilitate ingestion

All these issues belong to different layers.

Here, we analyze exclusively the conceptual level: what type of object is readable by a generative system and what is not.

1.3 Scope of disassembly

The scope of this article is negative and structural, not prescriptive.

Not proposed:

  • how to “make an AI read
  • how to adapt a website
  • how to improve results

It is demonstrated why the assumption “AI reads my web” does not describe the actual operation of the system and why maintaining such a framework leads to design, governance and expectation errors.

This dismantling is necessary in order to introduce the distinction between:

  • document
  • content
  • knowledge

and the notion of layers, central to the Shymow standard.

2. The assumption inherited from the classical model

To understand why the concept of “reading” fails when applied to AI systems, it is necessary to identify the assumption from which it comes.

This assumption is not born with AI. It comes from the classic web model.

2.1 The web as a set of documents

For decades, the web has been understood as a system composed of documents:

  • individual pages
  • with a beginning and an end
  • organized to be traveled by humans
  • consumed sequentially

In this model:

  • each URL represents a significant unit
  • the document is the container of meaning

Meaning emerges from the whole, not from the parts.

Human reading presupposes:

  • continuity
  • cumulative context
  • editorial intent

This framework is consistent for people. It is not necessarily consistent for automated systems.

2.2 Human reading vs. automatic processing

Human reading is:

  • sequential
  • contextual
  • dependent on intent
  • sensitive to form, emphasis and order

An automatic system does not reproduce this behavior.

It does not “move forward” through a text.

It does not interpret rhythm or narrative hierarchy.

It does not reconstruct intention based on editorial order.

Instead of reading, process:

  • fragments
  • patterns
  • relations
  • recurrences

What for an individual is a coherent document, for the system is fragmentable raw material.

2.3 Why this assumption is no longer valid

The error appears when the human model of reading is transferred to the operation of generative systems.

Assuming that an AI “reads webs” implies believing that:

  • the document is the operating unit
  • layout brings meaning
  • editorial order conditions comprehension
  • publishing is the same as exposing knowledge

In generative systems, none of these assumptions hold.

The document is not the final unit.

It is a presentation device, not an object of knowledge.

3. Conceptual separation: document, content and knowledge

The central error of the classical model is not technical, but conceptual: it treats three levels that fulfill different functions as equivalent.

3.1 Document as a presentation artifact

A document is an object designed for human reading.

  • has a specific format (HTML, PDF, page)
  • has an editorial structure
  • responds to a single communicative intention
  • assumes sequential reading

From the point of view of the generative system, the document:

  • is not a stable semantic unit
  • does not preserve meaning by fragmenting
  • cannot be reused as such

The document is not legible as a complete object in an operational sense.

3.2 Content as editorial material

The content is what is expressed within the document.

  • explanations
  • examples
  • arguments
  • comparisons
  • narratives

The content may be rigorous and valuable, but it remains tied to the document.

Not all content is reusable by generative systems.

3.3 Knowledge as a reusable semantic unit

Knowledge is a unit that:

  • has its own meaning outside the document
  • can be explicitly defined
  • maintains intention by fragmenting
  • can be attributed to an entity
  • has clear limits

It can be operated by automated systems because it can be identified, evaluated, and reused.

When an AI “reads”, the only thing it can really read is this level.

4. Why an AI does not “read” documents

The notion of reading applied to documents does not describe the actual behavior of a generative system.

4.1 Lack of reading experience

A generative system has no reading experience.

It does not exist for the system:

  • a beginning more important than an ending
  • a storyline to be followed
  • an editorial intention to interpret

It operates on abstract representations.

4.2 Fragmentation and operational decontextualization

In order to operate, the system must fragment.

  • the document is decomposed
  • the content loses editorial context
  • only fragments with autonomous meaning survive

Fragmentation is a necessary condition for reuse.

4.3 Layout and narrative inoperability

Core elements of human reading are inoperable for AI:

  • visual design
  • graphic hierarchy
  • narrative rhythm
  • editorial order

AI does not ignore the document due to technical limitations.

He ignores it because it is not the correct object.

5. What does it mean for an AI to “read” knowledge structures?

“Reading” means operating on previously abstracted and stabilized knowledge structures.

5.1 Identification of entities and relationships

The system recognizes entities and relationships.

Without entity, no reading is possible.

5.2 Use of definitions and limits

A structure is readable when it exposes:

  • clear definitions
  • explicit limits
  • application conditions
  • consistent patterns

Reading means reducing operational ambiguity.

5.3 Access vs. operational understanding

Accessing is not the same as understanding operationally.

To understand operationally is to be able to use without reconstructing.

6. The notion of layers

There is no single layer of reading.

6.1 Human layer

Document, narrative, and persuasion.

6.2 Machine layer

Structure, serialization, and stability.

6.3 Necessary separation

Layers can coexist, but they cannot condition each other.

7. Conceptual closure

AI does not read websites.

What they read are previously defined and stabilized structures of knowledge.

The correct question is no longer “How does my website read?”

and becomes: what knowledge exists in a readable form outside the document?