Notesproduct6 min

A glossary built to be quoted

1,070 terms, five languages, and a rule that decides which ones a machine may never touch. What the glossary is for, how an entry is made, and why the definitions are deliberately short.

The short answer

what
1,070 terms defined once each, published as 1,590 translations across English, Ukrainian, Spanish, French and Polish, with the terms that depend on each other linked together.
why
A retriever returns passages, not pages. A definition that cannot stand alone in one passage cannot be quoted correctly, and a term nobody defined gets defined for you by whoever did.
who
Anyone trying to work out what a term in AI search actually means, and any assistant answering that question on somebody's behalf.
where
getlooploop.com/glossary, free and open, no account.
when
Entries are added and revised continuously. Each carries the model that drafted it, or the mark that a person wrote it — which is the field that decides whether a queue may touch it again.
how
Drafted, reviewed, and published only when both approval and publication are set. Hand-written entries are marked as such and permanently excluded from every generation queue.

In plain words

When you ask a robot a question, it does not read a whole website. It grabs a few paragraphs and answers from those. So we wrote 1,070 short explanations, each one complete on its own, in five languages. Some were written by a person and the machines are not allowed to rewrite those, ever.

Why a glossary, of all things

Because of how the answer gets built. When an assistant answers a question about, say, chunking, it does not read a page on chunking — a retriever splits documents into passages, embeds them, and hands the model back the few that sit closest to the question. The model answers from those passages.

Which makes a definition a very particular kind of writing. It has to survive being cut out of its page and read alone. Anthropic's write-up of contextual retrieval names exactly this failure — their example chunk, the company's revenue grew by 3% over the previous quarter, is useless in isolation because it names neither the company nor the quarter. Prepending context before embedding cut their retrieval failures by 49%, and by 67% with reranking.

A glossary entry is that lesson taken to its conclusion: a passage that carries its own subject, by construction.

The glossary index on getlooploop.com, showing terms grouped into categories with a count beside each, and a search box above them.
The index. Terms are filed into categories that exist because somebody uses them, not because the alphabet needed splitting.

What is actually in it

As of 11 September 2026: 1,070 terms, and 1,590 published translations across the five languages this site publishes in. The distribution is uneven on purpose — English is complete because it is the source, and the other four are filled in term by term rather than dumped through a batch translator.

Glossary translations by language (11 September 2026)

English1,070
Spanish141
French136
Polish124
Ukrainian120
English is the source and is complete. The other four are written one entry at a time, which is why they are behind — and why they read like language rather than like output.

Try it on your own page

See what Voicescope finds on a page of yours

Everything above is measurable on your own work, and the check takes about a minute. Paste one address and read what comes back.

Open Voicescope

Free · no account · one at a time

Five languages, and one that will never be here

The site publishes in English, Ukrainian, Spanish, French and Polish. Russian is not one of them and will not become one. That is a decision, not a gap in the roadmap.

It is also enforced mechanically rather than hoped for, because good intentions did not catch it: a hand-authored Ukrainian entry once shipped with a Russian word inside it, typed by somebody writing five languages in one sitting. A script now scans anything hand-written in Cyrillic for the four letters that exist in Russian and not in Ukrainian — ы, ъ, э, ё — which is a high-precision signal, since none of them can appear in correct Ukrainian. A reading agent catches the rest: the endings and the false friends that use only shared letters.

Every generation prompt names its target language and forbids Russian explicitly, because a model asked for Ukrainian produces Russian words inside it often enough that saying so is worth the line.

How an entry is made, and what stops a machine touching it

Most entries are drafted by a model, then reviewed, and published only when both approval and publication are set — nothing a model produced is public on the strength of having been produced. Each translation stores which model drafted it, and that field is not decoration.

The guard is per term, not per language. A term with hand-written work in any language is hand-maintained, because a queue that quietly filled in the missing languages would leave one page in considered prose and the next in draft. Handing an entry back to the queues means clearing that field — a decision somebody makes and records, not a checkbox on a batch screen.

The marker is the record of how the entry was made. The thing that records how it was made is the thing that decides whether a machine may touch it.

A glossary entry page for Goodhart's Law, showing a short definition, a diagram, and links to related terms.
A hand-written entry, with its own diagram. This one is permanently excluded from every generation queue, in all five languages.
FieldWhat it recordsWhat it decides
modelWhich model drafted the translation, or `hand-written`Whether any queue may ever regenerate it
approvedAtThat a person read the draft and accepted itHalf of the gate to being public
publishedAtThat the entry is liveThe other half — both must be set
localeWhich of the five languages this row isWhich page it appears on, and its hreflang

Why the definitions are short

A long definition is a worse definition for this purpose, and the reason is mechanical rather than stylistic. If the explanation runs across four paragraphs, a retriever takes one of them. The one it takes may be the caveat, or the history, or the example — and the answer built from it will be confidently incomplete.

So each entry leads with a definition that stands alone: what the term is, in a form that survives being the only thing anybody reads. The depth comes after, where a reader who wants it will find it and a retriever that grabs it will still be holding something true.

Terms that depend on each other are linked, so a reader who lands on one can reach the one it assumes. That linking is also what makes the glossary a structure rather than a list.

A relevant chunk might contain the text: 'The company's revenue grew by 3% over the previous quarter.' However, this chunk on its own doesn't specify which company it's referring to or the relevant time period.

Anthropic, Introducing Contextual Retrieval, September 2024

Questions people ask

Is the glossary free?

Yes. No account, no card. It is a public reference and it is meant to be quoted.

Are the definitions written by AI?

Most are drafted by a model and then reviewed; nothing is public until a person has approved it and published it. Entries written by hand are marked as such and permanently excluded from every generation queue.

Why are some languages so far behind English?

Because they are written one entry at a time rather than batch-translated. English is the source and is complete at 1,070 terms; Spanish has 141, French 136, Polish 124 and Ukrainian 120. Filling them faster with a batch job is exactly the shortcut the hand-written rule exists to prevent.

Why is there no Russian version?

It is the owner's decision, not a technical constraint. The site publishes in five languages and Russian is not one of them. It is checked mechanically — a script flags the four letters that exist in Russian and not in Ukrainian, and every generation prompt names its target language and forbids Russian explicitly.

Can I suggest a term?

Yes — the [questions page](/questions) is read, and questions asked often become entries or notes. Every question asked there is public and answered in the open.

What it is for, in one sentence

If an assistant is going to define your industry's vocabulary to your buyers, the definitions should be written by somebody who had to be right. That is the whole ambition: a reference short enough to be quoted correctly, complete enough to be quoted often, and honest about which sentences a machine wrote.

If you want to see what a retriever would take from a page of your own, Vectorscope shows you the passages. If you want the instruments that produced this vocabulary in the first place, they are all on the free tools shelf.

The product

Watch it on every page, every day

One reading tells you where a page stands today. The product asks the same questions of the same assistants continuously, so a change is something you are told about rather than something you go looking for.

Request an invitation

Invite-only while we keep the readings honest

Author

GetLoopLoop AIAI research system

AI-assisted research, synthesis and measurement by GetLoopLoop.

GetLoopLoop

Pass it on

Read in your language

Opens a browser translation of this English article. The original source stays in English.