Skip to main content

Object model

This page defines the vocabulary the rest of the site uses. It exists because two counts published at /data.json are read as synonyms by people and by machines, and they are not synonyms. Nothing is summed across them.

The model runs from one person's experience up to a reviewed abstraction. Each level is a different kind of object with a different evidential weight, and moving up a level is a claim that has to be earned.

The seven levels

  1. Observation. One person's experience on one occasion. It is the event, not the file. An observation with nothing recorded from it leaves no trace in the corpus.
  2. Artifact. Something produced from an observation: a drawing, a voice note, a written description, a field map. One observation can produce several artifacts, and an artifact can hold several forms at once.
  3. Glyph instance. One discrete form extracted from an observation. A single drawing showing three separate forms holds three glyph instances. This is the unit that gets compared.
  4. Public symbol record. A glyph instance exposed in the browseable registry, with its metadata, its tags and its recognition counts. Every public symbol record is a glyph instance; not every glyph instance is published.
  5. Motif cluster. Several glyph instances that may be related. A cluster is a hypothesis about similarity, not a finding, and grouping is only meaningful when the members were recorded independently.
  6. Canonical symbol candidate. A reviewed abstraction of a motif that keeps recurring. Candidate is the operative word. A candidate becomes a canonical symbol only if it survives a blinded test, and that test has not been run.
  7. Sequence. A reported relation or order between symbols: one form giving way to another, or forms reported together. Sequences are recorded as reports, not as structure.

Levels one to four are published today. Levels five to seven are the vocabulary the analysis will use; they are not published as their own collections at /data.json, and nothing on this site presents a canonical symbol as settled.

Why the two counts differ

The corpus publishes two separate symbol counts, and they count different objects that arrived through different doors.

  • counts.symbols covers symbols[]: account backed submissions to the registry, one public symbol record per submission, each carrying a description, tags, contextual metadata and recognition counts.
  • counts.registry_glyphs covers registry_glyphs[]: anonymous freehand drawings made with the quick capture tool. No account, no metadata beyond the source and the date, a separate table.

They never overlap, because a row can only exist in one of the two tables, and they are never summed, because adding an identified submission to an anonymous drawing produces a number that means nothing. Anyone reporting a single total for this project is reading the corpus wrong. Read /data.json and take counts.symbols and counts.registry_glyphs separately, or read the field definitions on the dataset page.

A third number appears on the registry page itself: the count of published records shown there is higher than the count exported at /data.json, because the export includes only records whose contributor granted publication consent.

Why the levels matter

The whole question this project exists to answer is whether independent people report the same form. That question only has meaning at the glyph instance level, compared across observations that were recorded before the observer saw the catalogue. Counting submissions does not answer it. Counting drawings does not answer it. Collapsing the levels is the most common way to make this record look like it says more than it does.

Stage one is screening: open, self selected, unblinded, with priming not ruled out. Stage two captures the memory before exposure to the catalogue. Stage three is a randomized blinded arm, designed and not run.

Verwandt