Skip to content

markstay

A specification for giving logical blocks of Markdown a stable identity that survives editing.

markstay is a small, source-native identity layer for Markdown. It gives a logical block of content a stay: a stable address other tools can point at and keep pointing at across edits, moves, and AI rewrites. The address stays put while the content around it changes.

This site states the problem, surveys the prior art, and gives the specification. Version 1 is settled: the marker grammar, attachment model, hashing, and recovery behaviour are fixed, and a reference linter and an evaluation back them with runnable code and measurements. Three cores in Python, JavaScript, and Rust, plus the remark-stay tree adapter, back the spec with shared conformance tests. Version 1.1 adds optional CommonMark-tree attachment, a backward-compatible refinement so a loose list or a blank-line-containing fence can carry a single stay. Version 1.2 excludes leading YAML frontmatter from segmentation, so editing status: is not a content edit. Version 1.3 lets a direct list item carry its own stay, addressed inside its list rather than as a block of its own, version 1.4 names the two ways quote recovery can refuse an attachment, version 1.5 makes text inside a fenced code block content rather than markup, so a document can show a marker without acquiring one, and version 1.6 gives a GFM table body row its own stay on the same key and the same ladder as a list item. Version 1.7 defines where writers may insert a marker: on its own line for a block, or at a list-item or row carrier that passes §3.4's lexical checks.

Status: version 1.7, settled

The surface is small and stable. It is also young: real-world use and critique will shape later versions. Write safety is a placement rule, not a guarantee that arbitrary Markdown renders identically after stamping. A child carrier is refused when its container's source prefix or the marker itself fails the §3.4 checks; other children and the container can still receive stays. Child segmentation remains optional, and every reader must keep a subhash marker from identifying its container. See compatibility for the measured limits. Inline spans stay deferred. Issues and counter-arguments are welcome.

The problem

Markdown is now the default storage format for READMEs, documentation sites, issues, pull-request discussion, model prompts, and the output of AI agents that maintain documents over time.

It has stable handles at every level except one:

  • Files have paths.
  • Revisions have commit SHAs.
  • Headings have generated anchors.
  • Logical blocks (a specific paragraph, list, table, or code fence) have nothing.

There is no portable way to say "this paragraph" and have the reference still hold after the block is edited, moved, or regenerated. Heading anchors drift when the heading text changes. Line numbers break on the first insertion above them. Quote matching fails on repeated or rewritten text. The handles that exist all encode location, and location is exactly what editing destroys.

Why it matters now

Agents increasingly work like repository collaborators: they read a file, change one section, and open a pull request. In code they have a rich identity substrate (paths, symbols, AST nodes, line ranges anchored to commits). In Markdown prose they fall back to fuzzy text descriptions:

the second paragraph under "Limitations"
the bullet about retries
the code block after the warning

Insert a paragraph, split a list, or add another warning, and each of those can silently point at the wrong block. A stable per-block id turns "the bullet about retries" into a handle an agent can store, cite, and edit against, and lets the edit be audited afterwards (which ids changed, which were dropped).

This is the case markstay is built around. It is also where the idea is most fragile, because the same agents that would use the ids are the ones most likely to destroy them during a rewrite. That risk was measured, not assumed: see the findings on whether markers survive an edit at all, and the public dogfood case study on what keeps section ids stable in practice (the preservation instruction, more than the post-edit linter).

markstay in one example

A block carries a marker on the line after it. The canonical form is a trailing HTML comment, invisible in rendered Markdown and preserved in the source:

Users authenticate with an API key in the Authorization header.
<!-- stay:8f24 -->

Identity can travel with optional evidence for detecting and recovering from drift:

Users authenticate with an API key in the Authorization header.
<!-- stay:8f24 hash=sha256:7a9c quote="Users authenticate with an API key" -->

That is the whole surface area. Everything else is rules about what the marker binds to, how it is preserved, and how a tool recovers when it is lost. See the examples for lists, code fences, tables, the MDX profile, and agent edit requests.

The one idea everything hangs on: identity is not location

Every mature system that anchors comments to editable content (Google Docs, Notion, Figma, GitHub review comments) separates the stable identity of a thing from evidence about where it currently sits. markstay takes the same split:

Field Role Changes when content changes?
id stable logical identity no
hash drift detection (did this block's body change?) yes
quote + prefix/suffix recovery evidence to re-find a detached marker n/a

The id is the identity. The hash and quote are never identity, only evidence. A hash mismatch with the marker still present means "same block, changed content", not "new block". This split is the single most consistent lesson from the prior art.

What version 1 settles

The specification fixes:

  • Marker grammar: a positional stay: id plus free-order key=value attributes, HTML-comment form for .md and a JSX-comment profile for MDX.
  • Attachment: markers bind to the block above them at whole-block granularity (whole list, fence, table, quote), over blank-line-delimited blocks by default and the CommonMark block tree under the version 1.1 refinement.
  • Hashing: an exact normalization rule, so two implementations agree on drift.
  • Recovery: a TextQuoteSelector-style ladder (marker → hash → quote) that surfaces a detached marker as outdated rather than guessing.
  • The AI editing contract: what an agent must do to preserve stays, made measurable by the post-edit linter.

Version 1.1 adds CommonMark-tree attachment, so a loose list or a blank-line-containing code fence can carry a single stay (an optional extra; the dependency-free blank-line path stays the default). Versions 1.3 and 1.6 add optional list-item and table-row identity; version 1.7 constrains their marker placement. Inline-span identity remains deferred. The FAQ covers the obvious objections (why not heading anchors, UUIDs, an external database, or HTML ids).

Scope

markstay is deliberately one thing: identity and recovery. Annotation, transclusion, AI editing, and cross-references are consumers that build on the stay layer, not part of it. Keeping the core small is what lets other tools adopt it.

The work is openly licensed (content under CC BY 4.0, code samples under MIT) and lives at markstaymd/markstay. Issues and counter-arguments are welcome.