A mesh node's identity is the pair (namespace, id). Its path is derived, not stored independently — in Postgres it is literally a generated column:

CREATE TABLE IF NOT EXISTS mesh_nodes (
    namespace       TEXT        NOT NULL DEFAULT '',
    id              TEXT        NOT NULL,
    path            TEXT        GENERATED ALWAYS AS (
                        CASE WHEN namespace = '' THEN id ELSE namespace || '/' || id END
                    ) STORED,
    ...
    PRIMARY KEY (namespace, id)
);

Two consequences follow, and together they are the whole of this page.

An id MAY contain a slash

There is no constraint anywhere forbidding it. Every mesh_nodes DDL — the three Postgres variants, the satellite-table script, mesh_node_history, and the SQLite adapter — declares plain id TEXT NOT NULL. A repo-wide sweep for a SQL CHECK constraint across both src/ trees finds none at all.

Slash-bearing ids are not an accident to be tolerated; several node families depend on them. Every LanguageModel node's id is the provider's wire idz-ai/glm-5.3, anthropic/claude-opus-5, openai/gpt-5.2 — because that string is what the provider's API expects and what the model is known by. The Postgres adapter states the rule in its own words (issue #2212):

🚨 THERE IS NO POSITIONAL (namespace, id) SPLIT OF A PATH — an id may contain '/'.

Splitting a path positionally is path-invariant and key-destroying

Because path is namespace || '/' || id, moving a slash from the id into the namespace leaves the path byte-identical while changing the primary key. These two rows have the same path and different identities:

namespace id generated path
Provider/OpenRouter z-ai/glm-5.3 Provider/OpenRouter/z-ai/glm-5.3
Provider/OpenRouter/z-ai glm-5.3 Provider/OpenRouter/z-ai/glm-5.3

That invariance is what makes the corruption invisible: every path-addressed read, every log line, every URL keeps reading the same. Nothing looks wrong.

The read/write asymmetry that turns it into data loss

The two sides of the adapter address rows differently, and both are correct in isolation:

So a write carrying a re-keyed (namespace, id) finds no conflict and INSERTs a second row whose generated path collides with the first. From that moment:

  1. A read WHERE path = $1 matches two rows and resolves an arbitrary one — not reliably the same one twice.
  2. The versions table (PK (namespace, id, version)) holds two independent chains, so a node with a long history can read back as having none.
  3. A delete WHERE path = $1 removes BOTH rows. One delete, and the node is gone.

What went wrong (#3894)

MeshOperations.SanitizeNodeId split a slash-bearing id at its LAST slash and moved the prefix into the namespace, on the stated premise that "the DB has a CHECK constraint blocking slashes in id". That premise was false, and had presumably always been false. Create and Update both ran through it, so no MCP write could address a flat-keyed slash-id node — every write minted a duplicate.

On the production portal this split Provider/OpenRouter + z-ai/glm-5.3 across two rows on 2026-09-09; the following morning the node was gone from the model list entirely. The MCP create tool's own documentation stated the same false rule ("id — the node's own slug, NO slashes"), so an agent following its instructions produced the duplicate by hand.

Both writes reported "did not land within the confirmation window" — a true negative: the confirmation reads the flat-keyed path while the write had gone to a different primary key.

Patch was never affected. It reads the existing node and writes existing with { … }, so it inherits that node's keying by construction — which is why the split could not be reproduced through patch alone.

The rules

The adjacent failure: a cached miss that outlives its invalidation

Worth knowing when a node "does not exist" right after you created it. A negative read is cached by the storm breaker inside MeshNodeStreamCache (_negative, keyed by path) — not by the per-node hub, and not by the unrelated MessageStormBreaker, which is a per-hub rate breaker with no path negative. An open window fast-fails reads and writes alike, and MeshOperations.Get reports it as Not found, indistinguishable from genuine absence and produced without ever reaching the owner.

A create at that path does clear it: the post-commit MeshChangeEvent.Created publish reaches MeshNodeStreamCache.OnMeshChangeResetFailureState(path), which drops the negative entry and evicts a faulted read entry. That behaviour is pinned by a test.

The remaining race was fixed in #3954. Every read/write that can conclude NotFound now claims the path before opening its owner round-trip. A change event revokes the current claim and clears only the negative entry belonging to that exact generation. Publication is a pair-exact compare-and-swap: an older probe cannot overwrite a newer probe's genuine miss, and its retraction cannot remove that newer entry either. Readers and writers also refuse an entry whose claim is no longer current, so a stale verdict cannot fast-fail even during the small interval before its owner retracts it.

Claims are not a second permanent path cache: a successful or transient probe removes its claim, and teardown removes a pending probe that never reached a terminal, while a genuine miss keeps one only for the lifetime of the existing negative entry. Natural re-probes replace the claim and keep the established exponential backoff; no timer, retry, or sweeper was added. This mirrors PathResolutionService._pendingFills: invalidation is authoritative over work that began in the older failure era.

See also

Reconnecting…
The connection to the server was interrupted. Trying to restore it…
Trying again…
The connection could not be restored. Reloading the page…
The server was updated. Reloading the page to pick up the latest version.