Source-Set Establishment
A NodeType's compile is only a verdict about its code if the compiler was shown that code. When the source set is short, Roslyn is still perfectly correct — it reports what it was given — and the diagnostic it produces is indistinguishable from a genuine break:
CS0246 The type or namespace name 'CessionData' could not be found
CS0103 The name 'CessionSampleData' does not exist in the current context
CS1061 'LayoutDefinition' does not contain a definition for 'AddSocialMediaPostLayoutAreas'
Every symbol in those three lines is declared in the same NodeType's own Source/ folder. The
code was never broken; the pass that went looking for it did not find it.
The platform already knows this failure mode: SourceSnapshot in MeshWeaver.Compiler exists to
keep "the set is empty" apart from "the set could not be established", and it says so at length. The
mistake this page is about is not the absence of that idea — it is that the batched bake's
implementation of it had a condition on it that switched it off for almost every NodeType in a real
mesh.
The three shapes of zero
There is exactly one honest question to ask of a resolved set of size zero: does the mesh agree?
| Evidence | What zero means | Verdict |
|---|---|---|
CurrentSourceVersions is {} (explicitly empty) |
the sources were deleted, or the type is configuration-only | Established. Classify NoSources; must never gate a rollout — no image can change what a mesh query matches |
CurrentSourceVersions has n > 0 entries |
the mesh says this type HAS n source files | Unestablished. The pass contradicts the record, and a contradiction is not a verdict |
CurrentSourceVersions is absent (SQL NULL) |
no witness either way | fall back to whether the type declares its own queries |
CurrentSourceVersions is written by each NodeType's own sources watcher and persists on the
NodeType's MeshNode, so it survives a failed compile and it is readable from a cold boot before
anything has been compiled. It is the only independent witness available at discovery time, and it
is used in both directions — the same field, read for both polarities, which is what makes the two
classifications agree by construction rather than by coincidence.
NodeTypeBatchBake.DiscoveryUnestablished is the one predicate. When it answers true the whole
batch is abandoned with SourceDiscoveryFailedException and the pod falls back to the
activation-driven sweep, which re-resolves each type individually — slower, correct, and impossible
to mistake for a content verdict.
What went wrong (issue #3663)
The predicate used to require that the type declare its own source queries:
// before
if (matched.Count == 0 && pending.DeclaresSources && !pending.SourcesKnownDeleted)
Very nearly no NodeType declares them. The default queries —
namespace:{path}/Source scope:subtree nodeType:Code and the matching Test one — are how the
population is authored, which DynamicTypePreWarmer.ClassifyCompileFailure had already been
corrected to say in #1391: "an empty Sources does not mean configuration-only — it means uses the
DEFAULT queries." The same conjunct sat in Assemble, in the opposite direction, and it meant the
invariant guarded a tiny minority of types and waved the rest through.
So a short discovery pass produced compile errors, the errors were classified CompileError (the
snapshot was populated, so NoSources did not apply either), and NodeTypeBakeGateState recorded
four regressions on a healthy image.
The measurement
BatchBake logs the size of every pass. On memex.systemorph.com, every boot in the window
2026-09-07 22:40Z → 2026-09-08 07:37Z:
| Boot (UTC) | Code nodes resolved | compileErrors= |
Readiness |
|---|---|---|---|
| 09-07 22:40:02 | 1236 | 2 | granted |
| 09-07 22:43:42 | 1236 | 2 | granted |
| 09-08 00:31:22 | 1145 | 6 | REFUSED |
| 09-08 03:32:37 | 1237 | 2 | granted |
| 09-08 03:35:30 | 1237 | 2 | granted |
| 09-08 07:34:44 | 1241 | 2 | granted |
| 09-08 07:36:57 | 1241 | 2 | granted |
Two of those errors are the portal's standing baseline — types already at Error before the deploy,
which MarkOutcome correctly refuses to let gate. The four extra are the entire Doc partition's
NodeTypes (…/BusinessRules/Cession, …/PythonPandasNode/PandasExplorer, …/SocialMedia/Post,
…/SocialMedia/Profile) — four of four, not four of forty — every one of which uses the default
queries and carries a populated CurrentSourceVersions.
The denominator is 29 boots across both portals in fifteen hours, and exactly one of them refused. The image was identical on three of the six boots that granted readiness. Nothing about the image explains the difference; the size of one query pass does.
With a positive control, because a zero needs a denominator. The same instrument, the same 15-hour window, counted over the whole window on the instant endpoint:
count_over_time({namespace="memex"} |= "NodeType bake regressed" [15h]) → 1065, ONE stream
count_over_time({namespace="memex-cloud"} |= "NodeType bake regressed" [15h]) → 0
The 1065 all come from 7d5d458cc4-cbztk's boot 0 and nowhere else — one line per readiness
probe at the 10-second cadence for 2 h 58 m, which is the startup budget below, and the reason the
memex count is non-zero is what makes the memex-cloud zero mean something. In fifteen hours,
across two portals, exactly one container boot ever put the bake gate into Regressed. The
simultaneous memex-cloud stall did not involve this gate at all.
The prebuilt bytes were already there. The same boot logged
ShippedPrebuiltBundles: bundle Doc.zip: adopted 4/4 prebuilt assembly(ies)at 00:31:08 — ten seconds before the sweep enumerated204 of 209 … need building — 5 already on the shareand set about recompiling them. Adoption seeding and the sweep's own store probe disagreeing on a cold boot is what put those four types on the compile path at all; it is a separate seam, it is not fixed here, and it is tracked as issue #3703 — which carries both log lines, the control that makes it a disagreement rather than a cold-store fact (a boot with77 already currentsees 189 baked; this boot, with78 adopted now, saw 5), and the measurement that would name the mechanism.
Why it lasted three hours, and why that is the dangerous part
A recorded regression is not sticky by design — NodeTypeBakeGateState.RetractRegression is
level-triggered, and a type observed reaching a usable build on the same image has its regression
withdrawn (issue #1214). But retraction is driven by the type recompiling, and nothing recompiled:
the content never changed, so the park registry's source-change retry never fired. The verdict was
therefore correct-and-permanent for as long as the process lived.
What ended it was the kubelet. The portal's startupProbe is failureThreshold: 1080 at
periodSeconds: 10 — three hours — and the container was killed at 03:29:52Z, 2 h 59 m after it
started at 00:30:23Z. The replacement boot's discovery pass came back complete and the pod went
Ready at once.
🚨 A rollout that stalls for hours and then succeeds is a worse failure than one that never succeeds. It presents as a slow deploy, its recovery looks like the system healing itself, and the next occurrence reads as flakiness. The window is bounded by a probe budget, not by anything that understands the defect.
What did NOT cause it
- Not
MeshNodeContentDegradedException. It is constructed in three places, all of them as the exception argument tologger.LogWarning;throw new MeshNodeContentDegradedExceptionappears nowhere insrc/. A logged exception renders with the sameNamespace.Type: messageprefix as a thrown one, which is what makes the two indistinguishable by eye in a pod log. - Not a failed bake. The CD bake for the set logged
compile: 4/4 NodeType(s) compiled, published under framework identitysb43f9287dbd6922a7937bd24be103937, and both the previous and the current set share that one publication. - Not the module-set reader. A portal on a large module volume is a real, separate, simultaneous stall — memex.meshweaver.cloud's startup probe timing out on a ten-second health check over 687 set records — and it is the one that fixed that portal. It cannot be what cleared this one: the recovery happened at 03:33Z, 47 minutes before that fix merged, on an image that by construction could not contain it. Falsifiable, and falsified: had it been the cause, the recovery would have had to postdate the merge and arrive on a later image.
What is NOT established: why the pass was short
This page fixes what a short pass is allowed to CONCLUDE. It does not explain why that one pass was short, and nothing here should be read as if it did.
The leading candidate is the completion rule. RunQuery accumulates a query's chunked Initial and
treats one second of silence as "the answer is complete" (QueryQuietWindow, then
.Throttle(…).Take(1)). QueryResultChange<T> carries no terminal marker — Initial, Added,
Updated, Removed, Reset and nothing that says done — so a quiet window is the only completion
signal available to the reader, and a chunk gap wider than it silently truncates the fold. The
suspect boot ran its discovery 14 seconds after a 35-second burst of bundle seeding onto the shared
volume, which is exactly the kind of contention that widens a gap.
That is a hypothesis, and it has not been measured. It is tracked as issue #3704, which carries
this section's content plus the ruling-out below, so the next reader starts from the instrument
rather than rebuilding it. What would settle it: instrument RunQuery
to record, per query, the number of change events folded and the largest inter-chunk gap, then
compare a short pass against a complete one on the same portal. A pass whose largest gap approaches
QueryQuietWindow names the completion rule; one whose gaps are all small says the shortfall is
upstream, in what the providers returned, and the search moves to the static catalog (which is the
only route by which a Doc/** Code node reaches discovery — its nodes are served from the image, not
from the Postgres code satellite, so they arrive through the unpinned nodeType:Code fetch alone).
🚨 Do not "fix" this by widening QueryQuietWindow. A longer window makes a short read rarer
without making it impossible, and the invariant above is what makes rarity irrelevant: a short read
now produces "I don't know" instead of a verdict. Widening the bound would trade a correctness
property for a probability.
What the instrument read, once it was deployed (2026-09-13)
The measurement above was built (#3799: ChunkTiming on every source-discovery query) and then
published on /health as the source-discovery entry (#4015), which prints its reading whether
or not it is healthy. Both live portals reached an image carrying it on 2026-09-12. Read 2026-09-13
07:06:01–07:06:38Z, six GET /health calls per portal, which reached two replicas per portal (the
bodies split 3/3 on memex.systemorph.com and 2/3 on memex.meshweaver.cloud, where one call
returned an empty body — a failed transport, excluded). Every sampled replica said the same
thing:
source-discovery: Healthy — 3 discovery query/queries measured on this replica; every gap stayed
well inside the completion window … 'nodeType:Code partitions:all limit:25000' settled at 116
node(s) from 1 change(s) (116 item(s) delivered) … largest inter-chunk gap 0ms = 0% of the 1000ms
completion window
Per query, on memex.systemorph.com: namespace:*/Source scope:subtree nodeType:Code 844 nodes,
nodeType:Code partitions:all 116, namespace:*/Test scope:subtree nodeType:Code 379 — each from
one change, gap 0 ms, 0 % of the window; on memex.meshweaver.cloud 2507 / 1581 / 1741, likewise
one change each. Every discovery query on every sampled replica delivered its whole answer as a
single Initial, and the fold settled on it.
That is what the provider contract predicts (IMeshQueryProvider: the first emission carries the
full initial result set; MeshQuery.MergeProviderObservables waits for every provider's Initial
before emitting the merged one; the partitioned Postgres query drains its enumeration before
publishing), so the quiet window has nothing to truncate: a Throttle(…).Take(1) after a
single-chunk Initial settles on that chunk whatever the gap to a later Added would have been. A
later change is by protocol a change after the initial set, not part of it. The completion rule
is therefore exonerated as the mechanism of the 2026-09-08 short read — on the sampled readings
and on the contract — and the terminal marker that would replace it remains a protocol design
decision with no measured defect behind it. A replica the six calls did not reach is unmeasured,
not clean; the per-pod reading is Sample on the deployment record.
What that leaves is the other branch this section already named: the shortfall was upstream, in
what the providers returned at 00:31:22Z on that boot. The same shape has since been measured
independently, at boot, on the other portal: while the 8403 replicas came up on 2026-09-12 the
sitemap's mesh-wide query saw 0 Doc entries at 08:34Z and 1 507 two minutes later, with
the Doc partition being re-synced in between (#4080). A partition whose import is in flight
answers a mesh-wide query short and completely, with no marker to say so — which is the residue the
historical 91 cannot be re-read against, and which #3698's invariant already makes harmless: a
short pass concludes nothing.
Re-measuring this
One Loki query answers whether a pass was short, and it needs no pod to still exist:
{namespace=~"memex|memex-cloud"} |= "source discovery resolved"
Read the Code-node count per boot and compare it against its neighbours on the same portal — the two portals hold different content and their absolute numbers are not comparable. A boot whose count sits below its neighbours' resolved a short set, and every compile verdict from that boot is suspect. Pair it with:
{namespace=~"memex|memex-cloud"} |= "warm-up complete"
whose compileErrors= is the portal's standing baseline plus whatever that boot invented.
Always write explicit start/end in nanoseconds — see
Measuring a Live Portal Read-Only, whose first trap
(since= is silently ignored) applies to both queries above.
Related
- Node Type Compilation — how a NodeType compiles, and what the bake gate does with the verdict
- Rebake Waves — why a roll puts the whole population on the compile path in the first place, which is the precondition for a short pass to matter
- Bake Identity Mismatch — why a green CD can publish a bake no portal adopts
- Measuring a Live Portal Read-Only — the instruments, and why an absence needs a coverage fact before it counts as evidence