Why your architecture diagram changes completely after one small commit
Because the layout is re-solved from scratch every time. Layered graph drawing minimises edge crossings across the whole drawing, so one added node legitimately moves everything else. It is not the renderer being flaky, and no spacing setting will fix it, because spacing is not what is wrong.
We hit this building the map Backthread draws: a picture of a system that agents are changing faster than the team reads it, with the reasoning behind each merged change attached. A map that redraws itself differently every week teaches nobody anything, so we measured it.
What the instability actually is
A layered layout engine runs in phases: break cycles, assign nodes to layers, order the nodes within each layer to reduce crossings, then place and route. The ordering phase is the one that bites. It is a global optimisation over the whole graph, and it is heuristic, so a graph that differs by one node is a different problem with a different answer. The engine is behaving correctly. It was never asked to keep the previous drawing.
That is fine for a diagram you generate once. It is useless for a diagram you regenerate per commit and expect a person to recognise, which is exactly what a map of a codebase under agents has to be. If the reader has to re-find the ingest pipeline every Monday, the map costs more attention than it returns.
The measurement
We took our demo repository, produced eight successive versions of the same system, and rendered the map at each step in both of its grouping modes. The probe is deliberately dull: for every node that exists in both version n and version n+1, did its position change.
Then we ran the identical probe twice. Once with a canonical layout in place — one superset graph laid out a single time, with each version rendered as a subset of that fixed drawing — and once with the canonical layout removed, so each version was laid out on its own.
| Probe | Nodes surviving between versions | Nodes that moved | Share |
|---|---|---|---|
| Each version laid out independently | 26 | 26 | 100% |
| Canonical superset laid out once, versions rendered as subsets | 26 | 0 | 0% |
Zero at every step, in both grouping modes. Not "mostly stable", not "stable if the change is small". The one-off structural fix removed the entire class of movement, because the drawing is no longer being recomputed at all — only masked.
The cost is real and worth stating: you pay for it in a layout that is optimal for the union rather than for any single version, so an individual version is drawn slightly less tightly than it could be. We took that trade without much argument. A reader who can find things beats a drawing that packs well.
The suspect we exonerated first
Before any of that we were sure the problem was spacing. The map was mostly empty — 38 boxes spread over 28 layers, filling 8.1% of a 14,277×15,992 canvas. Every instinct says node size or padding constants.
We measured the padding instead of tuning it. The median gap between boxes was exactly the configured 120 pixels: the engine was doing precisely what it had been told. So the constants were not the cause. The cause was structural — 166 of 261 container pairs pointed at each other. A layered engine assumes a rough acyclic graph, and when most pairs are mutual, cycle-breaking smears a small graph across a large number of layers.
The two fixes are not comparable:
- Tuning the layered engine's spacing and compaction moved canvas fill from 8.1% to 8.6%.
- Replacing the top level with a packing layout moved it to 50.0% — the whole map legible at 3.3 times the zoom.
The second one also fixed stability as a side effect, before the canonical layout existed: under a single added edge, the layered version moved 37 of its 38 boxes and the packed version moved 0 of 38. We had spent days on the wrong suspect, and the measurement that cleared it took an hour.
What to do if you regenerate a diagram per commit
- Measure movement before you tune anything. Count the nodes that exist in both versions and how many changed position. Without that number you cannot tell a fix from a coincidence, and every layout knob feels like it helped.
- Check whether your graph is acyclic enough for a layered engine. Count reciprocal pairs among your top-level containers. If a large share point at each other, you are asking a hierarchy algorithm to draw something that is not a hierarchy, and no configuration recovers that.
- Lay out the union once, render each version as a subset. Positions come from the superset graph and stay fixed; a version that lacks a node hides it rather than re-solving around it.
- Keep the trade visible. Stable positions mean a drawing tuned for the union, not for today. Say so in the interface, or the first person to notice slack space will file it as a bug.
- Treat the diagram as a reading surface, not an artefact. The point is a reader who builds a durable mental picture across weeks. That constraint, not aesthetics, is what decides these calls.
What the numbers do not show
Three limits, honestly. First, this is one repository — ours — at one scale. The largest map we have measured in production is 75 nodes and 222 edges, and the maximum across all 348 versions we have rendered is 77. A graph of several thousand nodes may hit costs in the canonical approach we have not paid yet.
Second, 0 of 26 measures position stability and nothing else. It says the reader will find the same box in the same place. It says nothing about whether the map is a good description of the system, which is a separate problem with separate evidence.
Third, the packing result and the canonical-layout result come from the same programme of work but are not the same experiment, and the 37 of 38 figure predates the canonical layout. We are reporting both because the order in which we found them is the useful part: the structural cause was not where the obvious instinct pointed, twice.
A stable map matters because it is where the rest of the record hangs. The decisions captured from agent sessions, the areas where the reasoning is blank, the question of who on the team actually understands each part — all of it is read against positions a person has learned. That is what Backthread's map of the system is for, and why we spent this long on a drawing that holds still.
Connect one repo and you get the versioned map with positions that stay put between renders, plus the reasoning captured against each area; the trial runs fourteen days.
In short
- Layered layout re-solves the whole drawing, so one node moves everything
- Crossing minimisation is a global heuristic over the entire graph. A graph that differs by one node is a different optimisation problem with a different answer, so per-commit regeneration will keep reshuffling. The engine is not misbehaving.
- Laying out a canonical superset once removed the movement entirely
- Across eight versions of the same repository in two grouping modes, 26 of 26 surviving nodes moved when each version was laid out independently, and 0 of 26 moved when versions were rendered as subsets of one fixed superset layout.
- Spacing was the wrong suspect, and measuring cleared it in an hour
- The map filled 8.1% of its canvas, but the median gap was exactly the configured 120 pixels. The real cause was 166 of 261 container pairs pointing at each other. Tuning the layered engine reached 8.6% fill; switching the top level to packing reached 50.0%.
- Position stability is not the same as a correct map
- The measurement says a reader finds the same box in the same place between versions. Whether the map describes the system well is a different question needing different evidence.
Sources
Backthread shows how much of what your agents built your team really understands. See how it works