The map
that knows
what it is
looking at.
Seven years ago a 3D scene graph was a way to annotate a building scan. It is now the representation robotics is betting on for perception, planning and long-term memory. It still cannot tell you when to forget something.
Six pieces, and what breaks when you remove one
The survey's contribution starts by ending an argument: every 3D scene graph in the literature is the same tuple, and the variants differ only in which pieces they fill in. Switch a piece off and see what representation you are left with.
Seven years, five currents
Milestones by theme. Larger dots are the papers that changed what everyone else did next;
click any dot for why it mattered. Every entry here has its LaTeX source in
dl_papers/sources/.
The layers, and what they are called elsewhere
The hierarchy is a labelling function assigning each node a level. Indoors it maps onto architecture almost for free. Switch the environment and watch the structure survive while every layer's meaning is renegotiated. That renegotiation is one of the field's live open problems.
Search gets cheap. Finding "a mug" in a flat map is a scan over every object. In a hierarchy you rule out three floors and eleven rooms before looking at a single mug.
Planning gets tractable. Taskography's result was that classical planners choke on full-size graphs and recover once the graph is sparsified along its structure.
Inference has a bound. Hughes et al. showed indoor scene graphs have low treewidth, giving a structural reason why reasoning over them is fast.
Corrections propagate. When a loop closure fires, Hydra re-optimises the hierarchy along with the trajectory. Rooms move because the mesh moved.
The layer people skip when explaining this, and the one that makes it actionable. Places hold free space: a sparse graph over where the robot could stand, typically a generalised Voronoi diagram or a clustering of the ESDF.
That is what turns a description into a plan. "The mug is in the kitchen" is semantics. "…and here are the 14 poses connecting you to it" is navigation.
In the aerial case it stops being a floor plan at all: free space is a volume and the node becomes a viewpoint, meaning a pose you can see the target from. Same layer, different physics.
Five decisions make a system
Every pipeline in the literature is a point in the same five-dimensional design space. Pick your options and the panel names the systems closest to what you just described.
Batch is a different problem than online
Vision-side work assumes a finished reconstruction and asks what is in it. Robotics gets one frame at a time from a drifting sensor and has to commit. Below is what pose drift actually costs, measured, not simulated.
When does merging help?
One robot cannot cover a building, a harbour or a bridge span. Merging K partial graphs is not free: it only pays where robots co-observe, because agreement is the only thing that lets one robot's error be caught by another. This is the measured crossover.
The same structure, three different physics
The survey says it is an indoor-and-urban document. Held against marine and aerial deployment, the gaps are specific and nameable, and one of the three is close to open ground.
What the graph is asked to do
Planning, navigation, manipulation, question answering. The dominant interface today is blunt: serialise the graph to JSON, hand it to a language model. Try a query and watch where that breaks.
Two camps. LLM-native: SayPlan searches the graph semantically, then refines to a path. Fast and general, though it forfeits completeness and soundness. Formal: translate language into PDDL or LTL and hand it to a real planner, which keeps the guarantees.
Exactly one paper does joint task and motion planning on a scene graph. The field has barely started here.
Mostly the graph localises the target and then is ignored. FunGraph and its relatives add the layer that helps, with handles, knobs and buttons as nodes, so an instruction resolves to a graspable part.
The open question is VLA integration: a policy conditioned on pixels alone cannot know the bottle is full. The graph remembers that it is.
Grounding referring expressions, multi-hop question answering, cross-modal retrieval. The hierarchy is what makes "the light in the bathroom on the second floor" resolvable at all, since it needs three layers at once.
Cost is the constraint. Everything here is a patch on context length.
The most systemic problem in the field
Graph quality and task success are measured separately, usually on different datasets. So the question a roboticist has, how much does this perception error cost me downstream, goes unanswered by any standard protocol.
What nobody can do yet
Drawn from the survey's own per-section challenge blocks. Filter by who is blocked. The same gap reads differently to a perception person and a planning person.
Consensus over a fleet of imperfect graphs
Hydra-Multi merges robots by geometry: relative frames and loop closures. The open question one layer up is trust: given K graphs that disagree, which claims survive?
The claim
Given K robots that each mapped part of a space, produce one merged graph that is more trustworthy than naively unioning their outputs, using disagreement and informed silence (a peer was positioned to see it and did not report it) as evidence, plus physical plausibility to reject floating and interpenetrating objects.
A single-robot system never had to do this, because it never merged. That part of the problem is real.
Status
Proven offline on new scenes. Run on spaces with no prior graphs of any kind, with the reference built only afterwards for scoring. On a multi-room apartment no single robot could cover, fusing beat every individual robot on both the object and the hierarchy layer.
Unproven live. Fusion is still batch: it merges K finished runs. Real deployment wants the merged graph growing as the robots map. That is the one structural gap left, and it is the same gap the survey names as the divide between the vision and robotics halves of the field.
Conditional, and said so. Where robots do not co-observe, consensus loses to naive merging, because there is no second opinion to be had. The regime map in §06 is the finding itself.
Partly. Virtual fleets cut from one recording share a camera, so detector errors are correlated and the independence assumption is flattered. The real-hardware arena run is the answer, and it already exposed a genuine defect: coverage-based silence deletes objects a peer drove past but never detected.
For the recovery numbers, yes, and it should be labelled. The decoupled result is the one scored against simulator ground truth, and only that table should carry weight.
Bandwidth and frames. Graphs are ~100× smaller than the clouds they summarise and do not need a shared real-time frame. This is the argument that gets much stronger underwater.
Real-arena run: precision 1.000, recall 0.64. Coverage-based negative evidence is too aggressive Charging silence by viewing opportunity instead of cell membership is the fix, already working on one path and not yet wired through the other.
Odometry drift of 1.69 m with no loop closure smears revisited surfaces; the surface ballot then drops most vertices as single-witness.
The fused room layer over-segments badly, at 47 rooms on the arena. The agglomeration threshold is unfitted, which is exactly the "abstraction layers are hard-coded" challenge in §10 showing up as a number.