A rundown of the field · 2019–2026

The map
that knows
what it is
looking at.

Seven years ago a 3D scene graph was a way to annotate a building scan. It is now the representation robotics is betting on for perception, planning and long-term memory. It still cannot tell you when to forget something.

34papers on the timeline
150+catalogued at 3dscenegraphs.com
5layers, 3 domains
Fig. 00 · hierarchical scene graph, isometric –
01 · The definition

Six pieces, and what breaks when you remove one

The survey's contribution starts by ending an argument: every 3D scene graph in the literature is the same tuple, and the variants differ only in which pieces they fill in. Switch a piece off and see what representation you are left with.

Eq. 1 · the tuple
Live · what remains
You now have

Why this matters for a robot. The grounding Vg is the whole argument for the representation. Language models already give you semantics; SLAM already gives you geometry. The scene graph's claim is that binding them, so every symbol sits at a metric location, is what makes the map actionable.
02 · History

Seven years, five currents

Milestones by theme. Larger dots are the papers that changed what everyone else did next; click any dot for why it mattered. Every entry here has its LaTeX source in dl_papers/sources/.

Fig. 02 · milestone papers by theme and date
2019Annotation A scene graph is a way to label a building scan. Built semi-automatically, consumed by humans.
2021–2022Perception Kimera and Hydra derive the whole hierarchy live from a moving sensor. The graph becomes a SLAM output.
2023Interface Open vocabulary removes the fixed label set; SayPlan notices the graph serialises to JSON and an LLM can read it.
2024–2026Substrate Task-driven granularity, scene history, real shape, physical plausibility. The graph becomes something you compute with.
03 · Anatomy

The layers, and what they are called elsewhere

The hierarchy is a labelling function assigning each node a level. Indoors it maps onto architecture almost for free. Switch the environment and watch the structure survive while every layer's meaning is renegotiated. That renegotiation is one of the field's live open problems.

Fig. 03 · layer explorer
Layer separation 1.00
Rotation 32°
Layers · click to toggle
Edges
Why the hierarchy earns its keep

Search gets cheap. Finding "a mug" in a flat map is a scan over every object. In a hierarchy you rule out three floors and eleven rooms before looking at a single mug.

Planning gets tractable. Taskography's result was that classical planners choke on full-size graphs and recover once the graph is sparsified along its structure.

Inference has a bound. Hughes et al. showed indoor scene graphs have low treewidth, giving a structural reason why reasoning over them is fast.

Corrections propagate. When a loop closure fires, Hydra re-optimises the hierarchy along with the trajectory. Rooms move because the mesh moved.

The places layer, specifically

The layer people skip when explaining this, and the one that makes it actionable. Places hold free space: a sparse graph over where the robot could stand, typically a generalised Voronoi diagram or a clustering of the ESDF.

That is what turns a description into a plan. "The mug is in the kitchen" is semantics. "…and here are the 14 poses connecting you to it" is navigation.

In the aerial case it stops being a floor plan at all: free space is a volume and the node becomes a viewpoint, meaning a pose you can see the target from. Same layer, different physics.

04 · Construction

Five decisions make a system

Every pipeline in the literature is a point in the same five-dimensional design space. Pick your options and the panel names the systems closest to what you just described.

Fig. 04 · construction design space
Closest systems
05 · The central divide

Batch is a different problem than online

Vision-side work assumes a finished reconstruction and asks what is in it. Robotics gets one frame at a time from a drifting sensor and has to commit. Below is what pose drift actually costs, measured, not simulated.

Fig. 05 · measured drift ablation
Reading

What the measurement says. Drift inflates the map as well as blurring it. As pose error grows, the same physical object observed before and after a drifting loop lands in two places and data association has no reason to merge them, so node count climbs against a fixed reference while precision falls. Switch the panel to node count to see it directly. Loop closure is what lets the hierarchy be retroactively corrected, which is the property Hydra is built around.
06 · Many robots

When does merging help?

One robot cannot cover a building, a harbour or a bridge span. Merging K partial graphs is not free: it only pays where robots co-observe, because agreement is the only thing that lets one robot's error be caught by another. This is the measured crossover.

Fig. 06 · measured regime map
Reading

Read the crossover, and read the caveat next to it. Consensus starts below plain geometric merging and overtakes it as corroboration rises, on all three TUM sequences. That is the finding: the value is conditional on the deployment. The fleets here are synthetic, cut as overlapping windows from one recording, so detector errors are correlated across the "robots" and the independence assumption is flattered. The uHumans2 column behaves differently again, which is why the claim is stated for the TUM sequences and not in general.
07 · Ground, marine, aerial

The same structure, three different physics

The survey says it is an indoor-and-urban document. Held against marine and aerial deployment, the gaps are specific and nameable, and one of the three is close to open ground.

Tab. 07 · what changes
The marine observation worth making out loud. Nothing in the 3D scene graph literature targets underwater. Yet DRACo-SLAM2, which represents sonar maps as object graphs and matches them between vehicles to close loops over an acoustic link, arrived at graph-level map sharing independently, for the reason that matters most here: bandwidth. A graph is roughly two orders of magnitude smaller than the point cloud it summarises. On land that is a convenience. Underwater, where the link is kilobits per second, it is the only option there has ever been.
08 · Applications

What the graph is asked to do

Planning, navigation, manipulation, question answering. The dominant interface today is blunt: serialise the graph to JSON, hand it to a language model. Try a query and watch where that breaks.

Fig. 08 · serialised graph as LLM context
Ask the scene
Scene size
4 floors
Planning & navigation

Two camps. LLM-native: SayPlan searches the graph semantically, then refines to a path. Fast and general, though it forfeits completeness and soundness. Formal: translate language into PDDL or LTL and hand it to a real planner, which keeps the guarantees.

Exactly one paper does joint task and motion planning on a scene graph. The field has barely started here.

Manipulation

Mostly the graph localises the target and then is ignored. FunGraph and its relatives add the layer that helps, with handles, knobs and buttons as nodes, so an instruction resolves to a graspable part.

The open question is VLA integration: a policy conditioned on pixels alone cannot know the bottle is full. The graph remembers that it is.

Scene understanding

Grounding referring expressions, multi-hop question answering, cross-modal retrieval. The hierarchy is what makes "the light in the bathroom on the second floor" resolvable at all, since it needs three layers at once.

Cost is the constraint. Everything here is a patch on context length.

09 · Evaluation

The most systemic problem in the field

Graph quality and task success are measured separately, usually on different datasets. So the question a roboticist has, how much does this perception error cost me downstream, goes unanswered by any standard protocol.

Tab. 09 · what each metric sees, and what it cannot
The thing to name in a review. If the reference graph is produced by the same pipeline family being evaluated, a high recovery score only measures self-consistency. The defensible move is a reference built from independent ground truth, such as simulator segmentation and poses or hand annotation, plus a plain statement of which numbers carry the coupling.
10 · Open challenges

What nobody can do yet

Drawn from the survey's own per-section challenge blocks. Filter by who is blocked. The same gap reads differently to a perception person and a planning person.

11 · Where this project sits

Consensus over a fleet of imperfect graphs

Hydra-Multi merges robots by geometry: relative frames and loop closures. The open question one layer up is trust: given K graphs that disagree, which claims survive?

The claim

Given K robots that each mapped part of a space, produce one merged graph that is more trustworthy than naively unioning their outputs, using disagreement and informed silence (a peer was positioned to see it and did not report it) as evidence, plus physical plausibility to reject floating and interpenetrating objects.

A single-robot system never had to do this, because it never merged. That part of the problem is real.

Status

Proven offline on new scenes. Run on spaces with no prior graphs of any kind, with the reference built only afterwards for scoring. On a multi-room apartment no single robot could cover, fusing beat every individual robot on both the object and the hierarchy layer.

Unproven live. Fusion is still batch: it merges K finished runs. Real deployment wants the merged graph growing as the robots map. That is the one structural gap left, and it is the same gap the survey names as the divide between the vision and robotics halves of the field.

Conditional, and said so. Where robots do not co-observe, consensus loses to naive merging, because there is no second opinion to be had. The regime map in §06 is the finding itself.

The three questions to be ready for
"Isn't the fleet synthetic?"
Partly. Virtual fleets cut from one recording share a camera, so detector errors are correlated and the independence assumption is flattered. The real-hardware arena run is the answer, and it already exposed a genuine defect: coverage-based silence deletes objects a peer drove past but never detected.
"Isn't the reference circular?"
For the recovery numbers, yes, and it should be labelled. The decoupled result is the one scored against simulator ground truth, and only that table should carry weight.
"Why not just centralise the raw data?"
Bandwidth and frames. Graphs are ~100× smaller than the clouds they summarise and do not need a shared real-time frame. This is the argument that gets much stronger underwater.
Measured defects, currently open

Real-arena run: precision 1.000, recall 0.64. Coverage-based negative evidence is too aggressive Charging silence by viewing opportunity instead of cell membership is the fix, already working on one path and not yet wired through the other.

Odometry drift of 1.69 m with no loop closure smears revisited surfaces; the surface ballot then drops most vertices as single-witness.

The fused room layer over-segments badly, at 47 rooms on the arena. The agglomeration threshold is unfitted, which is exactly the "abstraction layers are hard-coded" challenge in §10 showing up as a number.

The one-sentence version. The field has proven it can build these graphs online, in the open vocabulary, on one robot. It has not shown how several robots come to agree on one, how a graph earns the right to forget, or that any of it works where the water is murky and the link is kilobits. All three come down to the same thing: deciding what a map is entitled to claim.