Coordinating on descriptions such as “the person in the red jacket near the jetty”, rather than on pixels or point clouds. Robots with different sensors (a drone’s camera, a ground robot’s lidar, a boat’s sonar) need a shared representation to agree on what they observe.
3D scene graphs are one such representation: objects, places and rooms as nodes, and their relations as edges. There’s an interactive explainer on 3D scene graphs with their history, how they’re built, and what changes on land, at sea and in the air.
Questions
- What does it mean for a semantic claim from another robot to be verifiable?
- Can resilient consensus work on labels and descriptions, or does everything have to be squashed into vectors first?
- Where do multimodal foundation models help, and where do they fail?
Related: multi-robot systems, human-robot trust.