MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
arXiv.orgEmbodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training.
Xun Huang +8 more
semanticscholar +1 more source
Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph
IEEE International Conference on Robotics and AutomationLocating objects described in natural language presents a significant challenge for autonomous agents. Existing CLIP-based open-vocabulary methods successfully perform 3D object grounding with simple (bare) queries, but cannot cope with ambiguous ...
Sergey Linok +6 more
semanticscholar +1 more source
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
AAAI Conference on Artificial IntelligenceOpen-world 3D scene understanding is fundamentally challenging for vision and robotics, due to the constraints of closed-vocabulary supervision and static annotations.
Fei Yu +4 more
semanticscholar +1 more source
CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learning
Computer Vision and Pattern Recognition3D Scene Graph Generation (3DSGG) aims to classify objects and their predicates within 3D point cloud scenes. However, current 3DSGG methods struggle with two main challenges. 1) The dependency on labor-intensive ground-truth annotations.
Lianggangxu Chen +5 more
semanticscholar +1 more source
SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
arXiv.orgAligning 3D scene graphs is a crucial initial step for several applications in robot navigation and embodied perception. Current methods in 3D scene graph alignment often rely on single-modality point cloud data and struggle with incomplete or noisy ...
B. Singh, Sayan Deb Sarkar, Iro Armeni
semanticscholar +1 more source
EmbodiedRAG: Dynamic 3D Scene Graph Retrieval for Efficient and Scalable Robot Task Planning
arXiv.orgRecent advances in Large Language Models (LLMs) have helped facilitate exciting progress for robotic planning in real, open-world environments. 3D scene graphs (3DSGs) offer a promising environment representation for grounding such LLM-based planners as ...
Meghan Booker +4 more
semanticscholar +1 more source
A Bottom-up Framework for Construction of Structured Semantic 3D Scene Graph
IEEE/RJS International Conference on Intelligent RObots and Systems, 2020For high-level human-robot interaction tasks, 3D scene understanding is important and non-trivial for autonomous robots. However, parsing and utilizing effective environment information of the 3D scene is not trivial due to the complexity of the 3D ...
Bangguo Yu +5 more
semanticscholar +1 more source
GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering
arXiv.orgStructured scene representations are a core component of embodied agents, helping to consolidate raw sensory streams into readable, modular, and searchable formats.
Muhammad Qasim Ali +4 more
semanticscholar +1 more source
Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning
arXiv.orgRepresenting and understanding 3D environments in a structured manner is crucial for autonomous agents to navigate and reason about their surroundings. While traditional Simultaneous Localization and Mapping (SLAM) methods generate metric reconstructions
Albert Gassol Puigjaner +2 more
semanticscholar +1 more source
SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences
IEEE Transactions on Visualization and Computer GraphicsWe introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D scene that ...
Seokyoung Kim +5 more
semanticscholar +1 more source

