MULTI-VIEW SPATIAL INTELLIGENCE

SpatialSpeak.

QA-Native Reconstruction with Local and Global Context
for Spatial Chain-of-Thought Reasoning

Yang Cao1Jiaxin Zhang3Dave Zhenyu Chen2Yingji Zhong1Ruiyuan Gao2Lanqing Hong2Dan Xu1,*

1 Hong Kong University of Science and Technology2 Huawei Noah’s Ark Lab3 Harbin Institute of Technology

Download demo video

Learning local geometry and global context makes spatial CoT more effective.

62.8 on ReVSI · +8.7 points. Reconstruction pretraining connects geometric understanding with spatial reasoning. View PDF ↗

THE SCENES BEHIND THE PAPER

Explore the 3D Scenes

Drag to rotate. Scroll or pinch to zoom.
Choose a scene to inspect its reconstruction and paper example.

Colored reconstructionLoading scene…
A colored point cloud of a classroom with two blackboards
Loading 3D scene…
↔ Drag to rotate · Scroll to zoom
INPUT VIEWS

OBJECT COUNTING

How many blackboards are in the scene?

A recorded example from the paper, paired with its reconstruction.

SpatialSpeak response
    SpatialSpeak2 blackboards
    Ground truth2

    Without QA-RP and CoT-VC: 3

    Metric Reconstruction

    The red arrows show exactly which span is measured.

    Original measurement annotations from the paper. Switch to the 3D view to inspect the corresponding reconstruction.

    TWO STAGES, ONE LANGUAGE INTERFACE

    Method

    01

    QA-RP · RECONSTRUCTION PRETRAINING

    Marked-point 3D queries teach fine-grained local geometry. Object-center queries teach global scene context across views. Both use text-based question answering.

    02

    CoT-VC · SPATIAL REASONING

    Spatial chain-of-thought connects estimated geometry to the requested answer. Reliability assessment and visual compensation support refinement when needed.

    QA-native reconstruction pretraining connects local and global scene understanding with spatial chain-of-thought learning. View PDF ↗

    EXPERIMENTS

    Experiments

    Results reported in the paper.
    Higher is better.

    62.8%

    ReVSI average

    +8.7pts

    Over the strongest compared baseline

    4B

    VLM backbone parameters

    ReVSI performance

    Average score (%)
    SpatialSpeak-4B
    62.8
    SpatialStack-4B
    54.1
    GeoThinker-8B
    54.0
    VLM-3R-7B
    50.1
    Cambrian-S-7B
    49.1

    Displayed score range: 20–65%. Selected comparisons from Table 6. SpatialSpeak, SpatialStack, and GeoThinker are evaluated under the same 32-frame setting. VLM-3R and Cambrian-S results are sourced from the ReVSI paper.

    WHY RECONSTRUCTION MATTERS

    Spatial CoT gains more
    after QA-RP.

    The improvement from CoT-VC grows from 2.6 to 6.9 points when the model first learns reconstruction.

    Without QA-RP
    +2.6
    With QA-RP
    +6.9

    ReVSI ablation, Table 1. Baseline: 52.4. CoT-VC only: 55.0. QA-RP only: 55.9. Full method: 62.8.

    Spatial reasoning across benchmarks
    BenchmarkSettingSpatialSpeak-4BBest compared baseline
    ReVSI32-frame evaluation62.854.1 SpatialStack-4B
    VSI-BenchNormal training63.355.4 Omni-View-7B
    VSI-BenchScaled training73.072.6 GeoThinker-8B
    SPAR-BenchOverall average76.072.0 SpatialStack-4B

    Scores and comparison settings follow Tables 6–9 of the paper. Normal and scaled training settings use different training data and should be read separately.

    Abstract

    Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction.

    We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed.

    On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.

    REFERENCE

    Citation

    @article{cao2026spatialspeak,
      title = {SpatialSpeak: QA-Native Reconstruction with Local and Global
               Context for Spatial Chain-of-Thought Reasoning},
      author = {Cao, Yang and Zhang, Jiaxin and Chen, Dave Zhenyu and
                Zhong, Yingji and Gao, Ruiyuan and Hong, Lanqing and Xu, Dan},
      journal = {arXiv preprint arXiv:2609.33616},
      year = {2026}
    }

    arXiv:2609.33616

    Paper figure

    View original PDF ↗