Graph-pruned semantic search

SemanticSplat

A hierarchical semantic graph that prunes scene search before grounding natural-language queries in captured evidence.

Research prototype. Current claims are deterministic, test-backed, and scoped below.

97team-captured RGB-D views
1,066indexed semantic items
150same-input benchmark queries
75.5%fewer views checked
68.2%fewer estimated context tokens

Research question

Can hierarchy reduce query context without changing the semantic map?

SemanticSplat organizes the same semantic records used by flat search into scene, zone, region, object, and view-evidence levels. A query retains likely branches, ranks their leaves, and expands only when confidence or geometry requires fallback.

Safest current claim

Hierarchy substantially reduces checked views and serialized query context. The present lexical ranker does not yet preserve flat-search hit@k on the internal track, so the evidence supports a measured quality-cost tradeoff, not universal superiority.

Method

Captured evidence to graph-pruned grounding

The figure is assembled from the real Conference Hall RGB-D capture, its ViewJSON record, the checked-in hierarchy, and frozen query qv2_010. No AI-generated imagery is used.

One-time construction is separated from per-query retrieval. The graph and flat lanes use identical records and scoring; only candidate selection changes.
01

Index evidence

Normalize objects, room cues, relations, optional 2D/3D boxes, and supporting views.

02

Build hierarchy

Attach deterministic entity records beneath scene, zone, and region branches.

03

Parse the query

Extract target, attributes, intent, optional anchor, and spatial relation.

04

Prune and rank

Gate branches, score retained entities, validate geometry, and expand on low confidence.

Hierarchical search path: retained target evidence and pruned siblings.
Controlled protocol: same map, query, labels, and scorer.

Evidence

Results with protocol boundaries visible

Internal graph-versus-flat numbers are directly comparable. Public-dataset and external-baseline tracks answer different questions and are not collapsed into a single ranking.

Graph4.77views/query
Flat19.4views/query
Graph1,014est. context tokens/query
Flat3,168est. context tokens/query

Five captured scenes, 150 queries

Graph reduces mean views by 75.5% and estimated context tokens by 68.2%. On 125 queries with verified view labels, hit@1 is 0.680 graph versus 0.768 flat; hit@3 is 0.808 versus 0.928.

Inspect the frozen metrics JSON
Five-scene graph versus flat summary charts.
Efficiency and quality are reported together.

Construction accounting

Measured graph-construction workload

The frozen audit separates one-time map preparation from per-query retrieval. It records the complete annotation payload, deterministic tree representation, serialized I/O, and clean-rebuild runtime for all five captured scenes.

Download construction audit
Manual records
97 ViewJSON annotations
Semantic items
1,066 manual payload entries
Annotation payload
66,937 est. tokens; 267,877 chars
Constructed tree
179,817 est. tokens
Serialized I/O
246,754 input + output est. tokens
Local rebuild
1.822 s sum of scene medians

Data

Exact evaluation sets and access boundaries

The page publishes scene IDs, manifests, derived metrics, and reproduction code. Licensed ScanNet and Replica scene files are not mirrored.

Team-ownedInternal track

Five manual captures

ConferenceHall-capture-pilot, Museume-capture, Theater-capture, outdoor-street-capture, and outdoor-drone-capture.

Views
97
Items
1,066
Queries
150
Browse scene records
Terms requiredPublic pilot

ScanNet v2

scene0011_00, scene0030_00, scene0046_00, scene0086_00, scene0222_00, scene0378_00, scene0389_00, scene0435_00.

Scenes
8
GT boxes
392
Queries
48
Request official ScanNet access
Official sourcePublic pilot

Replica

room0-room2 and office0-office4, aligned with the BBQ subset used in the paper.

Scenes
8
GT boxes
575
Queries
56
Open the official Replica repository

Machine-readable scene lists and access notes: dataset manifest JSON.

Visual audit

Captured scenes and a grounded case

Gallery of six captured digital-twin scenes.
Team-captured internal scenes. Public-dataset imagery is intentionally not redistributed here.
Qualitative projection-screen grounding case with graph and flat views.
Qualitative query evidence with a view-annotation box.

Reproduce

Code, artifacts, and validation gates

Core verification
python -B -m pytest -q
python -B scripts/check_academic_paper.py
python -B scripts/measure_graph_construction.py --repeats 7
python -B scripts/export_academic_figures.py

Limits

What this release does not claim

  • The five internal semantic indexes are manual reference annotations, not independent GT.
  • ScanNet and Replica SemanticSplat pilots use oracle semantic candidates; they do not measure semantic-perception accuracy.
  • ConceptGraphs and LangSplat use different map-construction and evaluation protocols, so their metrics are not a fair leaderboard.
  • Internal context tokens are deterministic text-size estimates. External baseline token counts are local tokenizer counts, not provider billing.
  • The current graph search saves context but trails flat lexical hit@k on the internal verified-view track.

Manuscript

SemanticSplat: Graph-Pruned Semantic Search

The public manuscript is a named-author academic preprint, independent of any workshop template. It includes the complete method, schemas, frozen metrics, uncertainty intervals, public-dataset pilots, external executions, and reproducibility contract.

Preprint citation
@article{mousatat2026semanticsplat,
  title = {SemanticSplat: Graph-Pruned Semantic Search},
  author = {Mousatat, Mahmoud and Vizan, Leo and
            Shankin, Nikita and Nuruzov, Telman and
            Medvedev, Alexandr},
  year = {2026},
  note = {Preprint}
}