R4DSG Retrieval+
+6.7 pts vs. EgoRAG-TextRelative 4D Scene Graph Memory
Object-centric question answering in long egocentric video
Remember the object, not just the sentence
Long video remembers events. Assistants need to remember things.
Where was an item moved? When did it last change state? Why was it relocated? Flat captions and clip summaries rarely preserve persistent identity or structured spatial change across hours of first-person video.
R4DSG stores a compact, queryable memory indexed by time, place, persistent objects, anchor-relative change, and local interaction context.
* Reported under the submitted protocol on the public EgoLifeQA A1_JAKE single-subject split.
From pixels to evidence
Five stages. One inspectable memory.
Semantic episodes constrain perception; persistent identity and anchor changes turn local observations into retrieval-ready evidence.
Bound the open-vocabulary search
Semantic episodes
Activity and location windows create a local time span, place, activity, and controlled object set.
Interactive reconstruction of the pipeline in Figure 2. Scientific labels, roles, distances, and states follow the camera-ready figure.
Relative, not global. “4D” denotes a time-indexed relative 3D scene-graph memory - not a globally consistent 4D reconstruction.
One bag · three moments
Scrub time. Rotate the evidence.
The same persistent bag is interpreted relative to the stable objects around it. Drag the 3D view and move across t1, t2, and t3.
Supermarket bagging area
Bagging purchased items near a shopping cart
- Target
- Bag · persistent ID
- Anchor
- Cart
- Distance
- 0.46 m
- State
- bagged groceries
supermarket bagging area
bagging purchased items near shopping cart
bag near cart; groceries packed
bag adjacent to cart; eggs and drinks co-occur
Retrieve first, then expose choices
Ask the memory, not the raw timeline.
R4DSG maps each question to an option-blind retrieval query, selects the top eight memory documents, then introduces answer choices for final synthesis.
Where did I leave the bag?
Reported results
The clearest gain is temporal.
On the public EgoLifeQA A1_JAKE single-subject split, relative object memory is most useful when the answer depends on when an object state changed.
72 when questions
+12.5 pts vs. EgoRAG-TextRetrieval+ documents
0.58 MB JSONAccuracy on 255 object-centric questions and the 72-question when subset. Gains are reported relative to EgoRAG-Text under the accepted-submission configuration.
63.6 seconds · ten curated cases
Inspect the evidence each method relies on.
The demo aligns the question, target-centered graph, retrieved memory, temporal trace, and comparative predictions in one interface. Select a case to jump directly to its chapter.
These ten cases are a curated qualitative set for evidence inspection; they are not a standalone accuracy benchmark.
R4DSG · ACM Multimedia 2026
Relative scene graphs as a memory substrate.
R4DSG reframes long egocentric QA as memory construction over persistent objects, static anchors, and retrieval-ready documents.
What this evidence supports
- Evaluation is limited to the public A1_JAKE single-subject split.
- Memory is retrospective and offline, with limited cross-day persistence.
- Adapted EMQA-style and AMEGO-inspired baselines are not official reproductions.
- WhyMemory is exploratory: the why subset contains 15 questions.
- Online updates, stronger identity maintenance, and multi-subject validation remain future work.
@inproceedings{ma2026r4dsg,
title={R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video},
author={Ma, Ke and Mao, Yamin and Li, Weiming and Tan, Shuai and Zhong, Yijie and Chen, Hao and Wang, Haofen and Wang, Meng},
booktitle={Proceedings of the 34th ACM International Conference on Multimedia},
year={2026},
doi={10.1145/3767308.3835995}
}