R4DSG Paper

ACM Multimedia 2026 · Long-horizon egocentric memory

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

Ke Ma1, Yamin Mao2, Weiming Li2, Shuai Tan3, Yijie Zhong1, Hao Chen4, Haofen Wang1, and Meng Wang1,5

1 College of Design and Innovation, Tongji University, Shanghai, China 2 Samsung R&D Institute China - Beijing, Beijing, China 3 Shanghai Jiao Tong University, Shanghai, China 4 Samsung Networks, Samsung Research America, Plano, Texas, United States 5 Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University, Shanghai, China
18:31:00 egocentric stream
bag
table
Object-centric question Where did I leave the bag?
19:12:30 relative scene memory
Persistent track cart → door → table

Relative 4D Scene Graph Memory

Object-centric question answering in long egocentric video

Remember the object, not just the sentence

Long video remembers events. Assistants need to remember things.

Where was an item moved? When did it last change state? Why was it relocated? Flat captions and clip summaries rarely preserve persistent identity or structured spatial change across hours of first-person video.

R4DSG stores a compact, queryable memory indexed by time, place, persistent objects, anchor-relative change, and local interaction context.

t118:31
bagnear cart
t219:10
bagnear door
t319:12
bagnear table
identityanchorchangetime
255object-centric questions
+6.7overall points vs. EgoRAG-Text*
+12.5when-question points*
0.58 MB134 Retrieval+ documents

* Reported under the submitted protocol on the public EgoLifeQA A1_JAKE single-subject split.

From pixels to evidence

Five stages. One inspectable memory.

Semantic episodes constrain perception; persistent identity and anchor changes turn local observations into retrieval-ready evidence.

01

Bound the open-vocabulary search

Semantic episodes

Activity and location windows create a local time span, place, activity, and controlled object set.

Interactive reconstruction of the pipeline in Figure 2. Scientific labels, roles, distances, and states follow the camera-ready figure.

Relative, not global. “4D” denotes a time-indexed relative 3D scene-graph memory - not a globally consistent 4D reconstruction.

One bag · three moments

Scrub time. Rotate the evidence.

The same persistent bag is interpreted relative to the stable objects around it. Drag the 3D view and move across t1, t2, and t3.

Unified spatial-temporal scene graph
dynamicanchorevidence
t118:31:00

Supermarket bagging area

Bagging purchased items near a shopping cart

Target
Bag · persistent ID
Anchor
Cart
Distance
0.46 m
State
bagged groceries
drag to orbit · wheel to zoom · arrow keys supported
m1

supermarket bagging area

Activity

bagging purchased items near shopping cart

Object states

bag near cart; groceries packed

Edge summary

bag adjacent to cart; eggs and drinks co-occur

Retrieve first, then expose choices

Ask the memory, not the raw timeline.

R4DSG maps each question to an option-blind retrieval query, selects the top eight memory documents, then introduces answer choices for final synthesis.

R4DSG / retrieval+
query only

Where did I leave the bag?

01Retrieve top-8mk + ei + local evidence
02Expose choicesA-D enter after retrieval
03Near the dining tableanswer + cited evidence

Reported results

The clearest gain is temporal.

On the public EgoLifeQA A1_JAKE single-subject split, relative object memory is most useful when the answer depends on when an object state changed.

Submitted protocol · overall39.6%

R4DSG Retrieval+

+6.7 pts vs. EgoRAG-Text
Submitted protocol · when43.1%

72 when questions

+12.5 pts vs. EgoRAG-Text
Offline memory scale134

Retrieval+ documents

0.58 MB JSON

Accuracy on 255 object-centric questions and the 72-question when subset. Gains are reported relative to EgoRAG-Text under the accepted-submission configuration.

500four-choice questions before object-centric filtering
7 daysin the public A1_JAKE file
828released Day1 clips across five sessions
top-8documents retrieved before answer selection

63.6 seconds · ten curated cases

Inspect the evidence each method relies on.

The demo aligns the question, target-centered graph, retrieved memory, temporal trace, and comparative predictions in one interface. Select a case to jump directly to its chapter.

Q3 · Clapperboard retrieval history00:00 / 01:04

These ten cases are a curated qualitative set for evidence inspection; they are not a standalone accuracy benchmark.

R4DSG · ACM Multimedia 2026

Relative scene graphs as a memory substrate.

R4DSG reframes long egocentric QA as memory construction over persistent objects, static anchors, and retrieval-ready documents.

Download paper DOI · forthcoming OpenReview

Ke Ma1 · Yamin Mao2 · Weiming Li2 · Shuai Tan3 · Yijie Zhong1 · Hao Chen4 · Haofen Wang1 · Meng Wang1,5

1 College of Design and Innovation, Tongji University, Shanghai, China 2 Samsung R&D Institute China - Beijing, Beijing, China 3 Shanghai Jiao Tong University, Shanghai, China 4 Samsung Networks, Samsung Research America, Plano, Texas, United States 5 Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University, Shanghai, China
Scope, not hype

What this evidence supports

  • Evaluation is limited to the public A1_JAKE single-subject split.
  • Memory is retrospective and offline, with limited cross-day persistence.
  • Adapted EMQA-style and AMEGO-inspired baselines are not official reproductions.
  • WhyMemory is exploratory: the why subset contains 15 questions.
  • Online updates, stronger identity maintenance, and multi-subject validation remain future work.
Cite R4DSG
@inproceedings{ma2026r4dsg,
  title={R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video},
  author={Ma, Ke and Mao, Yamin and Li, Weiming and Tan, Shuai and Zhong, Yijie and Chen, Hao and Wang, Haofen and Wang, Meng},
  booktitle={Proceedings of the 34th ACM International Conference on Multimedia},
  year={2026},
  doi={10.1145/3767308.3835995}
}