Action Scene Graphs for Long-Form Understanding of Egocentric Videos (CVPR 2024)
-
Updated
Apr 9, 2025 - Jupyter Notebook
Action Scene Graphs for Long-Form Understanding of Egocentric Videos (CVPR 2024)
NaQ: Leveraging Narrations as Queries to Supervise Episodic Memory. CVPR 2023.
Models of Mental Simulation
Winner of CVPR23 EGO4D STA challenge
Curated datasets, benchmarks, models, and tools for egocentric AI, embodied intelligence, VLA, world models, robotics, and wearable vision.
[NeurIPS 2026] Linguistic Trajectory Encoding (LTE): object-centric spatiotemporal memory for embodied agents. CPU and real-video demos with SAM3, ViPE, DINOv2, and Qwen3-VL. Spatial Memory Benchmark (SMB). arXiv:2609.04802.
A two-step framework for extracting textual answers from egocentric videos via NLQ, combining VSLNet for segment localization and Video-LLaVA for efficient answer generation.
Official implementation of FlowNar: Scalable Streaming Narration for Long-Form Videos
Efficient NLVL on Ego4D. Benchmark VSLBase/VSLNet with BERT/GloVE, EgoVLP/Omnivore, configurable FiLM conditioning layer (Perez et al., FiLM: Visual Reasoning with a General Conditioning Layer), explores compression via KD, including CBKD (Lan et al., Counterclockwise Block‑by‑Block Knowledge Distillation), prototype post‑training quantization.
Two-stage pipeline on the Ego4D NLQ benchmark: finding the moment in a first-person video that answers a question, then answering it — VSLNet/VSLBase ablation for temporal localisation, Video-LLaVA for answer generation, scored against hand-written ground truth.
To associate your repository with the ego4d topic, visit your repo's landing page and select "manage topics."