Dima Damen
@dimadamen
Professor of Computer Vision, @BristolUni. Senior Research Scientist @GoogleDeepMind - passionate about the temporal stream in our lives.
Whareformer @eccv.bsky.social #ECCV2026 fantastic collab bw @compscibristol.bsky.social @naverlabseurope.bsky.social by Jacob Chalk w/ @sinhasaptarshi.bsky.social @skamalas.bsky.social & @dlarlus.bsky.social Paper, Code &models public: jacobchalk.github.io/Whareformer/ arxiv.org/abs/2607.08537 N/N
We utilise the online DenStream algorithm to maintain persistent (recurrent viewpoints) and transient (new viewpoints) appearance micro-clusters within each track efficiently. Unlike applications in autonomous or multi-person tracking, tracks can be revisited after 30+mins 4/N
We train: * embedding network g(.) to combine appearance &location diff * NT token to learn to initialise new tracks * transformer encoder to contrast assignments Memory is updated online based on assignments. No tracks are removed - in ego you reuse objects after LONG gaps 3/N
Whareformer (What&Where former) tackles the OSNOM task - Out of Sight objects remain explicitly in memory so you know where objects are at all times. In natural ego videos, tracking moving objects offers a significant challenge, compared to tracking static parts of the scene. 2/N
Explore the dataset now at: sid2697.github.io/epic-contact... As we always do: Dataset, Code & models are public already along w ArXiv and full details to reproduce in code and supplementary EPIC-Cotact Dataset: huggingface.co/datasets/Sid... HOPformer Code: github.com/Sid2697/HOPf... 6/6
Work led by Siddhant Bansal w @zhifan-zhu.bsky.social Shashank Tripathi, Jiahe Zhao @michael-j-black.bsky.social and myself - a 16-months intense collaboration @compscibristol.bsky.social and MPI IS. [don't let anyone convince you good research can be done in a few months!] 5/6
+ We propose HOPformer (achieves new SOTA on Arctic) & eval on EPIC-Contact, feed-forward model that utilises hand prior to cross-attend to object pose backbone & decoder that estimates both poses of hands & object (relative to canonical). + generalisable to new instances of known classes 4/6
We release EPIC-Contact dataset - manually annotated in-contact hand vertices AND bijective annotations of corresponding object vertices. We contribute a full pipeline from hand painting to efficiently annotate objects, then fit posed hand&object meshes. Manually checked 3/6
TL;DR we predict bi-manual hand & object pose in a single forward pass from in-the-wild ego frames. Achieved through manually-collected high-quality annotations of 2.3K video clips (62.3K frames) of functional interactions during stable grasps with objects from 9 classes 2/6
*NEW* Our #ECCV2026 @eccv.bsky.social paper Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation Now on ArXiv w Dataset, Code&model sid2697.github.io/epic-contact/ arxiv.org/abs/2606.30598 Two contributions: 1. EPIC-Contact Dataset 2. HOPformer Method &Checkpoint 🧵 1/6
Congratulations @zhifan-zhu.bsky.social who passed his viva today @compscibristol.bsky.social, examined by Siyu Tang (ETH Zurich) &Andrew Calway on: Understanding 3D Dynamics of Hands, Objects and Bodies in Long Egocentric Videos zhifanzhu.github.io He is on the job market for positions in industry
What can you do with a 4-hour delay at Denver after @cvprconference.bsky.social? Create a collage of this week... 🧵 to tag all those who helped shape these captured moments... it was lovely to see you all and hope to catch up again soon...
Many congratulations Prof Jitendra Malik elected as a Fellow of the Royal Society @royalsociety.org in 2026 - well and truly deserved! Jitendra's contributions extend to inspiring, advising and mentoring many in our community, and just being a very nice person! royalsociety.org/news/2026/05...
From the V-Jepa2 presentation: Action anticipation in EPIC-KITCHENS is "a very challenging task and still there’s head room for improvement on this task". V-Jepa2 establishes a new SOTA but still <65% on both verb and noun acc. Mido Assran at the @iclr-conf.bsky.social workshop on World Models
Attending @iclr-conf.bsky.social #ICLR2026? Join us for the 2nd workshop on world models - full day 202 A/B @meng-yue-yang.bsky.social opening the workshop right now
Workshops are always more interesting than the main conference! @iclr-conf.bsky.social What's next in Next Token Prediction workshop @jcniebles.bsky.social starting the action,
Exciting 2 weeks ahead @csail.mit.edu thanks to B Freeman for hosting.. Sad I'll be missing Antonio - but what would be a better inspirational place to stay than his office!! Already caught up with @vincentsitzmann.bsky.social @sarameghanbeery.bsky.social ... will be a fun time!
At @brics-uob.bsky.social summit today @bristoluni.bsky.social. Happy faces of users all around. I &my team are proud users thanks to national DSIT grants. I took virtual pictures with the original Isambard and the new Isambard! Thanks @simonmcs.bsky.social, Sadaf and the team for the invitation.
Congrats Ahmad Darkhalil who passed his PhD today @compscibristol.bsky.social, examined by @csprofkgd.bsky.social & @chriswolfvision.bsky.social. Thesis includes his excellent work on: VISOR, EPIC Fields, HD-EPIC, EgoPoints, &Hand-Object Detection. Check works at: ahmaddarkhalil.github.io
anyone else keeps confusing overleaf with openreview when they type? I keep typing openreview when I mean overleaf, and overleaf when I mean openreview! #overleaf #openreview Ideas to solve my brain wiring are welcome!
Special thanks Ayush @ayusht.bsky.social @universitypress.cambridge.org for visiting us @bristoluni.bsky.social @compscibristol.bsky.social for a great #MaVi Seminar on "Physical Inductive Biases for World Models" and thoughtful 1-1s with the researchers. Have a good trip back &pls visit again soon,
3rd Egocentric Vision (EgoVis) workshop will be held as a full day workshop @cvprconference.bsky.social #CVPR2026 egovis.github.io/cvpr26/ CFP and challenge deadlines after the NY Great lineup of 7 keynote speakers... See you in Denver!
Call for Nominations EgoVis 2024/2025 Distinguished Paper Awards. Published a paper contributing to Ego Vision in 2024/25? Innovative &advancing Ego Vision? Worthy of a prize? DL for nominations 20 Feb 2026 Awards announced @cvprconference.bsky.social #CVPR2026 egovis.github.io/awards/2024_...
Work by @zhifan-zhu.bsky.social @bristoluni.bsky.social @compscibristol.bsky.social in close collaboration w Yifei Huang &Yoichi Sato @utokyoofficial.bsky.social as part of the ASPIRE collaboration b/w Japan-UK. Zhifan Zhu will be on industrial R&D job market in 2026 - reach out! 4/4
We introduce a structured prompting strategy to guide a VLM reasoning about the 3D environment, object usage, and temporal dependencies. Using public annotations, we measure performance on new eval metrics for parallel performance &analyse results thoroughly. 3/4
How? formulated as achieving speed⌛️&coverage⚒️ while respecting constraints for a feasible parallel execution around 3D spatial coverage of space🏘️, object constraints🥎(two people cannot use 1 object @ same time) &causal constraints➡️. Tested on 100 LONG egocentric videos 2/4
Preprint now on ArXiv 📢 The N-Body Problem: Parallel Execution from Single-Person Egocentric Video Input: Single-person egocentric video 👤 Out: imagine how these tasks can be performed faster by N > 1 people, correctly e.g. N=2 👥 📎 arxiv.org/abs/2512.11393 👀 zhifanzhu.github.io/ego-nbody/ 1/4
MaVi Faculty/Hosts: Peter Flach, Majid Mirmehdi, Tilo Burghardt, Raul Santos Rodriguez, Michael Wray, Zahraa S. Abdallah, Guosheng Hu, Wei-Hong Li, Xiang Lu, Nan Lu, Mengyue Yang, Telmo Silva Filho &myself