This repository is originating from our survey paper "Unifying Video Self-Supervised Learning across Families of Tasks: A Survey" and authors (Ishan Dave*, Malitha Gunawardhana*, Limalka Sadith, Honglu Zhou, Liel David, Daniel Harari, Mubarak Shah, Muhammad Haris Khan) will continue to update this over time.
Browse the searchable website for the same collection with paper search, repository statistics, and direct links to the research.
Abstract: Video self-supervised learning (VideoSSL) offers significant potential for reducing annotation costs and enhancing a wide range of downstream tasks in video understanding. The ultimate goal of VideoSSL is to achieve human-level video intelligence across a spectrum of tasks, from low-level tasks such as pixel temporal correspondence to high-level complex spatio-temporal tasks like action recognition. However, most existing VideoSSL methods focus on isolated aspects of this spectrum and fail to integrate different levels of task complexity. Our study presents the first comprehensive survey that connects all families of VideoSSL methods. We provide a detailed review of the full spectrum of VideoSSL, from low to high levels, by conceptually linking their self-supervised learning objectives and including a comprehensive categorization. Our extensive evaluation highlights the strengths and limitations of each SSL objective across various downstream task families. We also detail the challenges in current VideoSSL research such as data curation, interpretability, deployment, and privacy concerns, an area that previous surveys have not thoroughly explored. In addressing these challenges, we recognize the strengths of existing methods in addressing these challenges and outline future directions for research.
Overview of the three major families of video self-supervised learning methods. Dave and Gunawardhana et al. (2024)
This repository contains a collection of state-of-the-art self-supervised learning in video approaches for various downstream tasks, such as action recognition, video retrieval, etc. With the exponential growth of video data, there is an increasing need for automatic video analysis methods that can learn from large amounts of unlabeled data. Self-supervised learning provides an effective solution to this problem by allowing models to learn from the data itself without explicit supervision.
This research was supported by the joint grant P007 from Mohamed Bin Zayed University of Artificial Intelligence and the Weizmann Institute of Science. The authors would like to express their sincere gratitude for this generous support, which made the study possible.
If you find our work useful. Please consider giving a star ⭐ and a citation.
@article{dave2024unifying,
title={Unifying Video Self-Supervised Learning across Families of Tasks: A Survey},
author={Dave, Ishan and Gunawardhana, Malitha and Sadith, Limalka and Zhou, Honglu and David, Liel and Harari, Daniel and Shah, Mubarak and Khan, Muhammad Haris},
year={2024},
publisher={Preprints}
}
In this repository, we have gathered some of the most promising self-supervised learning approaches for video analysis and organized them based on their publication year. Whether you are new to self-supervised learning in videos or an experienced researcher in the field, we hope that this repository will serve as a valuable resource for exploring the latest advances in this exciting area of research.
Let's collaborate and enrich this list together! Reach out to me or submit a pull request. Your contributions are highly appreciated.
-
Unifying Video Self-Supervised Learning across Families of Tasks: A Survey (2024)
Preprint
Ishan Dave*, Malitha Gunawardhana*, Limalka Sadith, Honglu Zhou, Liel David, Daniel Harari, Mubarak Shah, Muhammad Hairs Khan
[Paper] -
Self-Supervised Learning for Videos: A Survey (2022)
ACM Computing Surveys
Madeline C. Schiappa, Yogesh S. Rawat, And Mubarak Shah
[Paper]
-
SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning (2024)
arXiv preprint
Fida Mohammad Thoker, Letian Jiang, Chen Zhao, Piyush Bagad, Hazel Doughty, Bernard Ghanem, Cees G. M. Snoek
[Paper] [Code] -
How Effective are Self-Supervised Models for Contact Identification in Videos (2024)
arXiv preprint
Malitha Gunawardhana, Limalka Sadith, Liel David, Daniel Harari, Muhammad Haris Khan
[Paper] [Code] -
Benchmarking self-supervised video representation learning (2023)
arXiv preprint arXiv:2306.06010
Akash Kumar, Ashlesha Kumar, Vibhav Vineet, Yogesh Singh Rawat
[Paper] [Page] -
A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action Recognition (2023)
arXiv preprint arXiv:2303.13505
Deng, A., Yang, T., & Chen, C.
[Paper] -
How Severe Is Benchmark-Sensitivity in Video Self-supervised Learning? (2022, October)
In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022
Fida Mohammad Thoker, Hazel Doughty, Piyush Bagad, Cees Snoek
[Paper] [Github] [Page]
-
Progressive Mask Distillation for Self-supervised Video Representation (2026)
CVPR 2026
Kewei Wu, Chong Liang, Zhao Xie, Dan Guo
[Paper] -
TrackMAE: Video Representation Learning via Track Mask and Predict (2026)
CVPR 2026
Renaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard Ghanem
[Paper] [Code] -
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning (2026)
arXiv preprint
Lorenzo Mur-Labadia, Matthew Muckley, Amir Bar, Mido Assran, Koustuv Sinha, Mike Rabbat, Yann LeCun, Nicolas Ballas, Adrien Bardes
[Paper] [Code] -
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning (2026)
CVPR 2026
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Xilin Zhao, Qingming Huang
[Paper] [Code] -
The TIME Machine: On The Power of Motion for Efficient Perception (2026)
arXiv preprint
Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
[Paper] [Project Page] -
TrAction: Action Recognition with Sparse Trajectories (2026)
arXiv preprint
Jan F. Meier, Felix B. Mueller, Alexander Ecker, Timo Lüddecke
[Paper] [Code] -
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence (2026)
arXiv preprint
Feilong Tang, Xiang An, Yunyao Yan, Yin Xie, Bin Qin, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Chunyuan Li, Shikun Feng, Changrui Chen, Huajie Tan, Ming Hu, Manyuan Zhang, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng
[Paper] [Code] -
Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives (2026)
arXiv preprint
Santosh Premi
[Paper] [Code] -
Self-Supervised Learning of Structured Dynamics from Videos (2026)
arXiv preprint
Lukas Knobel, Andrew Zisserman, Yuki M. Asano
[Paper] [Code] [Project Page] -
Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers (2026)
arXiv preprint
Jonas Römer, Timo Dickscheid
[Paper] [Code] -
Beyond reconstruction: Enhancing masked autoencoders with contrastive learning for video representation learning (2026)
Engineering Applications of Artificial Intelligence, volume 171, article 114283 (2026)
Yawei Feng, Lijun Guo, Guitao Yu, Rong Zhang, Jiangbo Qian, Chong Wang, Shangce Gao
[Paper] -
Structured-Noise Masked Modeling for Video, Audio and Beyond (2026)
ECCV 2026
Aritra Bhowmik, Fida Mohammad Thoker, Carlos Hinojosa, Bernard Ghanem, Cees G. M. Snoek
[Paper] -
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers (2026)
ICLR 2026
Xianhang Li, Chen Huang, Chun-Liang Li, Eran Malach, Josh Susskind, Vimal Thilak, Etai Littwin
[Paper] -
Recurrent Video Masked Autoencoders (2026)
CVPR 2026
Daniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson, Joao Carreira, Andrew Zisserman
[Paper] -
Dual Perspectives on Non-Contrastive Self-Supervised Learning (2026)
ICLR 2026
Jean Ponce, Martial Hebert, Basile Terver ;
[Paper] -
Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective (2026)
International Journal of Computer Vision 2026
Zeen Song, Jingyao Wang, Jianqi Zhang, Changwen Zheng, Wenwen Qiang
[Paper] -
BIMM: Brain Inspired Masked Modeling for Video Representation Learning (2026)
IEEE Transactions on Circuits and Systems for Video Technology 2026
Zhifan Wan, Jie Zhang, Changzhen Li, Shiguang Shan
[Paper]
-
Efficient VideoMAE via Temporal Progressive Training (2025)
CVPR Workshops 2025
Xianhang Li, Peng Wang, Xinyu Li, Heng Wang, Hongru Zhu, Cihang Xie
[Paper] -
An Empirical Study of Autoregressive Pre-training from Videos (2025)
ICCV 2025
Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravishankar, Yossi Gandelsman, Christoph Feichtenhofer, Jitendra Malik
[Paper] -
Reinforcement Learning Meets Masked Video Modeling: Trajectory-Guided Adaptive Token Selection (2025)
ICCV Workshops 2025
Ayush K. Rai, Kyle Min, Tarun Krishna, Feiyan Hu, Alan F. Smeaton, Noel E. O'Connor
[Paper] -
Entropy-Guided Masked Autoencoding for Self-Supervised Human Action Recognition Using Video Swin Transformer (2025)
ACROSET 2025
Kollu Praveen Kumar; Guduri Baby Harshitha; Koti Vijay; Angothu Sravika
[Paper] -
Privacy Preservation Using Superimposed 3D-Models for Self-Supervised Training in Action Recognition (2025)
ICCV Workshops 2025
Asfandyar Azhar, Nidhish Shah, Shaurjya Mandal, Yongjie Jessica Zhang;
[Paper] -
Learning Complementary Knowledge via Trusted Multi-view Space Decomposition for Self-Supervised Contrastive Learning (2025)
Machine Learning 2025
Jiangmeng Li, Yunze Zhao, Yifan Jin, Changwen Zheng & Wenwen Qiang;
[Paper] -
OSKAR: Omnimodal Self-supervised Knowledge Abstraction and Representation (2025)
NeurIPS 2025
Mohamed O Abdelfattah, Kaouther Messaoud, Alexandre Alahi;
[Paper] -
MME: Video Representation Learning as World Model for Understanding and Planning (2025)
TechRxiv preprint
Xinyu Sun, Changhao Li, Chen Jian, Chuang Gan, Peihao Chen, and Mingkui Tan;
[Paper] -
Hashtag2Action: Data Engineering and Self-Supervised Pre-Training for Action Recognition in Short-Form Videos (2025)
ICCV Workshops 2025
Yang Qian, Ali Kargarandehkordi, Yinan Sun, Parnian Azizian, Onur Cezmi Mutlu, Saimourya Surabhi, Zain Jabbar, Dennis Wall, Peter Washington, Huaijin Chen;
[Paper] -
Kdhiera: boosting self-supervised masked video modeling via hierarchical knowledge distillation (2025)
Cluster Computing 2025
Yunlong Wang, Hong Liang, Mingwen Shao & Qian Zhang;
[Paper] -
Self-supervised video representation learning based on foreground and temporal information. (2025)
Proceedings of SPIE, ETAI 2025
Zhongliang Zhou, Jiayong Fang ;
[Paper] -
Feature Hallucination for Self-supervised Action Recognition (2025)
International Journal of Computer Vision 2025
Lei Wang, Piotr Koniusz; ;
[Paper] -
ViDROP: Video Dense Representation through Spatio-Temporal Sparsity (2025)
CVPR Workshops 2025
Sepehr Sameni, Simon Jenni, Paolo Favaro; ;
[Paper] -
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding (2025)
CVPR 2025
Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang; ;
[Paper] -
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (2025)
arXiv preprint
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, Nicolas Ballas ;
[Paper] [Code] -
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning (2025)
CVPR 2025
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Qingming Huang;
[Paper] -
Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals (2025)
NeurIPS 2025
Stefan Stojanov, David Wendt, Seungwoo Kim, Rahul Venkatesh, Kevin Feigelis, Jiajun Wu, Daniel LK Yamins
[Paper] -
Label Ranker: Self-aware Preference for Classification Label Position in Visual Masked Self-supervised Pre-trained Model (2025)
ICMR 2025
Peihao Xiang, Ou Bai
[Paper] -
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing (2025)
CVPR 2025
Niu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo, Yaowei Wang, Shu-Tao Xia, Bin Chen
[Paper] [Code] -
Efficient Self-Supervised Video Hashing with Selective State Spaces (2025)
AAAI 2025
Jinpeng Wang, Niu Lian, Jun Li, Yuting Wang, Yan Feng, Bin Chen, Yongbing Zhang, Shu-Tao Xia1
[Paper] -
Exemplar-free class incremental action recognition based on self-supervised learning (2025)
Image and Vision Computing 2025
Chunyu Hou, Yonghong Hou, Jinyin Jiang, Gunel Abdullayeva
[Paper] -
Learning from Streaming Video with Orthogonal Gradients (2025)
CVPR 2025
Tengda Han⋄, Dilara Gokay, Joseph Heyward, Chuhan Zhang, Daniel Zoran, Viorica Patraucean, Joao Carreira, Dima Damen, Andrew Zisserman
[Paper] -
Intuitive physics understanding emerges from self-supervised pretraining on natural videos (2025)
arXiv preprint
Quentin Garrido, Nicolas Ballas, Mahmoud Assran, Adrien Bardes, Laurent Najman, Michael Rabbat, Emmanuel Dupoux, Yann LeCun
[Paper] -
ST-HViT: spatial-temporal hierarchical vision transformer for action recognition (2025)
Pattern Analysis and Applications 2025
Limin Xia, Weiye Fu
[Paper] -
Advancing video self-supervised learning via image foundation models (2025)
Pattern Recognition Letters 2025
Jingwei Wu, Zhewei Huang, Chang Liu
[Paper] [Code] -
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning (2025)
CVPR 2025
Fida Mohammad Thoker, Letian Jiang, Chen Zhao†, Bernard Ghanem
[Paper] [Code] -
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning (2025)
CVPR Workshops 2025
Akash Kumar, Ashlesha Kumar, Vibhav Vineet, Yogesh S Rawat
[Paper] -
Progressive self-supervised spatio-temporal feature learning based on video sequence saliency (2025)
Proceedings of SPIE, ICVIP 2024
Jinlong Kang, Tao Xu, Boting Qu, Xiang Wang, Xiaoli Lian, Jing Guo, Yuan Gao
[Paper] -
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders (2025)
arXiv preprint
Shihab Aaqil Ahamed∗, Malitha Gunawardhana∗, Liel David, Michael Sidorov, Daniel Harari, Muhammad Haris Khan
[Paper] -
Motion-driven Adaptive Frame Selection Strategy for Video Action Recognition (2025)
EURASIP Journal on Image and Video Processing 2025
Hao Ding, Chen Guo, Jing Sun, Xiaoping Jiang, Hongling Shi, Jianjin Li
[Paper] -
Mitigating background bias in self-supervised video representation learning (2025)
Signal, Image and Video Processing 2025
Arif Akar, Ufuk Umut Senturk & Nazli Ikizler-Cinbis
[Paper] -
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning (2025)
Transactions on Machine Learning Research 2025
Sucheng Ren, Hongru Zhu, Chen Wei, Yijiang Li, Alan Yuille, Cihang Xie
[Paper] -
Learning Video Representations without Natural Videos (2025)
ICCV Workshops 2025
Xueyang Yu, Xinlei Chen, Yossi Gandelsman;
[Paper] -
Collaboratively Self-supervised Video Representation Learning for Action Recognition (2025)
IEEE Transactions on Information Forensics and Security 2025
Jie Zhang, Zhifan Wan, Lanqing Hu, Stephen Lin, Shuzhe Wu, Shiguang Shan
[Paper]
-
Asymmetric Masked Distillation for Pre-Training Small Foundation Models (2024)
CVPR 2024
Zhiyu Zhao, Bingkun Huang, Sen Xing, Gangshan Wu, Yu Qiao, Limin Wang
[Paper] -
Data Collection-free Masked Video Modeling (2024)
ECCV 2024
Yuchi Ishikawa, Masayoshi Kondo, Yoshimitsu Aoki
[Paper] -
Text-Guided Video Masked Autoencoder (2024)
ECCV 2024
David Fan, Jue Wang, Shuai Liao, Zhikang Zhang, Vimal Bhat, Xinyu Li
[Paper] -
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space (2024)
BMVC 2024
Mona Ahmadian, Frank Guerin, Andrew Gilbert
[Paper] -
Extending Video Masked Autoencoders to 128 Frames (2024)
NeurIPS 2024
Nitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde, Eirikur Agustsson, Sagar M. Waghmare, Mikhail Sirotenko, Ming-Hsuan Yang, Tobias Weyand, Boqing Gong, Leonid Sigal
[Paper] -
Scaling 4D Representations (2024)
arXiv / Preprint
João Carreira et al.
[Paper] -
VideoMAC: Video Masked Autoencoders Meet ConvNets (2024)
CVPR 2024
Gensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu, Zeren Sun, Yazhou Yao
[Paper] -
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention (2024)
arXiv / Preprint
Quang-Trung Truong,Duc Thanh Nguyen, Binh-Son Hua, Sai-Kit Yeung
[Paper] -
Towards Latent Masked Image Modeling for Self-supervised Visual Representation Learning (2024)
ECCV 2024
Yibing Wei, Abhinav Gupta & Pedro Morgado
[Paper] [Code] -
SIGMA: Sinkhorn-Guided Masked Video Modeling (2024)
ECCV 2024
Mohammadreza Salehi, Michael Dorkenwald, Fida Mohammad Thoker, Efstratios Gavves, Cees G. M. Snoek & Yuki M. Asano
[Paper] [Code] -
ST2ST: Self-Supervised Test-time Adaptation for Video Action Recognition (2024)
CVPR Workshops 2024
Masud An-Nur Islam Fahim, Mohammed Innat, Jani Boutellier;
[Paper] -
Self-supervised learning of video representations from a child's perspective (2024)
CogSci 2024
A. Emin Orhan, Wentao Wang, Alex N. Wang, Mengye Ren, Brenden M. Lake;
[Paper] [Code] -
ViC-MAE: Self-supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders (2024)
ECCV 2024
Jefferson Hernandez, Ruben Villegas, Vicente Ordonez;
[Paper] [Code] -
Learning to Predict Activity Progress by Self-Supervised Video Alignment (2024)
CVPR 2024
Gerard Donahue, Ehsan Elhamifar;
[Paper] [Code] -
Repeat and learn: Self-supervised visual representations learning by Repeated Scene Localization (2024)
Pattern Recognition 2024
Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen;
[Paper] [Code] -
ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations (2024)
CVPR 2024
Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen;
[Paper] -
Self-supervised Learning of Semantic Correspondence Using Web Videos (2024)
WACV 2024
Donghyeon Kwon, Minsu Cho, Suha Kwak;
[Paper] -
Video Compression and Action Recognition in Self-supervised Learning (2024)
IPEC 2024
Zongbo Hao; Conghui Hao; Kecheng He
[Paper] -
CycleCL: Self-supervised Learning for Periodic Videos (2024)
WACV 2024
Matteo Destro, Michael Gygl
[Paper] -
Self-Supervised Learning via Multi-Transformation Classification for Action Recognition (2024)
ICME Workshops 2024
Duc-Quang Vu; Ngan Le; Jia-Ching Wang
[Paper] -
Motion-guided spatiotemporal multitask feature discrimination for self-supervised video representation learning (2024)
Pattern Recognition 2024
Shuai Bi, Zhengping Hu, Hehao Zhang, Jirui Di, Zhe Sun
[Paper] -
What When and Where? Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions (2024)
CVPR 2024
Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas, Shih-Fu Chang, Rogerio Feris, James Glass, Hilde Kuehne
[Paper] -
Clustering-based multi-featured self-supervised learning for human activities and video retrieval (2024)
Applied Intelligence 2024
Muhammad Hafeez Javed, Zeng Yu, Taha M. Rajeh, Fahad Rafique & Tianrui Li
[Paper] -
Positive and negative sampling strategies for self-supervised learning on audio-video data (2024)
ICASSP Workshops 2024
Shanshan Wang, Soumya Tripathy, Toni Heittola, Annamaria Mesaros
[Paper] -
No More Shortcuts: Realizing the Potential of Temporal Self-Supervision (2024)
AAAI 2024
Ishan Rajendrakumar Dave, Simon Jenni, Mubarak Shah.
[Paper] [Project Page] -
GLOCAL: A self-supervised learning framework for global and local motion estimation (2024)
Pattern Recognition Letters 2024
Yihao Zheng , Kunming Luo , Shuaicheng Liu , Zun Li , Ye Xiang , Lifang Wu , Bing Zeng , Chang Wen Chen
[Paper] -
Self-supervised Video Representation Learning via Capturing Semantic Changes Indicated by Saccades (2024)
IEEE Transactions on Circuits and Systems for Video Technology 2024
Qiuxia Lai, Ailing Zeng, Ye Wang, Lihong Cao, Yu Li, Qiang Xu, IEEE
[Paper] -
MAR: Masked Autoencoders for Efficient Action Recognition (2024)
IEEE Transactions on Multimedia 2024
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Xiang Wang, Yuehuan Wang, Yiliang Lv, Changxin Gao, Nong Sang
[Paper] [Code] -
VicTR: Video-conditioned Text Representations for Activity Recognition (2024)
CVPR 2024
Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo
[Paper] -
Self-Supervised Video Representation Learning by Video Incoherence Detection (2024)
IEEE Transactions on Cybernetics 2024
Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie, Jianxiong Yin, Simon See, Qianwen Xu, and Jianfei Yang
[Paper] -
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding (2024)
ICLR 2024
Xiong, Y., Zhao, L., Gong, B., Yang, M. H., Schroff, F., Liu, T., ... & Yuan, L.
[Paper] -
MotionMAE: Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders (2024)
BMVC 2024
Haosen Yang, Deng Huang, Bin Wen, Jiannan Wu, Hongxun Yao, Yi Jiang, Xiatian Zhu, Zehuan Yuan
[Paper] [Code] -
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens (2024)
ICML 2024
Sunil Hwang, Jaehong Yoon, Youngwan Lee, Sung Ju Hwan
[Paper] [Code] -
XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning (2024)
AAAI 2024
Pritam Sarkar, Ali Etemad
[Paper] [Code] -
Controllable Augmentations for Video Representation Learning (2024)
Visual Intelligence 2024
Rui Qian, Weiyao Lin, John See, Dian Li
[Paper]
-
Self-supervised object-centric learning for videos (2023)
NeurIPS 2023
Görkay Aydemir, Weidi Xie, Fatma Guney
[Paper] -
Language-based Action Concept Spaces Improve Video Self-Supervised Learning (2023)
NeurIPS 2023
Kanchana Ranasinghe, Michael S Ryoo
[Paper] -
Uncovering the Hidden Dynamics of Video Self-supervised Learning under Distribution Shifts (2023)
NeurIPS 2023
Pritam Sarkar, Ahmad Beirami, Ali Etemad
[Paper] [Project Page] -
Self-supervised video pretraining yields robust and more human-aligned visual representation (2023)
NeurIPS 2023
Nikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. Hénaff.
[Paper] -
AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders (2023)
CVPR 2023
Wele Gedara Chaminda Bandara, Naman Patel, Ali Gholami, Mehdi Nikkhah, Motilal Agrawal, Vishal M. Patel
[Paper] -
Spatio-Temporal Crop Aggregation for Video Representation Learning (2023)
ICCV 2023
Sepehr Sameni, Simon Jenni, Paolo Favaro
[Paper] -
Motion-Guided Masking for Spatiotemporal Representation Learning (2023)
ICCV 2023
David Fan, Jue Wang, Shuai Liao, Yi Zhu, Vimal Bhat, Hector Santos-Villalobos, Rohith MV, Xinyu Li
[Paper] -
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning (2023)
ACM Multimedia 2023
Minghao Zhu, Xiao Lin, Ronghao Dang, Chengju Liu, Qijun Chen
[Paper] -
Unmasked Teacher: Towards Training-Efficient Video Foundation Models (2023)
ICCV 2023
Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, Yu Qiao
[Paper] -
Concatenated Masked Autoencoders as Spatial-Temporal Learner (2023)
arXiv / Preprint
Zhouqiang Jiang, Bowen Wang, Tong Xiang, Zhaofeng Niu, Hong Tang, Guangshun Li, Liangzhi Li
[Paper] -
AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual Masked Autoencoder (2023)
ICTAI 2023
Xingjian Diao, Ming Cheng, Shitong Cheng
[Paper] -
OmniMAE: Single Model Masked Pretraining on Images and Videos (2023)
CVPR 2023
Rohit Girdhar, Alaaeldin El-Nouby, Mannat Singh,Kalyan Vasudev Alwala, Armand Joulin , Ishan Misra
[Paper] [Code] -
TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition (2023)
CVPR 2023
Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak Shah
[Paper] [Code] [Project Page] -
Attentive spatial-temporal contrastive learning for self-supervised video representation (2023)
Image and Vision Computing 2023
Xingming Yang, Sixuan Xiong, Kewei Wu, Dongfeng Shan, Zhao Xie
[Paper] -
MGMAE: Motion Guided Masking for Video Masked Autoencoding (2023)
ICCV 2023
Bingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao, Limin Wang
[Paper] [Code] -
Cross-modal Manifold Cutmix for Self-supervised Video Representation Learning (2023)
MVA 2023
Srijan Das; Michael Ryoo
[Paper] -
CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video Hashing (2023)
ACM Multimedia 2023
Rukai Wei, Yu Liu, Jingkuan Song, Heng Cui, Yanzhao Xie, Ke Zhou
[Paper] -
Data-Efficient Masked Video Modeling for Self-supervised Action Recognition (2023)
ACM Multimedia 2023
Qiankun Li, Xiaolong Huang, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang, Shiguang Shan, Zengfu Wang(
[Paper] -
Temporal Transformer Networks with Self-Supervision for Action Recognition (2023)
IEEE Internet of Things Journal 2023
Yongkang Zhang, Jun Li, Guoming Wu, Han Zhang, Zhiping Shi, Member, IEEE, Zhaoxun Liu, Zizhang Wu
[Paper] -
CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition (2023)
arXiv / Preprint
Cheng-Ze Lu, Xiaojie Jin, Zhicheng Huang, Qibin Hou, Ming-Ming Cheng, Jiashi Feng
[Paper] -
Learning Representational Invariances for Data-Efficient Action Recognition (2023)
Computer Vision and Image Understanding 2023
Yuliang Zou, Jinwoo Choi, Qitong Wang, Jia-Bin Huang
[Paper] [Code] -
SOR-TC: Self-attentive octave ResNet with temporal consistency for compressed video action recognition (2023)
Neurocomputing 2023
Junsan Zhang, Xiaomin Wang, Yao Wan, Leiquan Wang, Jian Wang, Philip S. Yu
[Paper] -
Masked Motion Encoding for Self-Supervised Video Representation Learning (2023)
CVPR 2023
Xinyu Sun, Peihao Chen, Liangwei Chen, Thomas H. Li, Mingkui Tan, Chuang Gan
[Paper] [Code] -
Spatiotemporal consistency enhancement self-supervised representation learning for action recognition (2023)
Signal, Image and Video Processing 2023
Shuai Bi, Zhengping Hu, Mengyao Zhao, Shufang Li & Zhe Sun
[Paper] -
Self-Supervised Video-Based Action Recognition With Disturbances (2023)
IEEE Transactions on Image Processing 2023
Wei Lin, Xinghao Ding, Yue Huang, Huanqiang Zeng
[Paper] -
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning (2023)
CVPR 2023
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Lu Yuan, Yu-Gang Jiang
[Paper] [Code] -
Enhancing motion visual cues for self-supervised video representation learning (2023)
Engineering Applications of Artificial Intelligence 2023
Mu Nie, Zhibin Quan, Weiping Ding, and Wankou Yang
[Paper] -
Continuous frame motion sensitive self-supervised collaborative network for video representation learning (2023)
Advanced Engineering Informatics 2023
Shuai Bi, Zhengping Hu, Mengyao Zhao, Hehao Zhang, Jirui Di, and Zhe Sun
[Paper] -
Self-supervised pretext task collaborative multi-view contrastive learning for video action recognition (2023)
Signal, Image and Video Processing 2023
Shuai Bi, Zhengping Hu, Mengyao Zhao, Hehao Zhang, Jirui Di, and Zhe Sun
[Paper] -
Self-Supervised Learning from Untrimmed Videos via Hierarchical Consistency (2023)
IEEE Transactions on Pattern Analysis and Machine Intelligence 2023
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Changxin Gao, Rong Jin, and Nong Sang
[Paper] -
Audio-Visual Contrastive Learning with Temporal Self-Supervision (2023)
AAAI 2023
Simon Jenni, Alexander Black, and John Collomosse
[Paper] -
Video Test-Time Adaptation for Action Recognition (2023)
CVPR 2023
Wei Lin, Muhammad Jehanzeb Mirza, Mateusz Kozinski, Horst Possegger, Hilde Kuehne, and Horst Bischof
[Paper] [Code] -
Self-Supervised Video Representation Learning via Latent Time Navigation (2023)
AAAI 2023
Di Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, and Francois Bremond
[Paper] -
Temporal Contrastive Learning with Curriculum (2023)
ICASSP 2023
Shuvendu Roy and Ali Etemad
[Paper] -
Nearest-Neighbor Inter-Intra Contrastive Learning from Unlabeled Videos (2023)
ICLR Workshops 2023
David Fan, Deyu Yang, Xinyu Li, Vimal Bhat, and Rohith MV
[Paper] -
Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization (2023)
ICCV 2023
Fida Mohammad Thoker, Hazel Doughty, and Cees Snoek
[Paper] -
Multi-scale Compositional Constraints for Representation Learning on Videos (2023)
ICASSP 2023
Georgios Paraskevopoulos, Chandrashekhar Lavania, Lovish Chum, and Shiva Sundaram
[Paper] -
Flavr: Flow-agnostic Video Representations for Fast Frame Interpolation (2023)
WACV 2023
Tarun Kalluri, Deepak Pathak, Manmohan Chandraker, and Du Tran
[Paper] -
HomE: Homography-Equivariant Video Representation Learning (2023)
arXiv / Preprint
Anirudh Sriram, Adrien Gaidon, Jiajun Wu, Juan Carlos Niebles, Li Fei-Fei, and Ehsan Adeli
[Paper] [Code] -
ViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints (2023)
WACV 2023
Srijan Das and Michael S Ryoo
[Paper] -
Videomae v2: Scaling Video Masked Autoencoders with Dual Masking (2023)
CVPR 2023
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao
[Paper] -
Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity (2023)
AAAI 2023
Pritam Sarkar, Ali Etemad
[Paper] [Code] -
Previts: contrastive pretraining with video tracking supervision (2023)
WACV 2023
Chen, B., Selvaraju, R. R., Chang, S. F., Niebles, J. C., & Naik, N.
[Paper] -
Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning (2023)
CVPR 2023
Zhang, H., Liu, D., Zheng, Q., & Su, B.
[Paper] -
Learning Fine-Grained Features for Pixel-wise Video Correspondences (2023)
ICCV 2023
Li, R., Zhou, S., & Liu, D.
[Paper] -
Cali-NCE: Boosting Cross-Modal Video Representation Learning With Calibrated Alignment (2023)
CVPR Workshops 2023
Zhao, N., Jiao, J., Xie, W., & Lin, D.
[Paper] -
Self-supervised motion perception for spatiotemporal representation learning (2023)
IEEE Transactions on Neural Networks and Learning Systems 2023
Chang Liu, Yuan Yao, Dezhao Luo, Yu Zhou, Qixiang Ye
[Paper] [Code] -
Similarity Contrastive Estimation for Image and Video Soft Contrastive Self-Supervised Learning (2023)
Machine Vision and Applications 2023
Julien Denize, Jaonary Rabarisoa, Astrid Orcesi, Romain H´erault
[Paper] -
Self-Supervised Contrastive Learning for Audio-Visual Action Recognition (2023)
ICIP 2023
Yang Liu, Ying Tan, Haoyuan Lan
[Paper] -
Self-Supervised Scene-Debiasing for Video Representation Learning via Background Patching (2023)
IEEE Transactions on Multimedia 2023
Maregu Assefa, Wei Jiang, Kumie Gedamu, Getinet Yilma, Bulbula Kumeda, Melese Ayalew
[Paper] -
LgNet: A local-global network for action recognition and beyond (2023)
IEEE Transactions on Multimedia 2023
Jiaqi Zhou, Zehua Fu, Qiuyu Huang, Qingjie Liu, Yunhong Wang
[Paper] -
Unsupervised Video-Based Action Recognition With Imagining Motion and Perceiving Appearance (2023)
IEEE Transactions on Circuits and Systems for Video Technology 2023
Wei Lin , Xiaoyu Liu , Yihong Zhuang , Xinghao Ding , Xiaotong Tu , Yue Huang , Huanqiang Zeng
[Paper] -
Spatiotemporal Augmentation on Selective Frequencies for Video Representation Learning (2023)
AAAI 2023
Jinhyung Kim, Taeoh Kim, Minho Shim, Dongyoon Han, Dongyoon Wee, Junmo Kim
[Paper] -
Consistent Intra-video Contrastive Learning with Asynchronous Long-term Memory Bank (2023)
IEEE Transactions on Circuits and Systems for Video Technology 2023
Zelin Chen, Kun-Yu Lin, Wei-Shi Zheng
[Paper]
-
BEVT: BERT Pretraining of Video Transformers (2022)
CVPR 2022
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, Lu Yuan
[Paper] -
Masked Autoencoders As Spatiotemporal Learners (2022)
NeurIPS 2022
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming He
[Paper] -
SPAct: Self-supervised Privacy Preservation for Action Recognition (2022)
CVPR 2022
Ishan Rajendrakumar Dave, Chen Chen, Mubarak Shah
[Paper] [Code] -
Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation Learning (2022)
AAAI 2022
Manlin Zhang, Jinpeng Wang, Andy J. Ma
[Paper] [Code] -
Self-supervised Video Transformer (2022)
CVPR 2022
Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, Michael S. Ryoo
[Paper] [Code] -
Exploring Relations in Untrimmed Videos for Self-Supervised Learning (2022)
ACM Transactions on Multimedia Computing, Communications, and Applications 2022
Dezhao Luo, Bo Fang, Yu Zhou, Yucan Zhou, Dayan Wu, Weiping Wang
[Paper] -
MaMiCo: Macro-to-Micro Semantic Correspondence for Self-supervised Video Representation Learning (2022)
ACM Multimedia 2022
Bo Fang, Wenhao Wu, Chang Liu, Yu Zhou, Dongliang He, Weiping Wang
[Paper] -
TCGL: Temporal Contrastive Graph for Self-Supervised Video Representation Learning (2022)
IEEE Transactions on Image Processing 2022
Yang Liu , Keze Wang , Lingbo Liu , Haoyuan Lan, and Liang Lin
[Paper] [Code] -
Cross-Architecture Self-supervised Video Representation Learning (2022)
CVPR 2022
Sheng Guo, Zihua Xiong, Yujie Zhong, Limin Wang, Xiaobo Guo, Bing Han, Weilin Huang
[Paper] -
Contrastive spatio-temporal pretext learning for self-supervised video representation (2022)
AAAI 2022
Yujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing-Yin Yu
[Paper] [Code] -
Transrank: Self-supervised video representation learning via ranking-based transformation recognition (2022)
CVPR 2022
Haodong Duan, Nanxuan Zhao, Kai Chen, Dahua Lin
[Paper] [Code] -
Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency (2022)
CVPR 2022
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Mingqian Tang, Changxin Gao, Rong Jin,Nong Sang
[Paper] [Code] -
Motion-aware contrastive video representation learning via foreground-background merging (2022)
CVPR 2022
Shuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian, Haohang Xu, Qingyi Chen, Jue Wang, Hongkai Xiong
[Paper] [Code] -
Self-Supervised Video Representation Learning with Motion-Contrastive Perception (2022)
ICME 2022
Jinyu Liu, Ying Cheng, Yuejie Zhang, Rui-Wei Zhao, Rui Feng
[Paper] -
Self-supervised video representation learning using improved instance-wise contrastive learning and deep clustering (2022)
IEEE Transactions on Circuits and Systems for Video Technology 2022
Yisheng Zhu, Hui Shuai, Guangcan Liu, Senior Member, Qingshan Liu
[Paper] -
TCLR: Temporal contrastive learning for video representation (2022)
Computer Vision and Image Understanding 2022
Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, Mubarak Shah
[Paper] [Code] -
Self-supervised spatiotemporal representation learning by exploiting video continuity (2022)
AAAI 2022
Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai, Juwei Lu, Yang Wang
[Paper] -
Probabilistic representations for video contrastive learning (2022)
CVPR 2022
Jungin Park, Jiyoung Lee, Ig-Jae Kim, Kwanghoon Sohn
[Paper] -
Contextualized spatio-temporal contrastive learning with self-supervision (2022)
CVPR 2022
Liangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong,Florian Schroff,Ming-Hsuan Yang, Hartwig Adam, Ting Liu
[Paper] [Code] -
VideoMAE: Masked Autoencoders Are Data-Efficient Learners for Self-Supervised Video Pre-Training (2022)
NeurIPS 2022
Zhan Tong, Yibing Song, Jue Wang, Limin Wang
[Paper] [Code] -
Self-supervised video representation learning with cross-stream prototypical contrasting (2022)
WACV 2022
Martine Toering, Ioannis Gatopoulos, Maarten Stol, Vincent Tao Hu
[Paper] [Code] -
SLIC: Self-supervised learning with iterative clustering for human action videos (2022)
CVPR 2022
Salar Hosseini Khorasgani, Yuxuan Chen, Florian Shkurti
[Paper] -
GOCA: guided online cluster assignment for self-supervised video representation Learning (2022)
ECCV 2022
Huseyin Coskun, Alireza Zareian, Joshua L. Moore, Federico Tombari, Chen Wang
[Paper] [Code] -
TCVM: Temporal Contrasting Video Montage Framework for Self-supervised Video Representation Learning (2022)
ACCV 2022
Fengrui Tian, Jiawei Fan, Xie Yu, Shaoyi Du, Meina Song, Yu Zhao
[Paper] -
Static and Dynamic Concepts for Self-supervised Video Representation Learning (2022)
ECCV 2022
Rui Qian, Shuangrui Ding, Xian Liu, Dahua Lin
[Paper] -
SOS! Self-supervised Learning over Sets of Handled Objects in Egocentric Action Recognition (2022)
ECCV 2022
Victor Escorcia, Ricardo Guerrero, Xiatian Zhu, Brais Martinez
[Paper] -
Self-Supervised Video Representation Learning with Cascade Positive Retrieval (2022)
CVPR Workshops 2022
Cheng-En Wu, Farley Lai, Yu Hen Hu, Asim Kadav
[Paper] [Code] -
Self-Supervised Learning of Audio Representations From Audio-Visual Data Using Spatial Alignment (2022)
IEEE Journal of Selected Topics in Signal Processing 2022
Shanshan Wang, Archontis Politis, Annamaria Mesaros
[Paper] -
Hierarchically decoupled spatial-temporal contrast for self-supervised video representation learning (2022)
WACV 2022
Zehua Zhang, David Crandall
[Paper] -
Spatio-temporal self-supervision enhanced transformer networks for action recognition (2022)
ICME 2022
Yongkang Zhang, Han Zhang, Guoming Wu, Jun Li
[Paper] -
Inter-Intra Cross-Modality Self-Supervised Video Representation Learning by Contrastive Clustering (2022)
ICPR 2022
Jiutong Wei. Guan Luo, Bing Li, Weiming Hu
[Paper] -
SCVRL: Shuffled Contrastive Video Representation Learning (2022)
CVPR Workshops 2022
Michael Dorkenwald, Fanyi Xiao, Biagio Brattoli, Joseph Tighe, Davide Modolo
[Paper] -
InternVideo: General Video Foundation Models via Generative and Discriminative Learning (2022)
arXiv / Preprint
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang,Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu,Yali Wang, Limin Wang, Yu Qiao
[Paper] [Code] -
Video Motion Perception for Self-supervised Representation Learning (2022)
ICANN 2022
Wei Li, Dezhao Luo, Bo Fang, Xiaoni Li, Yu Zhou, Weiping Wang
[Paper] -
An improved inter-intra contrastive learning framework on self-supervised video representation (2022)
IEEE Transactions on Circuits and Systems for Video Technology 2022
Li Tao, Xueting Wang, Toshihiko Yamasaki
[Paper] -
Auxiliary Learning for Self-Supervised Video Representation via Similarity-based Knowledge Distillation (2022)
CVPR Workshops 2022
Amirhossein Dadashzadeh, Alan Whone, Majid Mirmehdi
[Paper] [Code] -
Motion Sensitive Contrastive Learning for Self-supervised Video Representation (2022)
ECCV 2022
Jingcheng Ni, Nan Zhou, Jie Qin, Qian Wu, Junqi Liu, Boxun Li, Di Huang
[Paper] -
Unsupervised Learning of Spatio-Temporal Representation with Multi-Task Learning for Video Retrieval (2022)
National Conference on Communications 2022
Vidit Kumar
[Paper] -
Federated Self-supervised Learning for Video Understanding (2022)
ECCV 2022
Yasar Abbas Ur Rehman, Yan Gao, Jiajun Shen, Pedro Porto Buarque de Gusmão , Nicholas Lane
[Paper] [Code] -
Contrastive predictive coding with transformer for video representation learning (2022)
Neurocomputing 2022
Yue Liu, Junqi Ma, Yufei Xie, Xuefeng Yang, Xingzhen Tao, Lin Peng, Wei Gao
[Paper] [Code] -
Video representation learning by identifying spatio-temporal transformation (2022)
Applied Intelligence 2022
Sheng Geng, Shimin Zhao , Hu Liu
[Paper] -
On temporal granularity in self-supervised video representation learning (2022)
BMVC 2022
Rui Qian, Yeqing Li, Liangzhe Yuan, Boqing Gong, Ting Liu, Matthew Brown, Serge Belongie, Ming-Hsuan Yang, Hartwig Adam, and Yin Cui
[Paper] [Code] -
LAVA: Language Audio Vision Alignment for Data-Efficient Video Pre-Training (2022)
ICML Pre-training Workshop 2022
Sumanth Gurram , Andy Fang , David Chan , John Canny
[Paper] -
It Takes Two: Masked Appearance-Motion Modeling for Self-supervised Video Transformer Pre-training (2022)
arXiv / Preprint
Yuxin Song, Min Yang, Wenhao Wu, Dongliang He, Fu Li, Jingdong Wang
[Paper] -
MAC: Mask-Augmentation for Motion-Aware Video Representation Learning (2022)
BMVC 2022
Arif Akar, Ufuk Umut Senturk, and Nazli Ikizler-Cinbis.
[Paper] [Code] -
Temporal-Invariant Video Representation Learning with Dynamic Temporal Resolutions. (2022)
AVSS 2022
Seong-Yun Jeong, Ho-Joong Kim, Myeong-Seok Oh, Gun-Hee Lee, Seong-Whan Lee
[Paper] -
Dual Contrastive Learning for Spatio-temporal Representation (2022)
ACM Multimedia 2022
Shuangrui Ding,, Rui Qian, and Hongkai Xiongo
[Paper] -
MoQuad: Motion-focused Quadruple Construction for Video Contrastive Learning (2022)
ECCV Workshops 2022
Yuan Liu, Jiacheng Chen, Hao Wu
[Paper] -
On Negative Sampling for Audio-Visual Contrastive Learning from Movies (2022)
arXiv / Preprint
Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu, Nagendra Kamath, Ashok Chandrashekar
[Paper] -
Frame-wise Action Representations for Long Videos via Sequence Contrastive Learning (2022)
CVPR 2022
Minghao Chen, Fangyun Wei, Chong Li, Deng Cai
[Paper] [Code] -
Masked Feature Prediction for Self-Supervised Visual Pre-Training (2022)
CVPR 2022
Wei, C., Fan, H., Xie, S., Wu, C. Y., Yuille, A., & Feichtenhofer, C.
[Paper] -
Pixel-level Correspondence for Self-Supervised Learning from Video (2022)
arXiv / Preprint
Yash Sharma, Yanchao Zhu, Chris Russell, Thomas Brox
[Paper] -
Temporal Alignment Networks for Long-Term Video (2022)
CVPR 2022
Han, T., Xie, W., & Zisserman, A.
[Paper] -
SimVTP: Simple Video Text Pre-Training with Masked Autoencoders (2022)
arXiv / Preprint
Ma, Y., Yang, T., Shan, Y., & Li, X.
[Paper] -
Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction (2022)
ICLR 2022
Shi, B., Hsu, W. N., Lakhotia, K., & Mohamed, A.
[Paper] -
Self-supervised video representation learning by uncovering spatio-temporal statistics (2022)
IEEE Transactions on Pattern Analysis and Machine Intelligence 2022
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Wei Liu, Yun-hui Liu
[Paper] [Code]
-
Inter-intra Variant Dual Representations for Self-supervised Video Recognition (2021)
BMVC 2021
Lin Zhang, Qi She, Zhengyang Shen, Changhu Wang
[Paper] -
VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning (2021)
NeurIPS 2021
Hao Tan, Jie Lei, Thomas Wolf, Mohit Bansal
[Paper] -
Watching too much television is good: Self-supervised audio-visual representation learning from movies and tv shows (2021)
arXiv / Preprint
Mahdi M. Kalayeh, Nagendra Kamath, Lingyi Liu
[Paper] -
Temporally coherent embeddings for self-supervised video representation learning (2021)
ICPR 2020
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, Peyman Moghadam
[Paper] [Code] -
Audio-visual instance discrimination with cross-modal agreement (2021)
CVPR 2021
Pedro Morgado, Nuno Vasconcelos, Ishan Misra
[Paper] [Code] -
Removing the background by adding the background: Towards background robust self-supervised video representation learning (2021)
CVPR 2021
Jinpeng Wang, Yuting Gao, Ke Li, Yiqi Lin, Andy J. Ma, Hao Cheng, Pai Peng, Feiyue Huang, Rongrong Ji, Xing Sun
[Paper] [Code] -
Enhancing unsupervised video representation learning by decoupling the scene and the motion (2021)
AAAI 2021
Jinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu, Xinyang Jiang, Xiaowei Guo, Rongrong Ji, Xing Sun
[Paper] [Code] -
SeCo: Exploring Sequence Supervision for Unsupervised Representation Learning (2021)
AAAI 2021
Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, Tao Mei
[Paper] [Code] -
Enhancing self-supervised video representation learning via multi-level feature optimization (2021)
ICCV 2021
Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, Weiyao Lin
[Paper] [Code] -
RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning (2021)
AAAI 2021
Peihao Chen, Deng Huang, Dongliang He, Xiang Long, Runhao Zeng, Shilei Wen, Mingkui Tan, Chuang Gan
[Paper] [Code] -
VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples (2021)
CVPR 2021
Tian Pan, Yibing Song, Tianyu Yang, Wenhao Jiang, Wei Liu
[Paper] [Code] -
On compositions of transformations in contrastive self-supervised learning (2021)
ICCV 2021
Mandela Patrick, Yuki M. Asano, Polina Kuznetsova, Ruth Fong, João F. Henriques, Geoffrey Zweig, Andrea Vedaldi
[Paper] [Code] -
Unsupervised visual representation learning by tracking patches in video (2021)
CVPR 2021
Guangting Wang, Yizhou Zhou, Chong Luo, Wenxuan Xie, Wenjun Zeng, Zhiwei Xiong
[Paper] [Code] -
A large-scale study on unsupervised spatiotemporal representation learning (2021)
CVPR 2021
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, Kaiming He
[Paper] [Code] -
CoCon: Cooperative-Contrastive Learning (2021)
CVPR 2021
Nishant Rai, Ehsan Adeli ,Kuan-Hui Lee, Adrien Gaidon, Juan Carlos Niebles
[Paper] [Code] -
VATT: Transformers for multimodal self-supervised learning from raw video, audio and text (2021)
NeurIPS 2021
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, Boqing Gong
[Paper] [Code] -
ASCNet: Self-supervised video representation learning with appearance-speed consistency (2021)
ICCV 2021
Deng Huang, Wenhao Wu, Weiwen Hu, Xu Liu, Dongliang He, Zhihua Wu, Xiangmiao Wu, Mingkui Tan, Errui Ding
[Paper] -
Self-supervised visual learning by variable playback speeds prediction of a video (2021)
IEEE Access 2021
Hyeon Cho, Taehoon Kim, Hyungjin Chang, Wonjun Hwang
[Paper] [Code] -
Self-supervised video representation learning with meta-contrastive network (2021)
ICCV 2021
Yuanze Lin, Xun Guo, Yan Lu
[Paper] -
Long short view feature decomposition via contrastive video representation learning (2021)
ICCV 2021
Nadine Behrmann, Mohsen Fayyaz, Juergen Gall, Mehdi Noroozi
[Paper] -
Time-equivariant contrastive video representation learning (2021)
ICCV 2021
Simon Jenni, Hailin Jin
[Paper] -
Self-supervised video representation learning by context and motion decoupling (2021)
CVPR 2021
Lianghua Huang, Yu Liu, Bin Wang, Pan Pan, Yinghui Xu, Rong Jin
[Paper] -
Unsupervised video representation learning by bidirectional feature prediction (2021)
WACV 2021
Nadine Behrmann, Juergen Gall, Mehdi Noroozi
[Paper] -
Self-supervised learning of compressed video representations (2021)
ICLR 2021
Youngjae Yu, Sangho Lee, Gunhee Kim, Yale Song
[Paper] -
Spatiotemporal contrastive video representation learning (2021)
CVPR 2021
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, Yin Cui
[Paper] [Code] -
MoDist: Motion Distillation for Self-Supervised Video Representation Learning (2021)
arXiv / Preprint
Fanyi Xiao, Joseph Tighe, Davide Modolo
[Paper] -
Broaden your views for self-supervised video learning (2021)
ICCV 2021
Adria Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Ross Hemsley, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altche, Michal Valko, Jean-Bastien Grill, Aaron van den Oord, Andrew Zisserman
[Paper] [Code] -
Vi2CLR: Video and image for visual contrastive learning of representation (2021)
ICCV 2021
Ali Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi, M. Saquib Sarfraz,Rainer Stiefelhagen, Luc Van Gool,
[Paper] -
Contrast and order representations for video self-supervised learning (2021)
ICCV 2021
Kai Hu, Jie Shao, Yuan Liu, Bhiksha Raj, Marios Savvides, Zhiqiang Shen
[Paper] -
Motion-augmented self-training for video recognition at smaller scale (2021)
ICCV 2021
Kirill Gavrilyuk, Mihir Jain, Ilia Karmanov, Cees G. M. Snoek
[Paper] -
Video contrastive learning with global context (2021)
ICCV Workshops 2021
Haofei Kuang, Yi Zhu, Zhi Zhang, Xinyu Li, Joseph Tighe,Soren Schwertfeger, Cyrill Stachniss, Mu Li
[Paper] [Code] -
Motion-focused contrastive learning of video representations (2021)
ICCV 2021
Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao, Dong Liu, and Tao Mei
[Paper] [Code] -
Back to the Future: Cycle Encoding Prediction for Self-supervised Video Representation Learning (2021)
BMVC 2021
Xinyu Yang, Majid Mirmehdi,Tilo Burghardt
[Paper] [Code] -
Composable augmentation encoding for video representation learning (2021)
ICCV 2021
Sun, C., Nagrani, A., Tian, Y., & Schmid, C.
[Paper] -
Learning temporal dynamics from cycles in narrated video (2021)
ICCV 2021
Epstein, D., Wu, J., Schmid, C., & Sun, C.
[Paper] -
CrossCLR: Cross-Modal Contrastive Learning for Multi-Modal Video Representations (2021)
ICCV 2021
Zolfaghari, M., Zhu, Y., Gehler, P., & Brox, T.
[Paper] -
Watching the World Go By: Representation Learning from Unlabeled Videos (2021)
ICLR 2021
Daniel Gordon, Kiana Ehsani, Dieter Fox, Ali Farhadi
[Paper] -
Parameter Efficient Multimodal Transformers for Video Representation Learning (2021)
ICLR 2021
Lee, S., Yu, Y., Kim, G., Breuel, T., Kautz, J., & Song, Y.
[Paper] -
Active Contrastive Learning of Audio-Visual Video Representations (2021)
ICLR 2021
Ma, S., Zeng, Z., McDuff, D., & Song, Y.
[Paper]
-
Self-Supervised Learning to Detect Key Frames in Videos (2020)
Sensors 2020
Xiang Yan,Syed Zulqarnain Gilani,Mingtao Feng ,Liang Zhang,Hanlin Qin and Ajmal Mian
[Paper] -
Self-supervised motion representation via scattering local motion cues (2020)
ECCV 2020
Yuan Tian, Zhaohui Che, Wenbo Bao, Guangtao Zhai, Zhiyong Gao1
[Paper] -
Self-supervised video representation learning using inter-intra contrastive framework (2020)
ACM Multimedia 2020
Li Tao, Xueting Wang, Toshihiko Yamasaki
[Paper] [Code] -
Video representation learning with visual tempo consistency (2020)
arXiv / Preprint
Ceyuan Yang, Yinghao Xu, Bo Dai, Bolei Zhou
[Paper] [Code] -
Self-supervised temporal discriminative learning for video representation learning (2020)
arXiv / Preprint
Jinpeng Wang, Yiqi Lin, Andy J. Ma,Pong C. Yuen
[Paper] [Code] -
Self-supervised learning by cross-modal audio-video clustering (2020)
NeurIPS 2020
Humam Alwassel, Dhruv Mahajan, Bruno Korbar ,Lorenzo Torresani, Bernard Ghanem, Du Tran
[Paper] [Code] -
Self-supervised video representation learning by pace prediction (2020)
ECCV 2020
Jiangliu Wang, Jianbo Jiao, Yun-Hui Liu
[Paper] [Code] -
Unsupervised learning from video with deep neural embeddings (2020)
CVPR 2020
Chengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark, Daniel Yamins
[Paper] [Code] -
Unsupervised learning of video representations via dense trajectory clustering (2020)
ECCV Workshops 2020
Pavel Tokmakov, Martial Hebert, Cordelia Schmid
[Paper] [Code] -
Video representation learning by recognizing temporal transformations (2020)
ECCV 2020
Simon Jenni, Givi Meishvili, Paolo Favaro
[Paper] [Code] -
Video playback rate perception for self-supervised spatio-temporal representation learning (2020)
CVPR 2020
Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou, Qixiang Ye
[Paper] [Code] -
Self-supervised co-training for video representation learning (2020)
NeurIPS 2020
Tengda Han, Weidi Xie, Andrew Zisserman
[Paper] [Code] -
Video cloze procedure for self-supervised spatio-temporal learning (2020)
AAAI 2020
Dezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang, Can Ma, Qixiang Ye, Weiping Wang
[Paper] [Code] -
End-to-end learning of visual representations from uncurated instructional videos (2020)
CVPR 2020
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira,Ivan Laptev, Josef Sivic, Andrew Zisserman
[Paper] [Code] -
SpeedNet: Learning the Speediness in Videos (2020)
CVPR 2020
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, Tali Dekel
[Paper] [Code] -
Contrastive multiview coding (2020)
ECCV 2020
Yonglong Tian, Dilip Krishnan, Phillip Isola
[Paper] [Code] -
Self-supervised video representation learning by maximizing mutual information (2020)
Signal Processing: Image Communication 2020
Fei Xue, Hongbing Ji, Wenbo Zhang, Yi Cao
[Paper] -
Memory-augmented dense predictive coding for video representation learning (2020)
ECCV 2020
Tengda Han, Weidi Xie, Andrew Zisserman
[Paper] [Code] -
Evolving losses for unsupervised video representation learning (2020)
CVPR 2020
AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo
[Paper] -
AudioVisual SlowFast Networks for Video Recognition (2020)
arXiv / Preprint
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, Christoph Feichtenhofer
[Paper] [Code] -
Cycle-Contrast for Self-Supervised Video Representation Learning (2020)
NeurIPS 2020
Quan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga, Tomokazu Murakami
[Paper] -
Can temporal information help with contrastive self-supervised learning? (2020)
arXiv / Preprint
Yutong Bai, Haoqi Fan, Ishan Misra, Ganesh Venkatesh, Yongyi Lu
[Paper] -
Self-supervised multimodal versatile networks (2020)
NeurIPS 2020
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira Sander Dieleman, Andrew Zisserman
[Paper] -
Pretext-Contrastive Learning: Toward Good Practices in Self-Supervised Video Representation Learning (2020)
arXiv / Preprint
Li Tao, Xueting Wang, Toshihiko Yamasaki
[Paper] [Code] -
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation (2020)
arXiv / Preprint
Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., ... & Zhou, M.
[Paper] -
Self-supervised learning of audio-visual objects from video (2020)
ECCV 2020
Afouras, T., Owens, A., Chung, J. S., & Zisserman, A.
[Paper] -
Speech2Action: Cross-Modal Supervision for Action Recognition (2020)
CVPR 2020
Nagrani, A., Sun, C., Ross, D., Sukthankar, R., Schmid, C., & Zisserman, A.
[Paper] -
Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning (2020)
ACM Multimedia 2020
Cheng, Y., Wang, R., Pan, Z., Feng, R., & Zhang, Y.
[Paper]
-
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics (2019)
CVPR 2019
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, Wei Liu
[Paper] [Code] -
Video representation learning by dense predictive coding (2019)
ICCV Workshops 2019
Tengda Han, Weidi Xie, Andrew Zisserman
[Paper] [Code] -
Self-supervised spatiotemporal learning via video clip order prediction (2019)
CVPR 2019
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, Yueting Zhuang
[Paper] [Code] -
Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition (2019)
WACV 2019
Unaiza Ahsan, Rishi Madhok, Irfan Essa
[Paper] -
Self-supervised video representation learning with space-time cubic puzzles (2019)
AAAI 2019
Dahun Kim, Donghyeon Cho, In So Kweon
[Paper] -
Learning Video Representations Using Contrastive Bidirectional Transformer (2019)
arXiv / Preprint
Chen Sun, Fabien Baradel, Kevin Murphy, Cordelia Schmid
[Paper] -
DynamoNet: Dynamic Action and Motion Network (2019)
ICCV 2019
Ali Diba, Vivek Sharma, Luc Van Gool, Rainer Stiefelhagen
[Paper] -
Temporal Cycle-Consistency Learning (2019)
CVPR 2019
Dwibedi, D., Aytar, Y., Tompson, J., Sermanet, P., & Zisserman, A.
[Paper] -
VideoBERT: A Joint Model for Video and Language Representation Learning (2019)
ICCV 2019
Sun, C., Myers, A., Vondrick, C., Murphy, K., & Schmid, C.
[Paper]
-
Geometry Guided Convolutional Neural Networks for Self-Supervised Video Representation Learning (2018)
CVPR 2018
Chuang Gan, Boqing Gong, Kun Liu, Hao Su, Leonidas J. Guibas
[Paper] -
Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction (2018)
arXiv / Preprint
Longlong Jing, Xiaodong Yang, Jinggen Liu, Yingli Tian
[Paper] -
Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization (2018)
NeurIPS 2018
Bruno Korbar, Du Tran, Lorenzo Torresani
[Paper] -
Audio-Visual Scene Analysis with Self-Supervised Multisensory Features (2018)
ECCV 2018
Andrew Owens, Alexei A. Efros
[Paper] [Code] -
Compressed Video Action Recognition (2018)
CVPR 2018
Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R. Manmatha, Alexander J. Smola, Philipp Krahenb
[Paper] -
Improving Spatiotemporal Self-Supervision by Deep Reinforcement Learning (2018)
ECCV 2018
Uta Buchler, Biagio Brattoli, Bjorn Ommer
[Paper] -
Learning and Using the Arrow of Time (2018)
CVPR 2018
Donglai Wei, Joseph Lim, Andrew Zisserman, William T. Freeman
[Paper]
-
Unsupervised Representation Learning by Sorting Sequences (2017)
ICCV 2017
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, Ming-Hsuan Yang
[Paper] -
Self-Supervised Video Representation Learning With Odd-One-Out Networks (2017)
CVPR 2017
Basura Fernando, Hakan Bilen, Efstratios Gavves, Stephen Gould
[Paper]
- Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification (2016)
ECCV 2016
Ishan Misra, C. Lawrence Zitnick, Martial Hebert
[Paper]
Video self-supervised learning is a family of methods that learns useful spatial and temporal representations from videos without requiring a manually annotated label for every training example. Common objectives include masked reconstruction, contrastive learning, temporal prediction, latent feature prediction, motion modeling, cross-modal learning, and knowledge distillation.
Video SSL and VideoSSL are common abbreviations for video self-supervised learning. In research literature the same area is also described as self-supervised video representation learning, video representation pretraining, masked video modeling, and self-supervised action recognition.
Frequently used benchmarks include UCF101, HMDB51, Kinetics-400, Something-Something V1 and V2, Diving48, and EPIC-KITCHENS. Different benchmarks emphasize appearance, motion, temporal reasoning, egocentric activity understanding, or fine-grained actions.
Major families include contrastive and non-contrastive representation learning, masked video autoencoding and masked feature modeling, predictive and JEPA-style learning, motion-aware objectives, audio-visual or video-language self-supervision, temporal-order and transformation prediction, and fine-grained correspondence learning.
