top of page
UNC_Zoom_Backgrounds_C.jpg

Gedas Bertasius

Assistant Professor

x-social-media-black-icon.png

I am an Assistant Professor in the Computer Science department at the University of North Carolina, Chapel Hill. Before joining UNC, I was a postdoctoral researcher at Facebook AI Research (FAIR) working with Lorenzo Torresani. I received my Ph.D. at the University of Pennsylvania, where I was advised by Jianbo Shi, and my undergraduate degree at Dartmouth College.

Research

I lead the Multimodal Video Perception (MVP) group at UNC. The long-term goal of our research is to develop machines that understand human behavior from video. We pursue this goal across three levels: (1) building scalable representations of long, multimodal video; (2) developing methods that infer the intentions, capabilities, and decisions behind observed behavior; and (3) designing systems that translate these inferences into coaching, task assistance, and embodied action. Representative projects include TimeSformer, Video ReCap, LLoVi, VideoTree,ExAct, SVI-Bench, and WatchAct. ​​

Video Recognition

Designing the core architectures for video recognition, from short actions to long-form activities (e.g., TimeSformer, ViS4mer)

Multimodal AI

Building models that jointly understand video, audio, and language, from hour-long captioning to long-range question answering (e.g., Video ReCap, LLoVi, BIMBA).

Perceptual Assistants & Coaches

Strategic Video Intelligence

Developing AI assistants that guide people through everyday tasks and coach skill learning (e.g., VidAssist, Ego-Exo4D, and ExAct).

Video for Robotics

Building models that reason about goals, decisions, and outcomes, with team sports as a dynamic microworld (e.g., SVI-BenchBASKET).

Generative Video Modeling

Teaching robots to act by watching people, from behavior-grounded manipulation to robust long-horizon execution (e.g., WatchAct, BOSS, ReBot, and ARCADE)

Generating and editing video and audio together, from video-to-music generation to audio-visual editing (e.g., VMAsAvEDV2M-Zero, and TeDiO

Recent News

Contact

Selected Projects

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

Baiqi Li, Ce Zhang, Yu Fang, Yue Yang, Shangzhe Li, Mingyu Ding, Gedas Bertasius

arXiv 2026

[arxiv] [project page] [code] [data] [bibtex​​​

SiLVR: A Simple Language-based Video Reasoning Framework

Ce Zhang, Yan-Bo Lin, Ziyang Wang, Mohit Bansal, Gedas Bertasius

       TMLR 2026 (1st Place, CVPR Multi-Discipline Lecture Understanding Challenge)

[arxiv] [project page] [code] [bibtex

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Gedas Bertasius, ... , Michael Wray

CVPR 2024

[arxiv] [project website] [blog] [video] [bibtex

Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, Gedas Bertasius

      CVPR 2024 (Egocentric Vision Distinguished Paper Award)

[arxiv] [project website] [code] [dataset[bibtex

Is Space-Time Attention All You Need for Video Understanding?

Gedas Bertasius, Heng Wang, Lorenzo Torresani

      ICML 2021 (Top-5 Most Cited ICML 2021 Paper)

[arxiv] [code] [talk] [slides] [blog] [VentureBeat] [SiliconAngle] [bibtex]

Sponsors

We are grateful for the following agencies for supporting our research.

Contact

Prospective Graduate Students: I am recruiting motivated students in computer vision. Please email me a list of your prior publications and your CV.

Undergraduates at UNC: If you are interested in computer vision, especially its applications to sports, email me your CV and transcript with your GPA.

©2024 by Gedas Bertasius

bottom of page