šŸ“ Selected Publications

CVPRW 2026
sym

CAST: Training a Student Expert via Semi-Supervised Foundation Model Distillation Pardis Taghavi, Tian Liu, Renjie Li, Reza Langari, Zhengzhong Tu paper | arXiv | project page

  • CAST is a semi‐supervised knowledge distillation (SSKD) framework that compresses pretrained vision‐foundation models (VFMs) into compact expert networks by leveraging limited labeled data and abundant unlabeled data via stage‐wise fine‐tuning coupled with a contrastive self‐supervised loss.
IROS 2024
sym

SwinMTL: A Shared Architecture for Simultaneous Depth Estimation and Semantic Segmentation from Monocular Camera Images Pardis Taghavi, Reza Langari, Gaurav Pandey code | arXiv

  • A simple and effective multi-task learning framework that allows concurrent depth estimation and semantic segmentation using a single camera and without compromising computational efficiency.
arXiv 2026
sym

The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics Xiangbo Gao, Mingyang Wu, Siyuan Yang, Jiongze Yu, Pardis Taghavi, Fangzhou Lin, Zhengzhong Tu arXiv | project page

  • Proposes Visual Chronometer, a predictor that recovers Physical Frames Per Second (PhyFPS) from visual dynamics to address chronometric hallucination in generative video models and improve perceived motion naturalness.
arXiv 2026
sym

NaviDriveVLM: Decoupling High-Level Reasoning and Motion Planning for Autonomous Driving Ximeng Tao, Pardis Taghavi, Dimitar Filev, Reza Langari, Gaurav Pandey arXiv

  • A decoupled Navigator-Driver framework that separates semantic reasoning from waypoint prediction, preserving strong VLM reasoning while enabling efficient adaptation for end-to-end motion planning on nuScenes.