I am an Associate Professor at the School of Intelligence Science and Technology, Nanjing University. My research is currently on visual understanding in an open world, multimodal learning and domain-specific multimodal large models. Previously, I was a senior researcher (T11) in Tencent AI Lab from 2021 to 2023, and a research scientist in Inception Institute of Artificial Intelligence (IIAI) from 2018 to 2021, where I worked with Prof. Shengcai Liao. I was a research fellow in National University of Singapore (NUS) from 2015 to 2017, co-supervised by Prof. Jiashi Feng and Prof. Shuicheng Yan. I received my Ph.D. degree from Institute of Automation, Chinese Academy of Sciences (CASIA) under the supervision of Prof. Liang Wang.

News

  • 每年可招收博士生2名、硕士生3-4名,欢迎感兴趣的同学申请与报考!

Selected Papers

  • S. Zhang, W. Xu, Y. Fang, F. Lyu, L. Ma, G. Zhao, F. Zhao, C. Shan, L. Wang. Reconstructive Visual Tuning for Weakly Supervised Video Anomaly Detection. IEEE Transactions on Information Forensics and Security (TIFS), 2026.
  • X. Xu, C. Fu, X. Wang, S. Liu, C. Shan, F. Zhao. ACE-ing Video Corpus Moment Retrieval: An Automated Dataset Construction and Unified Retrieval-Localization Framework. ACM International Conference on Multimedia (ACM MM), 2026.
  • X. Liu, W. Li, T. Chen, H. Wang, G. Xie, C. Shan, F. Zhao. QST-SAM: Leveraging Cross-modal Instructions for Few-shot Referring Video Object Segmentation. European Conference on Computer Vision (ECCV), 2026.
  • W. Dong, Z. Wang, S. Zhang, K. Sun, B. Li, G. Xie, C. Shan, F. Zhao. CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection. European Conference on Computer Vision (ECCV), 2026.
  • R. Deng*, Y. Hu*, Y. Zhong*, Z. Wang, X. Liu, H. Wang, C. Shan, F. Zhao. Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection. European Conference on Computer Vision (ECCV), 2026.
  • S. Zhang, L. Ma, Z. Wang, W. Dong, X. Xu, G. Xie, C. Shan, F. Zhao. TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning. International Conference on Machine Learning (ICML), 2026.
  • H. Rao, Z. Wang, C. Si, Y. Lyu, Y. Duan, F. Zhao, C. Shan. One-to-More: High-Fidelity Training-Free Anomaly Generation with Attention Control. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2026. (Highlight)
  • X. Xu, H. Wang, G. Xie, C. Shan, F. Zhao. HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal Grounding. International Conference on Learning Representations (ICLR), 2026.
  • C. Zhang, Y. Lin, Y. Wei, H. Wang, C. Shan, F. Zhao. Matting Anything 2: Towards Video Matting for Anything. International Conference on Learning Representations (ICLR), 2026.
  • Y. Duan, W. Xu, Q. Wu, G. Xie, F. Zhao, C. Shan. AnomalyControl: Highly-Aligned Anomalous Image Generation with Controlled Diffusion Model. ACM International Conference on Multimedia (ACM MM), 2025.
  • H. Wang, W. Weng, J. Wang, F. Zhao, G. Xie, X. Geng, L. Wang. Foundation model for skeleton-based human action understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025.
  • F. Zhao, Z. Li, S. Huang, J. Weng, T. Zhou, G. Xie, J. Wang, Y. Shan. Learning anchor transformations for 3d garment animation. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  • W. Wang, F. Zhao, S. Liao, L. Shao. Attentive waveblock: complementarity-enhanced mutual networks for unsupervised domain adaptation in person re-identification and beyond. IEEE Transactions on Image Processing (TIP), 2022.
  • F. Zhao, W. Wang, S. Liao, L. Shao. Learning anchored unsigned distance functions with gradient direction alignment for single-view garment reconstruction. IEEE International Conference on Computer Vision (ICCV), 2021. (Oral)
  • F. Zhao, S. Liao, K. Zhang, L. Shao. Human parsing based texture transfer from single image to 3D human via cross-view consistency. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  • F. Zhao, S. Liao, G. Xie, J. Zhao, K. Zhang, L. Shao. Unsupervised domain adaptation with noise resistible mutual-training for person re-identification. European Conference on Computer Vision (ECCV), 2020.
  • F. Zhao, J. Li, J. Zhao, J. Feng. Weakly supervised phrase localization with multi-scale anchored transformer network. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  • F. Zhao*, J. Zhao*, S. Yan, J. Feng. Dynamic conditional networks for few-shot learning. European Conference on Computer Vision (ECCV), 2018.
  • F. Zhao, J. Feng, J. Zhao, W. Yang, S. Yan. Robust LSTM-autoencoders for face de-occlusion in the wild. IEEE Transactions on Image Processing (TIP), 2018.
  • F. Zhao, Y. Huang, L. Wang, T. Xiang, T. Tan. Learning relevance restricted Boltzmann machine for unstructured group activity and event understanding. International Journal of Computer Vision (IJCV), 2016.
  • F. Zhao, Y. Huang, L. Wang, T. Tan. Deep semantic ranking based hashing for multi-label image retrieval. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. (Google Scholar Citation 700+)
  • F. Zhao, Y. Huang, L. Wang, T. Tan. Relevance topic model for unstructured social group activity recognition. Advances in Neural Information Processing Systems (NeurIPS), 2013.
  • For more papers, please kindly refer to my Google Scholar page.