vix.ing · top · new · best · stats · spec

Fan, Zhaoxin

  1. MemOS: A Memory OS for AI System
    2025/07/04 by Zhiyu Li, Chunyan Xi, Li, Zhiyu +79 · 3 voices · 23 citations
    Computer Science · Engineering · #Embedded Systems Design Techniques #Parallel Computing and Optimization Techniques #Robotics and Automated Systems #cs.CL
  2. EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation
    2023/03/20 by Ziqiao Peng, Peng, Ziqiao, Haoyu Wu +13 · 24 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Facial Nerve Paralysis Treatment and Research #Generative Adversarial Networks and Image Synthesis #Sound (cs.SD) #electronic engineering #information engineering
  3. SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
    2023/11/29 by Ziqiao Peng, Wentao Hu, Peng, Ziqiao +15 · 21 citations
    Computer Science · Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Facial Nerve Paralysis Treatment and Research #Generative Adversarial Networks and Image Synthesis
  4. EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers
    2024/12/29 by Daiheng Gao, Shilin Lu, Gao, Daiheng +18 · 22 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning and Data Classification #Topic Modeling
  5. SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking Faces
    2023/06/19 by Peng, Ziqiao, Luo, Yihao, Shi, Yue +5 · 10 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  6. A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
    2025/04/22 by Kun Wang, Wang, Kun, Guibin Zhang +193 · 30 citations
    Computer Science · Health Professions · Medicine · #Adversarial Robustness in Machine Learning #Occupational Health and Safety Research #Artificial Intelligence in Healthcare and Education
  7. CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
    2024/12/09 by Gong, Zhefei, Ding, Pengxiang, Lyu, Shangke +5 · 12 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
  8. Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation From Monocular RGB Image
    2022/04/04 by Fan, Zhaoxin, Song, Zhenbo, Xu, Jian +4 · 4 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition
    2021/05/01 by Zhaoxin Fan, Fan, Zhaoxin, Zhenbo Song +9 · 3 citations
    Engineering · Earth and Planetary Sciences · #Robotics and Sensor-Based Localization #3D Surveying and Cultural Heritage #Indoor and Outdoor Localization Technologies
  10. Deep Learning on Monocular Object Pose Detection and Tracking: A Comprehensive Overview
    2021/05/29 by Fan, Zhaoxin, Zhu, Yazhi, He, Yulin +3 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  11. DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
    2025/05/23 by Peng, Ziqiao, Fan, Yanbo, Wu, Haoyu +4 · 10 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
    2024/08/18 by Xukun Zhou, Zhou, Xukun, Fengxin Li +13 · 4 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Graphics (cs.GR) #Human Motion and Animation #Sound (cs.SD) #electronic engineering #information engineering
  13. Dust to Tower: Coarse-to-Fine Photo-Realistic Scene Reconstruction from Sparse Uncalibrated Images
    2024/12/27 by Xudong Cai, Cai, Xudong, Yongcai Wang +18 · 1 voice · 4 citations
    Computer Science · Earth and Planetary Sciences · #3D Surveying and Cultural Heritage #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #cs.CV
  14. Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation
    2025/05/17 by Shuo Wang, Wang, Shuo, Yongcai Wang +17 · 8 citations
    Computer Science · #Multimodal Machine Learning Applications #Semantic Web and Ontologies #Speech and dialogue systems
  15. VGG-Tex: A Vivid Geometry-Guided Facial Texture Estimation Model for High Fidelity Monocular 3D Face Reconstruction
    2024/09/15 by Haoyu Wu, Wu, Haoyu, Ziqiao Peng +11 · 3 citations
    Computer Science · Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Facial Nerve Paralysis Treatment and Research
  16. Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs
    2024/04/05 by Junhao Chen, Chen, Junhao, Xiang Li +9 · 2 citations
    Engineering · Computer Science · #Human Motion and Animation #Speech and dialogue systems #Robotics and Automated Systems
  17. MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
    2025/08/04 by Shuo Wang, Wang, Shuo, Yongcai Wang +19 · 7 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotic Path Planning Algorithms #Robotics (cs.RO) #Robotics and Sensor-Based Localization
  18. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
    2025/06/17 by Peng, Ziqiao, Hu, Wentao, Ma, Junyuan +7 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  19. ACR-Pose: Adversarial Canonical Representation Reconstruction Network for Category Level 6D Object Pose Estimation
    2021/11/20 by Fan, Zhaoxin, Song, Zhengbo, Xu, Jian +4 · 1 citation
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  20. Reconstruction-Aware Prior Distillation for Semi-supervised Point Cloud Completion
    2022/04/20 by Zhaoxin Fan, Fan, Zhaoxin, Yulin He +9 · 1 citation
    Computer Science · Engineering · #Human Pose and Action Recognition #3D Shape Modeling and Analysis #Optical measurement and interference techniques
  21. Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
    2025/08/27 by Fan, Yiguo, Ding, Pengxiang, Bai, Shuanghao +10 · 10 citations
    #FOS: Computer and information sciences #Robotics (cs.RO)
  22. MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System
    2025/03/12 by Jihao Zhao, Zhiyuan Ji, Zhao, Jihao +13 · 4 citations
    Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Robotics and Automated Systems
  23. D-IF: Uncertainty-aware Human Digitization via Implicit Distribution Field
    2023/08/17 by Xueting Yang, Yang, Xueting, Yihao Luo +9 · 1 citation
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Pose and Action Recognition
  24. Unicorn: Text-Only Data Synthesis for Vision Language Model Training
    2025/03/28 by Yu, Xiaomin, Ding, Pengxiang, Zhang, Wenjie +7 · 2 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
  25. BeatDance: A Beat-Based Model-Agnostic Contrastive Learning Framework for Music-Dance Retrieval
    2023/10/16 by Yang, Kaixing, Zhou, Xukun, Tang, Xulong +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Sound (cs.SD) #electronic engineering #information engineering
  26. DenseMP: Unsupervised Dense Pre-training for Few-shot Medical Image Segmentation
    2023/07/13 by Fan, Zhaoxin, Pan, Puquan, Zhang, Zeren +4 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  27. DH-RAG: A Dynamic Historical Context-Powered Retrieval-Augmented Generation Method for Multi-Turn Dialogue
    2025/02/19 by Feiyuan Zhang, Zhang, Feiyuan, Zhu, Dezhi +12 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning #Machine Learning (cs.LG) #Speech and dialogue systems #Topic Modeling
  28. MLPHand: Real Time Multi-View 3D Hand Mesh Reconstruction via MLP Modeling
    2024/06/23 by Jian Yang, Yang, Jian, Jiakun Li +11 · 1 citation
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Hand Gesture Recognition Systems #Human Pose and Action Recognition
  29. GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer
    2024/08/03 by Yihong Lin, Zhaoxin Fan, Lin, Yihong +14 · 1 citation
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Motion and Animation #Human Pose and Action Recognition
  30. HF-VTON: High-Fidelity Virtual Try-On via Consistent Geometric and Semantic Alignment
    2025/05/26 by Ming Meng, Meng, Ming, Qi Dong +13 · 2 citations
    Engineering · #Manufacturing Process and Optimization
  31. CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
    2024/12/26 by Kaixing Yang, Yang, Kaixing, Xulong Tang +11 · 2 citations
    Computer Science · Engineering · #Music Technology and Sound Studies #Human Motion and Animation #Music and Audio Processing
  32. Ultraman: Single Image 3D Human Reconstruction with Ultra Speed and Detail
    2024/03/18 by Mingjin Chen, Junhao Chen, Chen, Mingjin +11 · 1 citation
    Computer Science · Physics and Astronomy · #Advanced Optical Sensing Technologies #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Optical measurement and interference techniques #electronic engineering #information engineering
  33. Enhancing Weakly Supervised 3D Medical Image Segmentation through Probabilistic-aware Learning
    2024/03/05 by Jiang, Runmin, Fan, Zhaoxin, Wu, Junhao +5 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #electronic engineering #information engineering
  34. Moderating the Generalization of Score-based Generative Model
    2024/12/10 by Jiang, Wan, Wang, He, Zhang, Xin +4 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  35. Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
    2024/12/12 by Jia, Bofang, Ding, Pengxiang, Cui, Can +5 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Robotics (cs.RO)
  36. Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
    2025/08/11 by Cai, Xudong, Wang, Shuo, Wang, Peng +7 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  37. DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
    2025/06/01 by Ming Meng, Ziyi Yang, Meng, Ming +9 · 2 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering