- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
2026/07/30 by Hanzhang Zhou, Panrong Tong, Xu Zhang +13 · 1 voice
Computer Science · #cs.AI #cs.CV
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
2026/07/30 by Qiushi Sun, Kanzhi Cheng, Yian Wang +20 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.CV
- Three-Photon Bayesian Imaging of Ortho-Positronium
2026/07/30 by L. Raczynski, W. Krzemien, A. Coussat +5 · 1 voice
Physics and Astronomy · Computer Science · #physics.med-ph #cs.CV #physics.comp-ph
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
2026/07/29 by Alexi Gladstone, Heng Ji, Yilun Du · 1 voice
Computer Science · #cs.LG #cs.AI #cs.CL #cs.CV
- Hearsay: Vision-Language Medical Diagnoses Without an Image
2026/07/29 by Siddharth Vohra · 1 voice
Computer Science · #cs.CV #cs.AI #cs.CL #cs.CY
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
2026/07/29 by Hengyi Xie, Chenfei Yao, Xianjin Wu +7 · 1 voice
Computer Science · #cs.CV #cs.RO
- Visual prompt engineering for video models
2026/07/28 by Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7 · 1 voice
#cs.CV #cs.AI
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
2026/07/28 by Zhaoyan Chen, Zhongxiu Cong, Zhuanfeng Jin +7 · 1 citation
#cs.CV
- 3D-Aware VLMs with Implicit and Explicit Geometries
2026/07/23 by Wenhao Li, Xueying Jiang, Quanhao Qian +4 · 1 voice
#cs.CV #cs.AI #cs.LG
- What Happens to Accuracy When Photo Lineups Contain Non-Mated Rank-One Images From Large Galleries?
2026/07/23 by Genesis Argueta, Kevin W. Bowyer, Michael King +1 · 1 voice
#cs.CV
- Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
2026/07/22 by Moshiur Rahman Tonmoy, Dunren Che, Haitham Y. Adarbah +1 · 1 citation
#cs.CV #cs.AI
- A Unified Tokenization Framework for Pain Recognition using Heterogeneous 3D Modalities
2026/07/22 by Stefanos Gkikas, Christian Arzate Cruz, Valentina Becchetti +3 · 3 citations
#cs.CV
- ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment
2026/07/22 by Stefanos Gkikas, Yu Fang, Christian Arzate Cruz +2 · 3 citations
#cs.CV
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
2026/07/21 by Xinjie Zhang, Peng Zhang, Shicheng Zheng +21 · 1 voice · 1 citation
#cs.CV #cs.AI #cs.LG #cs.MM #eess.IV
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
2026/07/21 by Fan Jiang, Zhaoxu Sun, Mengchao Wang +38 · 1 voice
#cs.CV #cs.AI #cs.LG
- Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
2026/07/21 by Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch +2 · 1 voice
#cs.CV #cs.AI #cs.GR
- Three-Body Scattering for Generative Modeling
2026/07/20 by Peng Sun, Zhenglin Cheng, Deyuan Liu +3 · 1 voice · 1 citation
#cs.LG #cs.CV
- BrainNext: A General-Purpose Self-Supervised Foundation Model for Brain MRI Analysis
2026/07/20 by Moona Mazher, Abdul Qayyum, Steven A. Niederer +1 · 1 voice
#cs.CV #cs.AI
- The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
2026/07/20 by Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann +4 · 1 voice
#cs.CV #cs.LG
- Generative Transmission: Rethinking Computation, Bandwidth, and Memory in Communication
2026/07/20 by Xiangyu Chen, Jixiang Luo, Yuankai Fan +3 · 1 citation
#cs.CV
- Semantic Context Matters: Analysis of Color Names Across Domains
2026/07/19 by Adilet Yerkin, Elnara Kadyrgali, Malika Ziyada +5 · 1 voice
#cs.CV #cs.HC #cs.MM
- TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
2026/07/19 by Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12 · 1 voice
#cs.CV
- HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
2026/07/19 by Lingwei Dang, Juntong Li, Zonghan Li +5 · 1 citation
#cs.CV
- Points as Tori: Fast Pointwise Signed Distance for Point Clouds
2026/07/18 by Nicole Feng, Ioannis Gkioulekas, Keenan Crane · 1 voice
#cs.GR #cs.CV
- Can Multimodal Large Language Models Understand OCT?
2026/07/18 by Baochen Fu, Wenzhi Deng, Baihao Jin +5 · 2 citations
#cs.CV #cs.CL
- Value-Monotonicity Matters: A Concordance Loss for Deep Survival Prediction
2026/07/18 by Meixu Chen, Kai Wang, Jing Wang · 1 citation
#cs.LG #cs.CV #stat.ML
- An Exam for Active Observers
2026/07/17 by Jiarui Zhang, Muzi Tao, Shangshang Wang +3 · 3 voices
#cs.CV #cs.AI #cs.CL #cs.LG
- CSS-BA: Gate-Guided Column Space Search for Bundle Adjustment
2026/07/17 by Ayano Kaneda, Takafumi Taketomi, Shugo Yamaguchi +1 · 1 voice
#cs.CV
- STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex
2026/07/17 by Ethan B. Trepka, Ruobing Xia, Shude Zhu +6 · 1 voice
#q-bio.NC #cs.CV
- MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
2026/07/16 by Yushi Huang, Xiangxin Zhou, Jun Zhang +2 · 1 voice
#cs.CV #cs.LG
more