Micrograph segmentations for DDEVD
2023/04/05 by Alexander M. Kirillov, Eric Mintun, Kirillov, Alexander +21 · 1171 citations
Computer Science · Medicine · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Visual Attention and Saliency Detection
paper · pdf · doi:10.48550/arxiv.2304.02643
openalex publication_date 2023/04/05 · openalex created_date 2023/04/07 · openalex updated_date 2026/07/28
Abstract
A dataset of segmented images of refractory metal micrographs segmented using SAM and cleaned to not contain micrograph boundaries. No cleaning of the SAM segmentations has been done at this stage.
Cited by
- InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting
- MedSAM-based lung masking for multi-label chest X-ray classification
- Towards Integrating Uncertainty for Domain-Agnostic Segmentation
- NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
- CountGD++: Generalized Prompting for Open-World Counting
- Contour Information Aware 2D Gaussian Splatting for Image Representation
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
- A Rapid GeoSAM-Based Workflow for Multi-Temporal Glacier Delineation: Case Study from Svalbard
- 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation
- Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- FinPercep-RM: A Fine-grained Reward Model and Co-evolutionary Curriculum for RL-based Real-world Super-Resolution
- Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
- Scalpel-SAM: A Semi-Supervised Paradigm for Adapting SAM to Infrared Small Object Detection
- SAM 3D for 3D Object Reconstruction from Remote Sensing Images
- A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation
- RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes
- Effect of User-Prompted Priors on Semi-Automated Cancer Lesion Segmentation in Whole-Body Computed Tomography
- DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection
- Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer
- EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations
- Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
- Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation
- Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
- StAR: Segment Anything Reasoner
- Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
- Patch-Discontinuity Mining for Generalized Deepfake Detection
- LVLM-Aided Alignment of Task-Specific Vision Models
- Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
- SDUM: A Scalable Deep Unrolled Model for Universal MRI Reconstruction
- UniStateDLO: Unified Generative State Estimation and Tracking of Deformable Linear Objects Under Occlusion for Constrained Manipulation
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- Fast SAM2 with Text-Driven Token Pruning
- Surgical Scene Segmentation using a Spike-Driven Video Transformer with Real-Time Potential
- ORCA: Object Recognition and Comprehension for Archiving Marine Species
- TGC-Net: A Structure-Aware and Semantically-Aligned Framework for Text-Guided Medical Image Segmentation
- Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
- UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images
- Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations
- Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation
- LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- A Flexible Field-Based Policy Learning Framework for Diverse Robotic Systems and Sensors
- Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
- HyGE-Occ: Hybrid View-Transformation with 3D Gaussian and Edge Priors for 3D Panoptic Occupancy Prediction
- VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference
- Multi-Part Object Representations via Graph Structures and Co-Part Discovery
- Enhancing 3D Semantic Scene Completion with a Refinement Module
- Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- SLIM: Semantic-based Low-bitrate Image compression for Machines by leveraging diffusion
- Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
- MatE: Material Extraction from Single-Image via Geometric Prior
- Name That Part: 3D Part Segmentation and Naming
- Dexterous World Models
- Keypoint Counting Classifiers: Turning Vision Transformers into Self-Explainable Models Without Training
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Dialectics for Artificial Intelligence
- InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
- The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
- SceneDiff: A Benchmark and Method for Multiview Object Change Detection
- Memory-Enhanced SAM3 for Occlusion-Robust Surgical Instrument Segmentation
- Pixel Seal: Adversarial-only training for invisible image and video watermarking
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
- AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
- CountZES: Counting via Zero-Shot Exemplar Selection
- PixelArena: A benchmark for Pixel-Precision Visual Intelligence
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
- Empirical aesthetics of bridges
- In Pursuit of Pixel Supervision for Visual Pre-training
- Multi-View Foundation Models
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- On the Effectiveness of Textual Prompting with Lightweight Fine-Tuning for SAM3 Remote Sensing Segmentation
- OMCL: Open-vocabulary Monte Carlo Localization
- Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
- Model Agnostic Preference Optimization for Medical Image Segmentation
- Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
- Unified Semantic Transformer for 3D Scene Understanding
- Particulate: Feed-Forward 3D Object Articulation
- Incentivizing Tool-augmented Thinking with Images for Medical Image Analysis
- Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
- MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Consistent Instance Field for Dynamic Scene Understanding
- GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners
- SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- RecTok: Reconstruction Distillation along Rectified Flow
- Learning to Generate Cross-Task Unexploitable Examples
- StarryGazer: Leveraging Monocular Depth Estimation Models for Domain-Agnostic Single Depth Image Completion
- Harmonizing Generalization and Specialization: Uncertainty-Informed Collaborative Learning for Semi-supervised Medical Image Segmentation
- UniVCD: A New Method for Unsupervised Change Detection in the Open-Vocabulary Era
- Light Field Based 6DoF Tracking of Previously Unobserved Objects
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- Generative Spatiotemporal Data Augmentation
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Referring Change Detection in Remote Sensing Imagery
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
- Architecting Large Action Models for Human-in-the-Loop Intelligent Robots
- SSL-MedSAM2: A Semi-supervised Medical Image Segmentation Framework Powered by Few-shot Learning of SAM2
- FreqDINO: Frequency-Guided Adaptation for Generalized Boundary-Aware Ultrasound Image Segmentation
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- Weak-to-Strong Generalization Enables Fully Automated De Novo Training of Multi-head Mask-RCNN Model for Segmenting Densely Overlapping Cell Nuclei in Multiplex Whole-slice Brain Images
- AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- Prior-Enhanced Gaussian Splatting for Dynamic Scene Reconstruction from Casual Video
- VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces
- Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
- Grounding Everything in Tokens for Multimodal Large Language Models
- Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- SAM-Body4D: Training-Free 4D Human Body Mesh Recovery from Videos
- Building Reasonable Inference for Vision-Language Models in Blind Image Quality Assessment
- Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation
- UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents
- Label-free Motion-Conditioned Diffusion Model for Cardiac Ultrasound Synthesis
- From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
- GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
- Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
- Contrast transfer functions help quantify neural network out-of-distribution generalization in HRTEM
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images
- Scale-invariant and View-relational Representation Learning for Full Surround Monocular Depth
- OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention
- Team-Aware Football Player Tracking with SAM: An Appearance-Based Approach to Occlusion Recovery
- LapFM: A Laparoscopic Segmentation Foundation Model via Hierarchical Concept Evolving Pre-training
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels
- Inferring Compositional 4D Scenes without Ever Seeing One
- Closed-Loop Robotic Manipulation of Transparent Substrates for Self-Driving Laboratories using Deep Learning Micro-Error Correction
- PAVAS: Physics-Aware Video-to-Audio Synthesis
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- Accuracy Does Not Guarantee Human-Likeness in Monocular Depth Estimators
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- PVeRA: Probabilistic Vector-Based Random Matrix Adaptation
- MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
- ESPADA: Execution Speedup via Semantics Aware Demonstration Data Downsampling for Imitation Learning
- Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
- Generalized Referring Expression Segmentation on Aerial Photos
- ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
- Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
- D3-Predictor: Noise-Free Deterministic Diffusion for Dense Prediction
- Online Segment Any 3D Thing as Instance Tracking
- A graph generation pipeline for critical infrastructures based on heuristics, images and depth data
- More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery
- MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
- Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
- Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
- Dynamic Visual SLAM using a General 3D Prior
- Hide-and-Seek Attribution: Weakly Supervised Segmentation of Vertebral Metastases in CT
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Pseudo-Label Refinement for Robust Wheat Head Segmentation via Two-Stage Hybrid Training
- NexusFlow: Unifying Disparate Tasks under Partial Supervision via Invertible Flow Networks
- Multi-Modal Zero-Shot Prediction of Color Trajectories in Food Drying
- Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection
- A Hyperspectral Imaging Guided Robotic Grasping System
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
- Self-Supervised Learning for Transparent Object Depth Completion Using Depth from Non-Transparent Objects
- Stable Single-Pixel Contrastive Learning for Semantic and Geometric Tasks
- The SAM2-to-SAM3 Gap in the Segment Anything Model Family: Why Prompt-Based Expertise Fails in Concept-Driven Image Segmentation
- ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
- Coordinated Humanoid Manipulation with Choice Policies
- Malicious Image Analysis via Vision-Language Segmentation Fusion: Detection, Element, and Location in One-shot
- SAM3-I: Segment Anything with Instructions
- Boundary-Aware Test-Time Adaptation for Zero-Shot Medical Image Segmentation
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- C3G: Learning Compact 3D Representations with 2K Gaussians
- Traffic Image Restoration under Adverse Weather via Frequency-Aware Mamba
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- HalluGen: Synthesizing Realistic and Controllable Hallucinations for Evaluating Image Restoration
- PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
- Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
- Flux4D: Flow-based Unsupervised 4D Reconstruction
- MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
- Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
- Object Counting with GPT-4o and GPT-5: A Comparative Study
- Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
- Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
- Prompt-based Consistent Video Colorization
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
- Learning Visual Affordance from Audio
- KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
- SAM3-UNet: Simplified Adaptation of Segment Anything Model 3
- Evaluating SAM2 for Video Semantic Segmentation
- Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledge
- ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- OpenBox: Annotate Any Bounding Boxes in 3D
- TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Tracking
- FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting
- VSRD++: Autolabeling for 3D Object Detection via Instance-Aware Volumetric Silhouette Rendering
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
- OmniFD: A Unified Model for Versatile Face Forgery Detection
- S2AM3D: Scale-controllable Part Segmentation of 3D Point Clouds
- TAP-CT: 3D Task-Agnostic Pretraining of Computed Tomography Foundation Models
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- Silhouette-based Gait Foundation Model
- Describe Anything Anywhere At Any Moment
- CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration
- EZ-SP: Fast and Lightweight Superpoint-Based 3D Segmentation
- PPBoost: Progressive Prompt Boosting for Text-Driven Medical Image Segmentation
- PAT3D: Physics-Augmented Text-to-3D Scene Generation
- GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
- Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data
- DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
- Language-guided 3D scene synthesis for fine-grained functionality understanding
- Fast Multi-view Consistent 3D Editing with Video Priors
- RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
- GSPN-2: Efficient Parallel Sequence Modeling
- Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
- Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Creating Blank Canvas Against AI-enabled Image Forgery
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- SemOD: Semantic Enabled Object Detection Network under Various Weather Conditions
- ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing Images
- SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- TR-Gaussians: High-fidelity Real-time Rendering of Planar Transmission and Reflection with 3D Gaussian Splatting
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation
- MarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environments
- PG-ControlNet: A Physics-Guided ControlNet for Generative Spatially Varying Image Deblurring
- RLM: A Vision-Language Model Approach for Radar Scene Understanding
- Open Vocabulary Compositional Explanations for Neuron Alignment
- GaINeR: Geometry-Aware Implicit Network Representation
- V2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- DINO-Tok: Adapting DINO for Visual Tokenizers
- Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning
- PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
- Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
- Fluid Intelligence: A Forward Look on AI Foundation Models in Computational Fluid Dynamics
- TaCo: Capturing Spatio-Temporal Semantic Consistency in Remote Sensing Change Detection
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
- DOGE: Differentiable Bezier Graph Optimization for Road Network Extraction
- SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
- Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorization
- In-Context Compositional Learning via Sparse Coding Transformer
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- Vision-Language Enhanced Foundation Model for Semi-supervised Medical Image Segmentation
- RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models
- SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation
- HABIT: Human Action Benchmark for Interactive Traffic in CARLA
- DEAP-3DSAM: Decoder Enhanced and Auto Prompt SAM for 3D Medical Image Segmentation
- Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation
- MedSAM3: Delving into Segment Anything with Medical Concepts
- Medal S: Spatio-Textual Prompt Model for Medical Segmentation
- Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
- DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
- Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
- Autonomous Surface Selection For Manipulator-Based UV Disinfection In Hospitals Using Foundation Models
- ObjectAlign: Neuro-Symbolic Object Consistency Verification and Correction
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Lightweight Transformer Framework for Weakly Supervised Semantic Segmentation
- CoD: A Diffusion Foundation Model for Image Compression
- Ref-SAM3D: Bridging SAM3D with Text for Reference 3D Reconstruction
- PhysGS: Bayesian-Inferred Gaussian Splatting for Physical Property Estimation
- ReCoGS: Real-time ReColoring for Gaussian Splatting scenes
- SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation
- Vision Token Masking Alone Cannot Prevent PHI Leakage in Medical Document OCR: A Systematic Evaluation
- SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors
- AVERY: Adaptive VLM Split Computing through Embodied Self-Awareness for Efficient Disaster Response Systems
- SCALER: SAM-Enhanced Collaborative Learning for Label-Deficient Concealed Object Segmentation
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- Not Quite Anything: Overcoming SAMs Limitations for 3D Medical Imaging
- CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation
- Attention Guided Alignment in Efficient Vision-Language Models
- CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- Illustrator's Depth: Monocular Layer Index Prediction for Image Decomposition
- Planning with Sketch-Guided Verification for Physics-Aware Video Generation
- Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
- Continual Alignment for SAM: Rethinking Foundation Models for Medical Image Segmentation in Continual Learning
- Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
- Trust-Aware Multimodal Data Fusion for Yield Estimation: A Case Study of the 2020 Beirut Explosion
- NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses
- PartUV: Part-Based UV Unwrapping of 3D Meshes
- You Only Forward Once: An Efficient Compositional Judging Paradigm
- Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation
- Decoupling Complexity from Scale in Latent Diffusion Model
- NaTex: Seamless Texture Generation as Latent Color Diffusion
- Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
- Rad-GS: Radar-Vision Integration for 3D Gaussian Splatting SLAM in Outdoor Environments
- Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM
- Box6D : Zero-shot Category-level 6D Pose Estimation of Warehouse Boxes
- UniUltra: Interactive Parameter-Efficient SAM2 for Universal Ultrasound Segmentation
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- Taming Generative Synthetic Data for X-ray Prohibited Item Detection
- UniSER: A Foundation Model for Unified Soft Effects Removal
- Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- Unbiased Semantic Decoding with Vision Foundation Models for Few-shot Segmentation
- EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
- Birth of a Painting: Differentiable Brushstroke Reconstruction
- Segmentation-Aware Latent Diffusion for Satellite Image Super-Resolution: Enabling Smallholder Farm Boundary Delineation
- Visionary Co-Driver: Enhancing Driver Perception of Potential Risks with LLM and HUD
- Segment Anything Across Shots: A Method and Benchmark
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
- Part-X-MLLM: Part-aware 3D Multimodal Large Language Model
- Lite ENSAM: a lightweight cancer segmentation model for 3D Computed Tomography
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
- Semantic BIM enrichment for firefighting assets: Fire-ART dataset and panoramic image-based 3D reconstruction
- R2Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection
- C3Net: Context-Contrast Network for Camouflaged Object Detection
- Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
- EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
- ClutterNav: Gradient-Guided Search for Efficient 3D Clutter Removal with Learned Costmaps
- CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
- LithoSeg: A Coarse-to-Fine Framework for High-Precision Lithography Segmentation
- Calibrated Decomposition of Aleatoric and Epistemic Uncertainty in Deep Features for Inference-Time Adaptation
- Changes in Real Time: Online Scene Change Detection with Multi-View Fusion
- Fast Reasoning Segmentation for Images and Videos
- Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
- MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
- Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
- Draft and Refine with Visual Experts
- Binary Verification for Zero-Shot Vision
- Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots
- MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics
- LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
- LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping
- High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection
- Exploring the Underwater World Segmentation without Extra Training
- ZeroSim: Zero-Shot Analog Circuit Evaluation with Unified Transformer Embeddings
- CAVER: Curious Audiovisual Exploring Robot
- Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
- Glioma C6: A Novel Dataset for Training and Benchmarking Cell Segmentation
- Semi-supervised Shelter Mapping for WASH Accessibility Assessment in Rohingya Refugee Camps
- ProcGen3D: Learning Neural Procedural Graph Representations for Image-to-3D Reconstruction
- CenterMamba-SAM: Center-Prioritized Scanning and Temporal Prototypes for Brain Lesion Segmentation
- From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
- DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
- NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
- DIAL-GS: Dynamic Instance Aware Reconstruction for Label-free Street Scenes with 4D Gaussian Splatting
- FlowFeat: Pixel-Dense Embedding of Motion Profiles
- Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
- A Two-Stage System for Layout-Controlled Image Generation using Large Language Models and Diffusion Models
- Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
- Training-Free Adaptive Quantization for Variable Rate Image Coding for Machines
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- Towards Better Ultrasound Video Segmentation Foundation Model: An Empirical study on SAM2 Finetuning from Data Perspective
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
- Walk the Lines 2: Contour Tracking for Detailed Segmentation
- SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation
- Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
- Real-World Adverse Weather Image Restoration via Dual-Level Reinforcement Learning with High-Quality Cold Start
- Tracking and Understanding Object Transformations
- Landslide Hazard Mapping with Geospatial Foundation Models: Geographical Generalizability, Data Scarcity, and Band Adaptability
- Covariance Descriptors Meet General Vision Encoders: Riemannian Deep Learning for Medical Image Classification
- Learning-based Cooperative Robotic Paper Wrapping: A Unified Control Policy with Residual Force Control
- Accelerating Physical Property Reasoning for Augmented Visual Cognition
- MIQ-SAM3D: From Single-Point Prompt to Multi-Instance Segmentation via Competitive Query Refinement
- Zero-Shot Multi-Animal Tracking in the Wild
- Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems
- OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control
- From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- From Instance Segmentation to 3D Growth Trajectory Reconstruction in Planktonic Foraminifera
- UniChange: Unifying Change Detection with Multimodal Large Language Model
- Anatomically Constrained Transformers for Echocardiogram Analysis
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- RefVTON: person-to-person Try on with Additional Unpaired Visual Reference
- Deep Generative Models for Enhanced Vitreous OCT Imaging
- Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- MIFO: Learning and Synthesizing Multi-Instance from One Image
- Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
- LGCA: Enhancing Semantic Representation via Progressive Expansion
- MapSAM2: Adapting SAM2 for Automatic Segmentation of Historical Map Images and Time Series
- FMint-SDE: A Multimodal Foundation Model for Accelerating Numerical Simulation of SDEs via Error Correction
- A Step Toward World Models: A Survey on Robotic Manipulation
- BlurGuard: A Simple Approach for Robustifying Image Protection Against AI-Powered Editing
- DM-QPMNET: Dual-modality fusion network for cell segmentation in quantitative phase microscopy
- AD-SAM: Fine-Tuning the Segment Anything Vision Foundation Model for Autonomous Driving Perception
- SAMRI: Segment Anything Model for MRI
- SpinalSAM-R1: A Vision-Language Multimodal Interactive System for Spine CT Segmentation
- Towards Realistic Earth-Observation Constellation Scheduling: Benchmark and Methodology
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- Self-localization on a 3D map by fusing global and local features from a monocular camera
- FullPart: Generating each 3D Part at Full Resolution
- FlexICL: A Flexible Visual In-context Learning Framework for Elbow and Wrist Ultrasound Segmentation
- AutoSurvey2: Empowering Researchers with Next Level Automated Literature Surveys
- Fine-tuning Segment Anything for Real-Time Tumor Tracking in Cine-MRI
- SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- STITCH 2.0: Extending Augmented Suturing with EKF Needle Estimation and Thread Management
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- ROGR: Relightable 3D Objects using Generative Relighting
- Instruction-based image editing: a survey on data, models, evaluation, and applications
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps
- Domain Generalization for Semantic Segmentation: A Survey
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- Advancing marine microplastic monitoring through deep learning-based image segmentation
- Unlimited OCR Works
- Scene-Centric Unsupervised Video Panoptic Segmentation
- SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
- Prototype-Grounded Concept Models for Verifiable Concept Alignment
- TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
- PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
- Instance-Aware Pseudo-Labeling and Class-Focused Contrastive Learning for Weakly Supervised Domain Adaptive Segmentation of Electron Microscopy
- Promptable Fire Segmentation: Unleashing SAM2's Potential for Real-Time Mobile Deployment with Strategic Bounding Box Guidance
- Quantification of plant trait data from herbarium scans in the DiSSCo Research Infrastructure
- Cataract-LMM: Large-Scale, Multi-Source, Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
- Aligning What You Separate: Denoised Patch Mixing for Source-Free Domain Adaptation in Medical Image Segmentation
- EA3D: Online Open-World 3D Object Extraction from Streaming Videos
- AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians
- UP2D: Uncertainty-aware Progressive Pseudo-label Denoising for Source-Free Domain Adaptive Medical Image Segmentation
- SAGE: Structure-Aware Generative Video Transitions between Diverse Clips
- Advancing site-specific disease and pest management in precision agriculture: From reasoning-driven foundation models to adaptive, feedback-based learning
- Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Vanish into Thin Air: Cross-prompt Universal Adversarial Attacks for SAM2
- LagMemo: Language 3D Gaussian Splatting Memory for Multi-modal Open-vocabulary Multi-goal Visual Navigation
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- PlanarGS: High-Fidelity Indoor 3D Gaussian Splatting Guided by Vision-Language Planar Priors
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
- Proactive Scene Decomposition and Reconstruction
- Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- USF-MAE: Ultrasound Self-Supervised Foundation Model with Masked Autoencoding
- Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges
- Gen-LangSplat: Generalized Language Gaussian Splatting with Pre-Trained Feature Compression
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- Optimal Spatial Anomaly Detection
- Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- Why Registration Quality Matters: Enhancing sCT Synthesis with IMPACT-Based Registration
- TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- Chain of Execution Supervision Promotes General Reasoning in Large Language Models
- BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies
- Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI
- Video-As-Prompt: Unified Semantic Control for Video Generation
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
- Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization
- COS3D: Collaborative Open-Vocabulary 3D Segmentation
- Monocular Visual 8D Pose Estimation for Articulated Bicycles and Cyclists
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
- Curvilinear Structure-preserving Unpaired Cross-domain Medical Image Translation
- Latent Space Factorization in LoRA
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- [De|Re]constructing VLMs' Reasoning in Counting
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Advances in 4D Representation: Geometry, Motion, and Interaction
- Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation
- Seg the HAB: Language-Guided Geospatial Algae Bloom Reasoning and Segmentation
- RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- EVER: Edge-Assisted Auto-Verification for Mobile MR-Aided Operation
- EMA-SAM: Exponential Moving-average for SAM-based PTMC Segmentation
- RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- SAM 2++: Tracking Anything at Any Granularity
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- DeepSeek-OCR: Contexts Optical Compression
- Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
- Botany-Bot: Digital Twin Monitoring of Occluded and Underleaf Plant Structures with Gaussian Splats
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Morphology-Aware KOA Classification: Integrating Graph Priors with Vision Models
- Automatic Classification of Circulating Blood Cell Clusters based on Multi-channel Flow Cytometry Imaging
- Towards 3D Objectness Learning in an Open World
- Intelligent Communication Mixture-of-Experts Boosted-Medical Image Segmentation Foundation Model
- Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction
- From Pixels to People: Satellite-Based Mapping and Quantification of Riverbank Erosion and Lost Villages in Bangladesh
- GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- Optimizing DINOv2 with Registers for Face Anti-Spoofing
- Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
- Personalized Image Filter: Mastering Your Photographic Style
- Geospatial Machine Learning Libraries
- Experience-Driven Exploration for Efficient API-Free AI Agents
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- ChangingGrounding: 3D Visual Grounding in Changing Scenes
- DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- Talking Points: Describing and Localizing Pixels
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation
- Salient Concept-Aware Generative Data Augmentation
- MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- AnyUp: Universal Feature Upsampling
- GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents
- Assessing the Potential for Catastrophic Failure in Dynamic Post-Training Quantization
- Unlocking Zero-Shot Plant Segmentation with Pl@ntNet Intelligence
- BEEP3D: Box-Supervised End-to-End Pseudo-Mask Generation for 3D Instance Segmentation
- G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- Inferring Dynamic Physical Properties from Video Foundation Models
- SNAP: Towards Segmenting Anything in Any Point Cloud
- Robust Ego-Exo Correspondence with Long-Term Memory
- When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
- Generalisation of automatic tumour segmentation in histopathological whole-slide images across multiple cancer types
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
- XGrasp: Gripper-Aware Grasp Detection with Multi-Gripper Data Generation
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- High-Fidelity Speech Enhancement via Discrete Audio Tokens
- Real2USD: Scene Representations in Universal Scene Description Language
- Fast Vision in the Dark: A Case for Single-Photon Imaging in Planetary Navigation
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Unified Open-World Segmentation with Multi-Modal Prompts
- FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
- MSM-Seg: A Modality-and-Slice Memory Framework with Category-Agnostic Prompting for Multi-Modal Brain Tumor Segmentation
- SAM2LoRA: Composite Loss-Guided, Parameter-Efficient Finetuning of SAM2 for Retinal Fundus Segmentation
- Sketch Animation: State-of-the-art Report
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer
- Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
- Tracking the Spatiotemporal Evolution of Landslide Scars Using a Vision Foundation Model: A Novel and Universal Framework
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- VG-Mapping: Variation-Aware 3D Gaussians for Online Semi-static Scene Mapping
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- MemPromptTSS: Persistent Prompt Memory for Iterative Multi-Granularity Time Series State Segmentation
- J-RAS: Mutual Adaptation for Medical Image Segmentation via Contrastive Retrieval-Augmented Joint Optimization
- AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
- SSeg: Active Sparse Point-Label Augmentation for Semantic Segmentation
- An uncertainty-aware framework for data-efficient multi-view animal pose estimation
- Holistic Order Prediction in Natural Scenes
- Cell Instance Segmentation: The Devil Is in the Boundaries
- Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenes
- TARO: Toward Semantically Rich Open-World Object Detection
- SAM2-3dMed: Empowering SAM2 for 3D Medical Image Segmentation
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- A methodology for clinically driven interactive segmentation evaluation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
- Vision Language Models: A Survey of 26K Papers
- LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
- Towards Precise Channel Knowledge Map: Exploiting Environmental Information from 2D Visuals to 3D Point Clouds
- Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- Geometry-aware Policy Imitation
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- DADO: A Depth-Attention framework for Object Discovery
- HTMformer: Hybrid Time and Multivariate Transformer for Time Series Forecasting
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- MATCH: Multi-faceted Adaptive Topo-Consistency for Semi-Supervised Histopathology Segmentation
- Extreme Amodal Face Detection
- Avi: Action from Volumetric Inference
- We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
- SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- Efficient Universal Models for Medical Image Segmentation via Weakly Supervised In-Context Learning
- StereoSync: Spatially-Aware Stereo Audio Generation from Video
- ALISE: Annotation-Free LiDAR Instance Segmentation for Autonomous Driving
- TFM Dataset: A Novel Multi-task Dataset and Integrated Pipeline for Automated Tear Film Break-Up Segmentation
- HoloScene: Simulation-Ready Interactive 3D Worlds from a Single Video
- Human3R: Everyone Everywhere All at Once
- Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis
- SegMASt3R: Geometry Grounded Segment Matching
- MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- SPEGNet: Synergistic Perception-Guided Network for Camouflaged Object Detection
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Spatially Grounded Concept-Based Image Classification
- Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
- Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- The Overlooked Value of Test-time Reference Sets in Visual Place Recognition
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Dynamic Prompt Generation for Interactive 3D Medical Image Segmentation Training
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- SAMSOD: Rethinking SAM Optimization for RGB-T Salient Object Detection
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis
- OTR: Synthesizing Overlay Text Dataset for Text Removal
- Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
- KeySG: Hierarchical Keyframe-Based 3D Scene Graphs
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- ProtoMask: Segmentation-Guided Prototype Learning
- Multi-Domain Brain Vessel Segmentation Through Feature Disentanglement
- Robust Context-Aware Object Recognition
- Multi-level Dynamic Style Transfer for NeRFs
- Advances in Medical Image Segmentation: A Comprehensive Survey with a Focus on Lumbar Spine Applications
- Domain-Specialized Interactive Segmentation Framework for Meningioma Radiotherapy Planning
- Photorealistic Inpainting for Perturbation-based Explanations in Ecological Monitoring
- Solar PV Installation Potential Assessment on Building Facades Based on Vision and Language Foundation Models
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
- Evaluating New AI Cell Foundation Models on Challenging Kidney Pathology Cases Unaddressed by Previous Foundation Models
- Drones that Think on their Feet: Sudden Landing Decisions with Embodied AI
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Video Object Segmentation-Aware Audio Generation
- AiDE-Q: Synthetic Labeled Datasets Can Enhance Learning Models for Quantum Property Estimation
- EasyOcc: 3D Pseudo-Label Supervision for Fully Self-Supervised Semantic Occupancy Prediction Models
- SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies
- Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
- A Multi-purpose Tracking Framework for Salmon Welfare Monitoring in Challenging Environments
- Kairos: Towards Adaptive and Generalizable Time Series Foundation Models
- Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
- DescribeEarth: Describe Anything for Remote Sensing Images
- PinPoint3D: Fine-Grained 3D Part Segmentation from a Few Clicks
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- LayerD: Decomposing Raster Graphic Designs into Layers
- Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
- Evaluation of Polarimetric Fusion for Semantic Segmentation in Aquatic Environments
- CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
- Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency
- RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
- Mask Clustering-based Annotation Engine for Large-Scale Submeter Land Cover Mapping
- TP-MVCC: Tri-plane Multi-view Fusion Model for Silkie Chicken Counting
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning
- BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
- K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
- Adaptive Canonicalization with Application to Invariant Anisotropic Geometric Networks
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- Personalized Vision via Visual In-Context Learning
- CrashSplat: 2D to 3D Vehicle Damage Segmentation in Gaussian Splatting
- Revisit the Imbalance Optimization in Multi-task Learning: An Experimental Analysis
- AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
- A Weather Foundation Model for the Power Grid
- Color-Pair Guided Robust Zero-Shot 6D Pose Estimation and Tracking of Cluttered Objects on Edge Devices
- Efficient Domain-Adaptive Multi-Task Dense Prediction with Vision Foundation Models
- BioVessel-Net and RetinaMix: Unsupervised Retinal Vessel Segmentation from OCTA Images
- ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
- From Fields to Splats: A Cross-Domain Survey of Real-Time Neural Scene Representations
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
- Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
- Confidence-Calibrating Regularization for Robust Brain MRI Segmentation Under Domain Shift
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- RefAM: Attention Magnets for Zero-Shot Referral Segmentation
- LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- Polysemous Language Gaussian Splatting via Matching-based Mask Lifting
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
- KG-SAM: Injecting Anatomical Knowledge into Segment Anything Models via Conditional Random Fields
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
- Unsupervised Defect Detection for Surgical Instruments
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- SLAM-Free Visual Navigation with Hierarchical Vision-Language Perception and Coarse-to-Fine Semantic Topological Planning
- Large Pre-Trained Models for Bimanual Manipulation in 3D
- LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
- Optical Ocean Recipes: Creating Realistic Datasets to Facilitate Underwater Vision Research
- LLM Trainer: Automated Robotic Data Generating via Demonstration Augmentation using LLMs
- Embodied AI: From LLMs to World Models
- AJAHR: Amputated Joint Aware 3D Human Mesh Recovery
- Where Did I Leave My Glasses? Open-Vocabulary Semantic Exploration in Real-World Semi-Static Environments
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
- Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
- VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
- ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
- Cross-Embodiment Transfer via Behavior-Aligned Representations
- Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
- Articulated Object Reconstruction from Rest-State Observation
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
- LivePyxel: accelerating image annotations with a Python-integrated webcam live streaming
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- Citizen Centered Climate Intelligence: Operationalizing Open Tree Data for Urban Cooling and Eco-Routing in Indian Cities
- Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation
- Deep-learning deconvolution and segmentation of fluorescent membranes for high-precision bacterial cell-size profiling
- EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- A Contrastive Learning-Guided Confident Meta-learning for Zero Shot Anomaly Detection
- RoSe: Robust Self-supervised Stereo Matching under Adverse Weather Conditions
- Spectral Signature Mapping from RGB Imagery for Terrain-Aware Navigation
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection
- MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving
- Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model
- Prompt-DAS: Annotation-Efficient Prompt Learning for Domain Adaptive Semantic Segmentation of Electron Microscopy Images
- The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model
- OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
- iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
- Learning Geometry-Aware Nonprehensile Pushing and Pulling with Dexterous Hands
- A Single Image Is All You Need: Zero-Shot Anomaly Localization Without Training Data
- Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- Towards Seeing Bones at Radio Frequency
- VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
- Few-Shot Pattern Detection via Template Matching and Regression
- From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge
- Depth Edge Alignment Loss: DEALing with Depth in Weakly Supervised Semantic Segmentation
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
- StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models
- Neural-MMGS: Multi-modal Neural Gaussian Splats for Large-Scale Scene Reconstruction
- Computational Scaffolding of Composition, Value, and Color for Disciplined Drawing
- Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- SAM-DCE: Addressing Token Uniformity and Semantic Over-Smoothing in Medical Segmentation
- MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation
- Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- Right-Side-Out: Learning Zero-Shot Sim-to-Real Garment Reversal
- ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
- Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
- Overview of PlantCLEF 2024: multi-species plant identification in vegetation plot images
- pFedSAM: Personalized Federated Learning of Segment Anything Model for Medical Image Segmentation
- Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
- TASAM: Terrain-and-Aware Segment Anything Model for Temporal-Scale Remote Sensing Segmentation
- FloorSAM: SAM-Guided Floorplan Reconstruction with Semantic-Geometric Fusion
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Sparse Multiview Open-Vocabulary 3D Detection
- Introducing Resizable Region Packing Problem in Image Generation, with a Heuristic Solution
- Compose by Focus: Scene Graph-based Atomic Skills
- RangeSAM: On the Potential of Visual Foundation Models for Range-View represented LiDAR segmentation
- Region-Aware Deformable Convolutions
- Geometric Image Synchronization with Deep Watermarking
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- Transplant-Ready? Evaluating AI Lung Segmentation Models in Candidates with Severe Lung Disease
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
- AutoEdit: Automatic Hyperparameter Tuning for Image Editing
- Variational Shape Inference for Grasp Diffusion on SE(3)
- Fracture interactive geodesic active contours for bone segmentation
- Trade-offs in Cross-Domain Generalization of Foundation Model Fine-Tuned for Biometric Applications
- Pseudo-Label Enhanced Cascaded Framework: 2nd Technical Report for LSVOS 2025 VOS Track
- FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction
- Frequency-Aware Ensemble Learning for BraTS 2025 Pediatric Brain Tumor Segmentation
- E-BayesSAM: Efficient Bayesian Adaptation of SAM with Self-Optimizing KAN-Based Interpretation for Uncertainty-Aware Ultrasonic Segmentation
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
- Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
- MSGFusion: Multimodal Scene Graph-Guided Infrared and Visible Image Fusion
- When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective
- Superpixel Anything: A general object-based framework for accurate yet regular superpixel segmentation
- CLIFF: Continual Learning for Incremental Flake Features in 2D Material Identification
- A biological vision inspired framework for machine perception of abutting grating illusory contours
- Advancing Real-World Parking Slot Detection with Large-Scale Dataset and Semi-Supervised Baseline
- Beyond Averages: Open-Vocabulary 3D Scene Understanding with Gaussian Splatting and Bag of Embeddings
- Advancing Weakly-Supervised Change Detection in Satellite Images via Adversarial Class Prompting
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- Road Obstacle Video Segmentation
- LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
- RailSafeNet: Visual Scene Understanding for Tram Safety
- FS-SAM2: Adapting Segment Anything Model 2 for Few-Shot Semantic Segmentation via Low-Rank Adaptation
- A Computer Vision Pipeline for Individual-Level Behavior Analysis: Benchmarking on the Edinburgh Pig Dataset
- SERES: Semantic-aware neural reconstruction from sparse views
- Segmentation-Driven Initialization for Sparse-view 3D Gaussian Splatting
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Synthetic vs. Real Training Data for Visual Navigation
- Seg2Track-SAM2: SAM2-based Multi-object Tracking and Segmentation
- A Controllable 3D Deepfake Generation Framework with Gaussian Splatting
- Multi-animal tracking in Transition: Comparative Insights into Established and Emerging Methods
- IMD: A 6-DoF Pose Estimation Benchmark for Industrial Metallic Objects
- MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic Segmentation
- Joint-octamamba:an octa joint segmentation network based on feature enhanced mamba
- WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild
- M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments
- Leveraging Geometric Priors for Unaligned Scene Change Detection
- OpenUrban3D: Annotation-Free Open-Vocabulary Semantic Segmentation of Large-Scale Urban Point Clouds
- Multimodal SAM-adapter for Semantic Segmentation
- SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation
- SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition
- A Comparison and Evaluation of Fine-tuned Convolutional Neural Networks to Large Language Models for Image Classification and Segmentation of Brain Tumors on MRI
- WebSight: A Vision-First Architecture for Robust Web Agents
- Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States
- GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
- Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
- Segment Anything for Cell Tracking
- Towards Understanding Visual Grounding in Visual Language Models
- ObjectReact: Learning Object-Relative Control for Visual Navigation
- PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
- Region-Wise Correspondence Prediction between Manga Line Art Images
- Mixture of Semantics Transmission for Generative AI-Enabled Semantic Communication Systems
- Image Recognition with Vision and Language Embeddings of VLMs
- Modular, On-Site Solutions with Lightweight Anomaly Detection for Sustainable Nutrient Management in Agriculture
- OCELOT 2023: Cell Detection from Cell-Tissue Interaction Challenge
- Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- Live(r) Die: Predicting Survival in Colorectal Liver Metastasis
- SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular Video
- CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging
- Implicit Shape-Prior for Few-Shot Assisted 3D Segmentation
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- X-Part: high fidelity and structure coherent shape decomposition
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- Visual Representation Alignment for Multimodal Large Language Models
- MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
- Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- Privacy Preserving Semantic Communications Using Vision Language Models: A Segmentation and Generation Approach
- XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning
- CellEcoNet: Decoding the Cellular Language of Pathology with Deep Learning for Invasive Lung Adenocarcinoma Recurrence Prediction
- P3-SAM: Native 3D Part Segmentation
- VIM-GS: Visual-Inertial Monocular Gaussian Splatting via Object-level Guidance in Large Scenes
- Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- Co-Seg: Mutual Prompt-Guided Collaborative Learning for Tissue and Nuclei Segmentation
- Event Spectroscopy: Event-based Multispectral and Depth Sensing using Structured Light
- Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models
- Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control
- A Probabilistic Segment Anything Model for Ambiguity-Aware Medical Image Segmentation
- Visibility-Aware Language Aggregation for Open-Vocabulary Segmentation in 3D Gaussian Splatting
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Enhancing Self-Driving Segmentation in Adverse Weather Conditions: A Dual Uncertainty-Aware Training Approach to SAM Optimization
- Towards Open World Detection: A Survey
- Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping
- SSGaussian: Semantic-Aware and Structure-Preserving 3D Style Transfer
- TensoIS: A Step Towards Feed-Forward Tensorial Inverse Subsurface Scattering for Perlin Distributed Heterogeneous Media
- A Synthetic-to-Real Dehazing Method based on Domain Unification
- Weakly-Supervised Learning of Dense Functional Correspondences
- Reactive In-Air Clothing Manipulation with Confidence-Aware Dense Correspondence and Visuotactile Affordance
- A Multidimensional AI-powered Framework for Analyzing Tourist Perception in Historic Urban Quarters: A Case Study in Shanghai
- 20 years of microfluidic technology for advancing plant sciences
- SLENet: A Guidance-Enhanced Network for Underwater Camouflaged Object Detection
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- Sample-efficient Integration of New Modalities into Large Language Models
- AutoDetect: Designing an Autoencoder-based Detection Method for Poisoning Attacks on Object Detection Applications in the Military Domain
- Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- PointAD+: Learning Hierarchical Representations for Zero-shot 3D Anomaly Detection
- Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
- IRSAMap:Towards Large-Scale, High-Resolution Land Cover Map Vectorization
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- An Investigation of Visual Foundation Models Robustness
- Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
- Self-Validated Learning for Particle Separation: A Correctness-Based Self-Training Framework Without Human Labels
- First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
- Fail2Progress: Learning from Real-World Robot Failures with Stein Variational Inference
- Im2Haircut: Single-view Strand-based Hair Reconstruction for Human Avatars
- Image Quality Enhancement and Detection of Small and Dense Objects in Industrial Recycling Processes
- Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation
- SegAssess: Panoramic quality mapping for robust and transferable unsupervised segmentation assessment
- MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
- Cross-Domain Few-Shot Segmentation via Ordinary Differential Equations over Time Intervals
- NeuralMeshing: Complete Object Mesh Extraction from Casual Captures
- No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- Can General-Purpose Omnimodels Compete with Specialists? A Case Study in Medical Image Segmentation
- DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
- Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
- Promptable Longitudinal Lesion Segmentation in Whole-Body CT
- A Modality-agnostic Multi-task Foundation Model for Human Brain Imaging
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank
- CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Representation Learning with Adaptive Superpixel Coding
- Towards Interactive Lesion Segmentation in Whole-Body PET/CT with Promptable Models
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation
- CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
- SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
- PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
- NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation
- OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
- Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
- LabelGS: Label-Aware 3D Gaussian Splatting for 3D Scene Segmentation
- Interact-Custom: Customized Human Object Interaction Image Generation
- SDiFL: Stable Diffusion-Driven Framework for Image Forgery Localization
- Autoregressive Universal Video Segmentation Model
- MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction
- The point is the mask: scaling coral reef segmentation with weak supervision
- OpenTie: Open-vocabulary Sequential Rebar Tying System
- Feature-Space Planes Searcher: A Universal Domain Adaptation Framework for Interpretability and Computational Efficiency
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- ROSE: Remove Objects with Side Effects in Videos
- eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images
- ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation
- GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
- Deep learning in plant phenotyping: the first ten years
- GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects
- Understanding Data Influence with Differential Approximation
- Locality-aware Concept Bottleneck Model
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images
- GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
- Adapting Biological Reflexes for Dynamic Reorientation in Space Manipulator Systems
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- ViT-FIQA: Assessing Face Image Quality using Vision Transformers
- PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis
- Diversity-enhanced Collaborative Mamba for Semi-supervised Medical Image Segmentation
- subCellSAM: Zero-Shot (Sub-)Cellular Segmentation for Hit Validation in Drug Discovery
- DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction
- OmniTry: Virtual Try-On Anything without Masks
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- AIM 2025 Rip Current Segmentation (RipSeg) Challenge Report
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop
- Splat Feature Solver
- S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- PEdger++: Practical Edge Detection via Assembling Cross Information
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
- StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
- Visuomotor Grasping with World Models for Surgical Robots
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
- Ovis2.5 Technical Report
- SlicerMorph photogrammetry: an open-source photogrammetry workflow for reconstructing 3D models
- Utilizing Vision-Language Models as Action Models for Intent Recognition and Assistance
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
- MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation
- Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Adapting SAM via Cross-Entropy Masking for Class Imbalance in Remote Sensing Change Detection
- Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Towards Efficient Prompt-based Continual Learning in Distributed Medical AI
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single Image
- COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound Datasets
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- Automated Segmentation of Coronal Brain Tissue Slabs for 3D Neuropathology
- Multi-Sequence Parotid Gland Lesion Segmentation via Expert Text-Guided Segment Anything Model
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
- ViPE: Video Pose Engine for 3D Geometric Perception
- HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View Synthesis
- Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
- MADPromptS: Unlocking Zero-Shot Morphing Attack Detection with Multiple Prompt Aggregation
- GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing Environments
- Exploring Palette based Color Guidance in Diffusion Models
- Think as Cardiac Sonographers: Marrying SAM with Left Ventricular Indicators Measurements According to Clinical Guidelines
- CObL: Toward Zero-Shot Ordinal Layering without User Prompting
- ReferSplat: Referring Segmentation in 3D Gaussian Splatting
- SAGOnline: Segment Any Gaussians Online
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild
- NeeCo: Image Synthesis of Novel Instrument States Based on Dynamic and Deformable 3D Gaussian Reconstruction
- Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
- A Spin Glass Characterization of Neural Networks
- 3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction
- ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack
- Membership Inference Attacks with False Discovery Rate Control
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
- Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
- Text-guided Visual Prompt DINO for Generic Segmentation
- SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures
- ETA: Energy-based Test-time Adaptation for Depth Completion
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Improving Diagnostic Accuracy for Oral Cancer with inpainting Synthesis Lesions Generated Using Diffusion Models
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- Real-Time 3D Vision-Language Embedding Mapping
- SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- User-Intent-Driven Semantic Communication via Adaptive Deep Understanding
- Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction
- Improving Masked Style Transfer using Blended Partial Convolution
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- SMOL-MapSeg: Show Me One Label as prompt
- F2PASeg: Feature Fusion for Pituitary Anatomy Segmentation in Endoscopic Surgery
- SGDFuse: SAM-Guided Diffusion for High-Fidelity Infrared and Visible Image Fusion
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Decoupling Continual Semantic Segmentation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Modeling Rapid Contextual Learning in the Visual Cortex with Fast-Weight Deep Autoencoder Networks
- CF3: Compact and Fast 3D Feature Fields
- HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
- Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
- Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- Open Scene Graphs for Open-World Object-Goal Navigation
- A Scalable Pretraining Framework for Link Prediction with Efficient Adaptation
- MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Benchmarking Foundation Models for Mitotic Figure Classification
- Composed Object Retrieval: Object-level Retrieval via Composed Expressions
- Deep Learning-based Scalable Image-to-3D Facade Parser for Generating Thermal 3D Building Models
- Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
- Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
- Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification
- RPCANet++: Deep Interpretable Robust PCA for Sparse Object Segmentation
- Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
- OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
- SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
- SoilNet: A Multimodal Multitask Model for Hierarchical Classification of Soil Horizons
- MAUP: Training-free Multi-center Adaptive Uncertainty-aware Prompting for Cross-domain Few-shot Medical Image Segmentation
- Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
- ParticleSAM: Small Particle Segmentation for Material Quality Monitoring in Recycling Processes
- Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
- Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping
- ADSeeker: A Knowledge-Infused Framework for Anomaly Detection and Reasoning
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
- Rethinking Transparent Object Grasping: Depth Completion with Monocular Depth Estimation and Instance Mask
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- Data-driven RF Tomography via Cross-modal Sensing and Continual Learning
- DreamPainter: Image Background Inpainting for E-commerce Scenarios
- ScrewSplat: An End-to-End Method for Articulated Object Recognition
- Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
- SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- TopoImages: Incorporating Local Topology Encoding into Deep Learning Models for Medical Image Classification
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- Register Anything: Estimating "Corresponding Prompts" for Segment Anything Model
- Learning to Perform Low-Contact Autonomous Nasotracheal Intubation by Recurrent Action-Confidence Chunking with Transformer
- AG2aussian: Anchor-Graph Structured Gaussian Splatting for Instance-Level 3D Scene Understanding and Editing
- M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
- Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
- Physically-based Lighting Generation for Robotic Manipulation
- ReMu: Reconstructing Multi-layer 3D Clothed Human from Image Layers
- Integrating Disparity Confidence Estimation into Relative Depth Prior-Guided Unsupervised Stereo Matching
- OCSplats: Observation Completeness Quantification and Label Noise Separation in 3DGS
- Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
- OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding
- MASIV: Toward Material-Agnostic System Identification from Videos
- Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- Video Color Grading via Look-Up Table Generation
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- LesiOnTime -- Joint Temporal and Clinical Modeling for Small Breast Lesion Segmentation in Longitudinal DCE-MRI
- SDMatte: Grafting Diffusion Models for Interactive Matting
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
- Multimodal Referring Segmentation: A Survey
- PointGauss: Point Cloud-Guided Multi-Object Segmentation for Gaussian Splatting
- Omni-Scan: Creating Visually-Accurate Digital Twin Object Models Using a Bimanual Robot with Handover and Gaussian Splat Merging
- Object-Centric Cropping for Visual Few-Shot Classification
- Topology Optimization in Medical Image Segmentation with Fast Euler Characteristic
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
- Mamba-based Efficient Spatio-Frequency Motion Perception for Video Camouflaged Object Detection
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- ST-SAM: SAM-Driven Self-Training Framework for Semi-Supervised Camouflaged Object Detection
- Training-free Geometric Image Editing on Diffusion Models
- PixNerd: Pixel Neural Field Diffusion
Related