Segment Anything
2023/10/01 by Alexander M. Kirillov, Eric Mintun, Nikhila Ravi +9 · 819 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Visual Attention and Saliency Detection
paper · doi:10.1109/iccv51070.2023.00371
openalex publication_date 2023/10/01 · openalex created_date 2024/01/16 · openalex updated_date 2026/07/31
Abstract
We introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest segmentation dataset to date (by far), with over 1 billion masks on 11M licensed and privacy respecting images. The model is designed and trained to be promptable, so it can transfer zero-shot to new image distributions and tasks. We evaluate its capabilities on numerous tasks and find that its zero-shot performance is impressive – often competitive with or even superior to prior fully supervised results. We are releasing the Segment Anything Model (SAM) and corresponding dataset (SA-1B) of 1B masks and 11M images at segment-anything.com to foster research into foundation models for computer vision. We recommend reading the full paper at: arxiv.org/abs/2304.02643.
Citations
Cited by
- ROGR: Relightable 3D Objects using Generative Relighting
- Quantification of plant trait data from herbarium scans in the DiSSCo Research Infrastructure
- Impact on Cost and Expert Time of Data-Efficient Deep Learning for Medical Image Segmentation
- FedBone: Towards Large-Scale Federated Multi-Task Learning
- FROM-GLC Plus 3.0: Multimodal Land Change Mapping with SAM and Dense Surface Observations
- Clinic-aligned Dual Distillation of Video and Image Foundation Models for Automated Breast Cancer US Diagnosis
- Domain Generalization for Semantic Segmentation: A Survey
- GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis
- Advancing marine microplastic monitoring through deep learning-based image segmentation
- Unlimited OCR Works
- High‐throughput recognition of small invertebrates via computer vision
- Scene-Centric Unsupervised Video Panoptic Segmentation
- SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
- A New Approach for Calculating Texture Coefficients of Different Rocks With Image Segmentation and Image Processing Techniques
- SHIC-XE: Viewpoint-Invariant Explainability via Dense 2D-3D Correspondences: an Application to Equine Pain Recognition
- Towards Stable Source-Free Domain Adaptive Semantic Segmentation
- Prototype-Grounded Concept Models for Verifiable Concept Alignment
- ReScale4DL: Balancing Pixel and Contextual Information for Enhanced Bioimage Segmentation
- DINOSim: Zero-Shot Object Detection and Semantic Segmentation on Microscopy Images
- LSKNet: A Foundation Lightweight Backbone for Remote Sensing
- TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
- PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
- Instance-Aware Pseudo-Labeling and Class-Focused Contrastive Learning for Weakly Supervised Domain Adaptive Segmentation of Electron Microscopy
- Promptable Fire Segmentation: Unleashing SAM2's Potential for Real-Time Mobile Deployment with Strategic Bounding Box Guidance
- Cataract-LMM: Large-Scale, Multi-Source, Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
- Aligning What You Separate: Denoised Patch Mixing for Source-Free Domain Adaptation in Medical Image Segmentation
- EA3D: Online Open-World 3D Object Extraction from Streaming Videos
- AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians
- UP2D: Uncertainty-aware Progressive Pseudo-label Denoising for Source-Free Domain Adaptive Medical Image Segmentation
- SAGE: Structure-Aware Generative Video Transitions between Diverse Clips
- Advancing site-specific disease and pest management in precision agriculture: From reasoning-driven foundation models to adaptive, feedback-based learning
- Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- Vanish into Thin Air: Cross-prompt Universal Adversarial Attacks for SAM2
- LagMemo: Language 3D Gaussian Splatting Memory for Multi-modal Open-vocabulary Multi-goal Visual Navigation
- Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
- PlanarGS: High-Fidelity Indoor 3D Gaussian Splatting Guided by Vision-Language Planar Priors
- A simple method for rapid reconstruction of 3D animal trajectory from monocular video
- Detecting Bees in Cherry Flowers Using Timelapse Images and Foundational Models
- SARATR-X: Toward Building a Foundation Model for SAR Target Recognition
- Context-CAM: Context-Level Weight-Based CAM With Sequential Denoising to Generate High-Quality Class Activation Maps
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
- Proactive Scene Decomposition and Reconstruction
- Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- USF-MAE: Ultrasound Self-Supervised Foundation Model with Masked Autoencoding
- Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges
- Gen-LangSplat: Generalized Language Gaussian Splatting with Pre-Trained Feature Compression
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- Optimal Spatial Anomaly Detection
- Diffusion-Driven Two-Stage Active Learning for Low-Budget Semantic Segmentation
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- Why Registration Quality Matters: Enhancing sCT Synthesis with IMPACT-Based Registration
- TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
- Controllable-LPMoE: Adapting to Challenging Object Segmentation via Dynamic Local Priors from Mixture-of-Experts
- Chain of Execution Supervision Promotes General Reasoning in Large Language Models
- BioDet: Boosting Industrial Object Detection with Image Preprocessing Strategies
- Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI
- Video-As-Prompt: Unified Semantic Control for Video Generation
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
- Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization
- COS3D: Collaborative Open-Vocabulary 3D Segmentation
- Monocular Visual 8D Pose Estimation for Articulated Bicycles and Cyclists
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
- Curvilinear Structure-preserving Unpaired Cross-domain Medical Image Translation
- Latent Space Factorization in LoRA
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- [De|Re]constructing VLMs' Reasoning in Counting
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Advances in 4D Representation: Geometry, Motion, and Interaction
- Face-MakeUpV2: Facial Consistency Learning for Controllable Text-to-Image Generation
- Seg the HAB: Language-Guided Geospatial Algae Bloom Reasoning and Segmentation
- RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- EVER: Edge-Assisted Auto-Verification for Mobile MR-Aided Operation
- EMA-SAM: Exponential Moving-average for SAM-based PTMC Segmentation
- RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- SAM 2++: Tracking Anything at Any Granularity
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- DeepSeek-OCR: Contexts Optical Compression
- Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
- Botany-Bot: Digital Twin Monitoring of Occluded and Underleaf Plant Structures with Gaussian Splats
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Morphology-Aware KOA Classification: Integrating Graph Priors with Vision Models
- Automatic Classification of Circulating Blood Cell Clusters based on Multi-channel Flow Cytometry Imaging
- Towards 3D Objectness Learning in an Open World
- Intelligent Communication Mixture-of-Experts Boosted-Medical Image Segmentation Foundation Model
- The predictive ability of GBVS feature channels on infants’ fixations of natural scenes
- Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction
- From Pixels to People: Satellite-Based Mapping and Quantification of Riverbank Erosion and Lost Villages in Bangladesh
- GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- Universal forged image detection and localization via self-supervised data generation and large-scale model adaptation
- Optimizing DINOv2 with Registers for Face Anti-Spoofing
- Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
- Personalized Image Filter: Mastering Your Photographic Style
- Geospatial Machine Learning Libraries
- Experience-Driven Exploration for Efficient API-Free AI Agents
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- ChangingGrounding: 3D Visual Grounding in Changing Scenes
- DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- Talking Points: Describing and Localizing Pixels
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation
- Salient Concept-Aware Generative Data Augmentation
- MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation
- Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- AnyUp: Universal Feature Upsampling
- GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents
- Assessing the Potential for Catastrophic Failure in Dynamic Post-Training Quantization
- Unlocking Zero-Shot Plant Segmentation with Pl@ntNet Intelligence
- BEEP3D: Box-Supervised End-to-End Pseudo-Mask Generation for 3D Instance Segmentation
- G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation
- CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- Inferring Dynamic Physical Properties from Video Foundation Models
- SNAP: Towards Segmenting Anything in Any Point Cloud
- Robust Ego-Exo Correspondence with Long-Term Memory
- When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
- Generalisation of automatic tumour segmentation in histopathological whole-slide images across multiple cancer types
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
- XGrasp: Gripper-Aware Grasp Detection with Multi-Gripper Data Generation
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- High-Fidelity Speech Enhancement via Discrete Audio Tokens
- Real2USD: Scene Representations in Universal Scene Description Language
- Fast Vision in the Dark: A Case for Single-Photon Imaging in Planetary Navigation
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Unified Open-World Segmentation with Multi-Modal Prompts
- FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
- MSM-Seg: A Modality-and-Slice Memory Framework with Category-Agnostic Prompting for Multi-Modal Brain Tumor Segmentation
- SAM2LoRA: Composite Loss-Guided, Parameter-Efficient Finetuning of SAM2 for Retinal Fundus Segmentation
- Sketch Animation: State-of-the-art Report
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer
- Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
- Tracking the Spatiotemporal Evolution of Landslide Scars Using a Vision Foundation Model: A Novel and Universal Framework
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- VG-Mapping: Variation-Aware 3D Gaussians for Online Semi-static Scene Mapping
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- MemPromptTSS: Persistent Prompt Memory for Iterative Multi-Granularity Time Series State Segmentation
- J-RAS: Mutual Adaptation for Medical Image Segmentation via Contrastive Retrieval-Augmented Joint Optimization
- AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
- SSeg: Active Sparse Point-Label Augmentation for Semantic Segmentation
- An uncertainty-aware framework for data-efficient multi-view animal pose estimation
- Holistic Order Prediction in Natural Scenes
- Cell Instance Segmentation: The Devil Is in the Boundaries
- Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenes
- TARO: Toward Semantically Rich Open-World Object Detection
- SAM2-3dMed: Empowering SAM2 for 3D Medical Image Segmentation
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- A methodology for clinically driven interactive segmentation evaluation
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
- Vision Language Models: A Survey of 26K Papers
- LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
- Towards Precise Channel Knowledge Map: Exploiting Environmental Information from 2D Visuals to 3D Point Clouds
- Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- Geometry-aware Policy Imitation
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- DADO: A Depth-Attention framework for Object Discovery
- HTMformer: Hybrid Time and Multivariate Transformer for Time Series Forecasting
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- MATCH: Multi-faceted Adaptive Topo-Consistency for Semi-Supervised Histopathology Segmentation
- Extreme Amodal Face Detection
- Avi: Action from Volumetric Inference
- We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
- SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- Efficient Universal Models for Medical Image Segmentation via Weakly Supervised In-Context Learning
- StereoSync: Spatially-Aware Stereo Audio Generation from Video
- ALISE: Annotation-Free LiDAR Instance Segmentation for Autonomous Driving
- TFM Dataset: A Novel Multi-task Dataset and Integrated Pipeline for Automated Tear Film Break-Up Segmentation
- HoloScene: Simulation-Ready Interactive 3D Worlds from a Single Video
- Human3R: Everyone Everywhere All at Once
- Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis
- SegMASt3R: Geometry Grounded Segment Matching
- MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- SPEGNet: Synergistic Perception-Guided Network for Camouflaged Object Detection
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Spatially Grounded Concept-Based Image Classification
- Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
- Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- The Overlooked Value of Test-time Reference Sets in Visual Place Recognition
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Dynamic Prompt Generation for Interactive 3D Medical Image Segmentation Training
- A Hybrid Co-Finetuning Approach for Visual Bug Detection in Video Games
- SAMSOD: Rethinking SAM Optimization for RGB-T Salient Object Detection
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis
- OTR: Synthesizing Overlay Text Dataset for Text Removal
- Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
- KeySG: Hierarchical Keyframe-Based 3D Scene Graphs
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- ProtoMask: Segmentation-Guided Prototype Learning
- Multi-Domain Brain Vessel Segmentation Through Feature Disentanglement
- Robust Context-Aware Object Recognition
- Multi-level Dynamic Style Transfer for NeRFs
- Advances in Medical Image Segmentation: A Comprehensive Survey with a Focus on Lumbar Spine Applications
- Domain-Specialized Interactive Segmentation Framework for Meningioma Radiotherapy Planning
- Photorealistic Inpainting for Perturbation-based Explanations in Ecological Monitoring
- Solar PV Installation Potential Assessment on Building Facades Based on Vision and Language Foundation Models
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
- Evaluating New AI Cell Foundation Models on Challenging Kidney Pathology Cases Unaddressed by Previous Foundation Models
- Drones that Think on their Feet: Sudden Landing Decisions with Embodied AI
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Video Object Segmentation-Aware Audio Generation
- AiDE-Q: Synthetic Labeled Datasets Can Enhance Learning Models for Quantum Property Estimation
- EasyOcc: 3D Pseudo-Label Supervision for Fully Self-Supervised Semantic Occupancy Prediction Models
- SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies
- Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
- A Multi-purpose Tracking Framework for Salmon Welfare Monitoring in Challenging Environments
- Kairos: Towards Adaptive and Generalizable Time Series Foundation Models
- Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
- DescribeEarth: Describe Anything for Remote Sensing Images
- PinPoint3D: Fine-Grained 3D Part Segmentation from a Few Clicks
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- LayerD: Decomposing Raster Graphic Designs into Layers
- Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
- Evaluation of Polarimetric Fusion for Semantic Segmentation in Aquatic Environments
- CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
- Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency
- RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
- Mask Clustering-based Annotation Engine for Large-Scale Submeter Land Cover Mapping
- TP-MVCC: Tri-plane Multi-view Fusion Model for Silkie Chicken Counting
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning
- BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
- PodNet: Pod real-time instance segmentation in pre-harvest soybean fields
- K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
- <i>Samplify</i> : A versatile tool for image-based segmentation and annotation of seed abortion phenotypes
- Segment anything in medical images
- Adaptive Canonicalization with Application to Invariant Anisotropic Geometric Networks
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- Personalized Vision via Visual In-Context Learning
- CrashSplat: 2D to 3D Vehicle Damage Segmentation in Gaussian Splatting
- Revisit the Imbalance Optimization in Multi-task Learning: An Experimental Analysis
- AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
- A Weather Foundation Model for the Power Grid
- Color-Pair Guided Robust Zero-Shot 6D Pose Estimation and Tracking of Cluttered Objects on Edge Devices
- Efficient Domain-Adaptive Multi-Task Dense Prediction with Vision Foundation Models
- BioVessel-Net and RetinaMix: Unsupervised Retinal Vessel Segmentation from OCTA Images
- ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
- From Fields to Splats: A Cross-Domain Survey of Real-Time Neural Scene Representations
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- Machine Learning on Blockchain (MLOB): A New Paradigm for Computational Security in Engineering
- Robotic integration for end-stations at scientific user facilities
- Foundation models in bioinformatics
- GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
- Segment Anything Model Can Not Segment Anything: Assessing AI Foundation Model’s Generalizability in Permafrost Mapping
- Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
- Confidence-Calibrating Regularization for Robust Brain MRI Segmentation Under Domain Shift
- RefAM: Attention Magnets for Zero-Shot Referral Segmentation
- LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- Polysemous Language Gaussian Splatting via Matching-based Mask Lifting
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
- Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation
- CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
- KG-SAM: Injecting Anatomical Knowledge into Segment Anything Models via Conditional Random Fields
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
- Unsupervised Defect Detection for Surgical Instruments
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- SLAM-Free Visual Navigation with Hierarchical Vision-Language Perception and Coarse-to-Fine Semantic Topological Planning
- Large Pre-Trained Models for Bimanual Manipulation in 3D
- LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
- Optical Ocean Recipes: Creating Realistic Datasets to Facilitate Underwater Vision Research
- LLM Trainer: Automated Robotic Data Generating via Demonstration Augmentation using LLMs
- Embodied AI: From LLMs to World Models
- AJAHR: Amputated Joint Aware 3D Human Mesh Recovery
- Where Did I Leave My Glasses? Open-Vocabulary Semantic Exploration in Real-World Semi-Static Environments
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
- Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
- VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
- ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
- Cross-Embodiment Transfer via Behavior-Aligned Representations
- Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- Drawing out What They Struggle to say: AI-Augmented Analysis of Projective Techniques in Qualitative Health Research
- TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
- Articulated Object Reconstruction from Rest-State Observation
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
- Deep learning in plant phenotyping: the first ten years
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- FADE: A Task-Agnostic Upsampling Operator for Encoder–Decoder Architectures
- Embodied large language models enable robots to complete complex tasks in unpredictable environments
- Domain-specific AI segmentation of IMPDH2 rod/ring structures in mouse embryonic stem cells
- Research Note: A deep learning method segments chicken keel bones from whole-body X-ray images
- Towards generalizable Federated Learning in medical imaging: A real-world case study on mammography data
- Towards cognition-augmented human-centric assembly: A visual computation perspective
- Unlocking the power of AI for phenotyping fruit morphology in Arabidopsis
- Enhancing augmented reality with machine learning for hands-on origami training
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- BioImageIT: A novel python‐based architecture for reproducible bio‐image workflows
- Citizen Centered Climate Intelligence: Operationalizing Open Tree Data for Urban Cooling and Eco-Routing in Indian Cities
- Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation
- Deep-learning deconvolution and segmentation of fluorescent membranes for high-precision bacterial cell-size profiling
- EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images
- Domain and Task-Focused Example Selection for Data-Efficient Contrastive Medical Image Segmentation
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- A Contrastive Learning-Guided Confident Meta-learning for Zero Shot Anomaly Detection
- RoSe: Robust Self-supervised Stereo Matching under Adverse Weather Conditions
- Spectral Signature Mapping from RGB Imagery for Terrain-Aware Navigation
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection
- MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving
- Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model
- Prompt-DAS: Annotation-Efficient Prompt Learning for Domain Adaptive Semantic Segmentation of Electron Microscopy Images
- The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Model
- OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
- iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
- Learning Geometry-Aware Nonprehensile Pushing and Pulling with Dexterous Hands
- A Single Image Is All You Need: Zero-Shot Anomaly Localization Without Training Data
- Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- Towards Seeing Bones at Radio Frequency
- VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
- Few-Shot Pattern Detection via Template Matching and Regression
- From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge
- Depth Edge Alignment Loss: DEALing with Depth in Weakly Supervised Semantic Segmentation
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
- StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models
- Neural-MMGS: Multi-modal Neural Gaussian Splats for Large-Scale Scene Reconstruction
- Computational Scaffolding of Composition, Value, and Color for Disciplined Drawing
- Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- SAM-DCE: Addressing Token Uniformity and Semantic Over-Smoothing in Medical Segmentation
- MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation
- Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- Right-Side-Out: Learning Zero-Shot Sim-to-Real Garment Reversal
- ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
- Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
- Overview of PlantCLEF 2024: multi-species plant identification in vegetation plot images
- pFedSAM: Personalized Federated Learning of Segment Anything Model for Medical Image Segmentation
- Towards Size-invariant Salient Object Detection: A Generic Evaluation and Optimization Approach
- TASAM: Terrain-and-Aware Segment Anything Model for Temporal-Scale Remote Sensing Segmentation
- FloorSAM: SAM-Guided Floorplan Reconstruction with Semantic-Geometric Fusion
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Sparse Multiview Open-Vocabulary 3D Detection
- Introducing Resizable Region Packing Problem in Image Generation, with a Heuristic Solution
- Compose by Focus: Scene Graph-based Atomic Skills
- RangeSAM: On the Potential of Visual Foundation Models for Range-View represented LiDAR segmentation
- Region-Aware Deformable Convolutions
- Geometric Image Synchronization with Deep Watermarking
- WorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free Guidance
- Transplant-Ready? Evaluating AI Lung Segmentation Models in Candidates with Severe Lung Disease
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
- AutoEdit: Automatic Hyperparameter Tuning for Image Editing
- Variational Shape Inference for Grasp Diffusion on SE(3)
- Fracture interactive geodesic active contours for bone segmentation
- Trade-offs in Cross-Domain Generalization of Foundation Model Fine-Tuned for Biometric Applications
- Pseudo-Label Enhanced Cascaded Framework: 2nd Technical Report for LSVOS 2025 VOS Track
- FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction
- Frequency-Aware Ensemble Learning for BraTS 2025 Pediatric Brain Tumor Segmentation
- Lightweight Progressive Multilevel Feature Collaborative Network for Remote Sensing Image Salient Object Detection
- E-BayesSAM: Efficient Bayesian Adaptation of SAM with Self-Optimizing KAN-Based Interpretation for Uncertainty-Aware Ultrasonic Segmentation
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
- Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
- Human-Centric Fine-Grained Action Quality Assessment
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
- MSGFusion: Multimodal Scene Graph-Guided Infrared and Visible Image Fusion
- When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective
- Superpixel Anything: A general object-based framework for accurate yet regular superpixel segmentation
- CLIFF: Continual Learning for Incremental Flake Features in 2D Material Identification
- A biological vision inspired framework for machine perception of abutting grating illusory contours
- Advancing Real-World Parking Slot Detection with Large-Scale Dataset and Semi-Supervised Baseline
- Beyond Averages: Open-Vocabulary 3D Scene Understanding with Gaussian Splatting and Bag of Embeddings
- Advancing Weakly-Supervised Change Detection in Satellite Images via Adversarial Class Prompting
- MMMS: Multi-Modal Multi-Surface Interactive Segmentation
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- Road Obstacle Video Segmentation
- LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
- RailSafeNet: Visual Scene Understanding for Tram Safety
- FS-SAM2: Adapting Segment Anything Model 2 for Few-Shot Semantic Segmentation via Low-Rank Adaptation
- A Computer Vision Pipeline for Individual-Level Behavior Analysis: Benchmarking on the Edinburgh Pig Dataset
- SERES: Semantic-aware neural reconstruction from sparse views
- Segmentation-Driven Initialization for Sparse-view 3D Gaussian Splatting
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Synthetic vs. Real Training Data for Visual Navigation
- Seg2Track-SAM2: SAM2-based Multi-object Tracking and Segmentation
- A Controllable 3D Deepfake Generation Framework with Gaussian Splatting
- Multi-animal tracking in Transition: Comparative Insights into Established and Emerging Methods
- IMD: A 6-DoF Pose Estimation Benchmark for Industrial Metallic Objects
- MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic Segmentation
- Joint-octamamba:an octa joint segmentation network based on feature enhanced mamba
- WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild
- M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments
- Leveraging Geometric Priors for Unaligned Scene Change Detection
- OpenUrban3D: Annotation-Free Open-Vocabulary Semantic Segmentation of Large-Scale Urban Point Clouds
- Multimodal SAM-adapter for Semantic Segmentation
- SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation
- SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition
- A Comparison and Evaluation of Fine-tuned Convolutional Neural Networks to Large Language Models for Image Classification and Segmentation of Brain Tumors on MRI
- WebSight: A Vision-First Architecture for Robust Web Agents
- Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States
- GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
- Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
- Segment Anything for Cell Tracking
- Towards Understanding Visual Grounding in Visual Language Models
- ObjectReact: Learning Object-Relative Control for Visual Navigation
- PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
- Region-Wise Correspondence Prediction between Manga Line Art Images
- Mixture of Semantics Transmission for Generative AI-Enabled Semantic Communication Systems
- Image Recognition with Vision and Language Embeddings of VLMs
- Modular, On-Site Solutions with Lightweight Anomaly Detection for Sustainable Nutrient Management in Agriculture
- OCELOT 2023: Cell Detection from Cell-Tissue Interaction Challenge
- Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- Live(r) Die: Predicting Survival in Colorectal Liver Metastasis
- SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular Video
- CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging
- Implicit Shape-Prior for Few-Shot Assisted 3D Segmentation
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- X-Part: high fidelity and structure coherent shape decomposition
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- MAE-SAM2: Mask Autoencoder-Enhanced SAM2 for Clinical Retinal Vascular Leakage Segmentation
- Visual Representation Alignment for Multimodal Large Language Models
- MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
- Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- Privacy Preserving Semantic Communications Using Vision Language Models: A Segmentation and Generation Approach
- XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning
- CellEcoNet: Decoding the Cellular Language of Pathology with Deep Learning for Invasive Lung Adenocarcinoma Recurrence Prediction
- P3-SAM: Native 3D Part Segmentation
- VIM-GS: Visual-Inertial Monocular Gaussian Splatting via Object-level Guidance in Large Scenes
- Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
- Co-Seg: Mutual Prompt-Guided Collaborative Learning for Tissue and Nuclei Segmentation
- Event Spectroscopy: Event-based Multispectral and Depth Sensing using Structured Light
- Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models
- Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control
- A Probabilistic Segment Anything Model for Ambiguity-Aware Medical Image Segmentation
- Visibility-Aware Language Aggregation for Open-Vocabulary Segmentation in 3D Gaussian Splatting
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
- Enhancing Self-Driving Segmentation in Adverse Weather Conditions: A Dual Uncertainty-Aware Training Approach to SAM Optimization
- Towards Open World Detection: A Survey
- Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional Warping
- SSGaussian: Semantic-Aware and Structure-Preserving 3D Style Transfer
- TensoIS: A Step Towards Feed-Forward Tensorial Inverse Subsurface Scattering for Perlin Distributed Heterogeneous Media
- A Synthetic-to-Real Dehazing Method based on Domain Unification
- Weakly-Supervised Learning of Dense Functional Correspondences
- Reactive In-Air Clothing Manipulation with Confidence-Aware Dense Correspondence and Visuotactile Affordance
- A Multidimensional AI-powered Framework for Analyzing Tourist Perception in Historic Urban Quarters: A Case Study in Shanghai
- SLENet: A Guidance-Enhanced Network for Underwater Camouflaged Object Detection
- DisPatch: Disarming Adversarial Patches in Object Detection with Diffusion Models
- Sample-efficient Integration of New Modalities into Large Language Models
- AutoDetect: Designing an Autoencoder-based Detection Method for Poisoning Attacks on Object Detection Applications in the Military Domain
- Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- PointAD+: Learning Hierarchical Representations for Zero-shot 3D Anomaly Detection
- Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
- IRSAMap:Towards Large-Scale, High-Resolution Land Cover Map Vectorization
- EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- An Investigation of Visual Foundation Models Robustness
- Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
- Self-Validated Learning for Particle Separation: A Correctness-Based Self-Training Framework Without Human Labels
- First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
- Fail2Progress: Learning from Real-World Robot Failures with Stein Variational Inference
- Im2Haircut: Single-view Strand-based Hair Reconstruction for Human Avatars
- Image Quality Enhancement and Detection of Small and Dense Objects in Industrial Recycling Processes
- GaussianObject: High-Quality 3D Object Reconstruction from Four Views with Gaussian Splatting
- Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation
- SegAssess: Panoramic quality mapping for robust and transferable unsupervised segmentation assessment
- MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
- Cross-Domain Few-Shot Segmentation via Ordinary Differential Equations over Time Intervals
- NeuralMeshing: Complete Object Mesh Extraction from Casual Captures
- No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- Can General-Purpose Omnimodels Compete with Specialists? A Case Study in Medical Image Segmentation
- DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
- Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
- Promptable Longitudinal Lesion Segmentation in Whole-Body CT
- A Modality-agnostic Multi-task Foundation Model for Human Brain Imaging
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank
- CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Representation Learning with Adaptive Superpixel Coding
- Towards Interactive Lesion Segmentation in Whole-Body PET/CT with Promptable Models
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
- Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation
- CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
- SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
- PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
- NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary Segmentation
- OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
- Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
- LabelGS: Label-Aware 3D Gaussian Splatting for 3D Scene Segmentation
- Interact-Custom: Customized Human Object Interaction Image Generation
- SDiFL: Stable Diffusion-Driven Framework for Image Forgery Localization
- Autoregressive Universal Video Segmentation Model
- MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction
- The point is the mask: scaling coral reef segmentation with weak supervision
- OpenTie: Open-vocabulary Sequential Rebar Tying System
- Feature-Space Planes Searcher: A Universal Domain Adaptation Framework for Interpretability and Computational Efficiency
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- ROSE: Remove Objects with Side Effects in Videos
- eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images
- ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation
- GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
- GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects
- Understanding Data Influence with Differential Approximation
- Locality-aware Concept Bottleneck Model
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Towards PerSense++: Advancing Training-Free Personalized Instance Segmentation in Dense Images
- GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
- Adapting Biological Reflexes for Dynamic Reorientation in Space Manipulator Systems
- Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
- ViT-FIQA: Assessing Face Image Quality using Vision Transformers
- PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis
- Diversity-enhanced Collaborative Mamba for Semi-supervised Medical Image Segmentation
- subCellSAM: Zero-Shot (Sub-)Cellular Segmentation for Hit Validation in Drug Discovery
- DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction
- OmniTry: Virtual Try-On Anything without Masks
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- AIM 2025 Rip Current Segmentation (RipSeg) Challenge Report
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop
- Splat Feature Solver
- S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- PEdger++: Practical Edge Detection via Assembling Cross Information
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
- StyleMM: Stylized 3D Morphable Face Model via Text-Driven Aligned Image Translation
- Visuomotor Grasping with World Models for Surgical Robots
- VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection
- Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
- Ovis2.5 Technical Report
- SlicerMorph photogrammetry: an open-source photogrammetry workflow for reconstructing 3D models
- Utilizing Vision-Language Models as Action Models for Intent Recognition and Assistance
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning
- MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation
- MSAF-UNet: A Robust Segmentation Method Integrating Traditional Image Processing and Multi-Scale Attention Fusion
- Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- A Segmentation-driven Editing Method for Bolt Defect Augmentation and Detection
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- Adapting SAM via Cross-Entropy Masking for Class Imbalance in Remote Sensing Change Detection
- Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Towards Efficient Prompt-based Continual Learning in Distributed Medical AI
- From Pixel to Mask: A Survey of Out-of-Distribution Segmentation
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- SynSpill: Improved Industrial Spill Detection With Synthetic Data
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single Image
- COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound Datasets
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- Automated Segmentation of Coronal Brain Tissue Slabs for 3D Neuropathology
- Multi-Sequence Parotid Gland Lesion Segmentation via Expert Text-Guided Segment Anything Model
- Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
- ViPE: Video Pose Engine for 3D Geometric Perception
- HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View Synthesis
- Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
- MADPromptS: Unlocking Zero-Shot Morphing Attack Detection with Multiple Prompt Aggregation
- GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing Environments
- Exploring Palette based Color Guidance in Diffusion Models
- Think as Cardiac Sonographers: Marrying SAM with Left Ventricular Indicators Measurements According to Clinical Guidelines
- CObL: Toward Zero-Shot Ordinal Layering without User Prompting
- ReferSplat: Referring Segmentation in 3D Gaussian Splatting
- SAGOnline: Segment Any Gaussians Online
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild
- NeeCo: Image Synthesis of Novel Instrument States Based on Dynamic and Deformable 3D Gaussian Reconstruction
- Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
- A Spin Glass Characterization of Neural Networks
- 3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction
- ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack
- Membership Inference Attacks with False Discovery Rate Control
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
- Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
- Text-guided Visual Prompt DINO for Generic Segmentation
- SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures
- ETA: Energy-based Test-time Adaptation for Depth Completion
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Improving Diagnostic Accuracy for Oral Cancer with inpainting Synthesis Lesions Generated Using Diffusion Models
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- Real-Time 3D Vision-Language Embedding Mapping
- SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- User-Intent-Driven Semantic Communication via Adaptive Deep Understanding
- Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction
- Improving Masked Style Transfer using Blended Partial Convolution
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- SMOL-MapSeg: Show Me One Label as prompt
- F2PASeg: Feature Fusion for Pituitary Anatomy Segmentation in Endoscopic Surgery
- SGDFuse: SAM-Guided Diffusion for High-Fidelity Infrared and Visible Image Fusion
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- Contextual Object Detection with Multimodal Large Language Models
- Decoupling Continual Semantic Segmentation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Modeling Rapid Contextual Learning in the Visual Cortex with Fast-Weight Deep Autoencoder Networks
- CF3: Compact and Fast 3D Feature Fields
- HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
- Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
- Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- Open Scene Graphs for Open-World Object-Goal Navigation
- A Scalable Pretraining Framework for Link Prediction with Efficient Adaptation
- MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Benchmarking Foundation Models for Mitotic Figure Classification
- Composed Object Retrieval: Object-level Retrieval via Composed Expressions
- Deep Learning-based Scalable Image-to-3D Facade Parser for Generating Thermal 3D Building Models
- Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
- Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
- Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification
- RPCANet++: Deep Interpretable Robust PCA for Sparse Object Segmentation
- Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation
- DOMR: Establishing Cross-View Segmentation via Dense Object Matching
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
- OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
- SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
- SoilNet: A Multimodal Multitask Model for Hierarchical Classification of Soil Horizons
- MAUP: Training-free Multi-center Adaptive Uncertainty-aware Prompting for Cross-domain Few-shot Medical Image Segmentation
- Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
- ParticleSAM: Small Particle Segmentation for Material Quality Monitoring in Recycling Processes
- Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
- Trace3D: Consistent Segmentation Lifting via Gaussian Instance Tracing
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping
- ADSeeker: A Knowledge-Infused Framework for Anomaly Detection and Reasoning
- Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation
- FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
- Rethinking Transparent Object Grasping: Depth Completion with Monocular Depth Estimation and Instance Mask
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- Data-driven RF Tomography via Cross-modal Sensing and Continual Learning
- DreamPainter: Image Background Inpainting for E-commerce Scenarios
- ScrewSplat: An End-to-End Method for Articulated Object Recognition
- Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
- SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- TopoImages: Incorporating Local Topology Encoding into Deep Learning Models for Medical Image Classification
- Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- Register Anything: Estimating "Corresponding Prompts" for Segment Anything Model
- Learning to Perform Low-Contact Autonomous Nasotracheal Intubation by Recurrent Action-Confidence Chunking with Transformer
- AG2aussian: Anchor-Graph Structured Gaussian Splatting for Instance-Level 3D Scene Understanding and Editing
- M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
- Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
- Physically-based Lighting Generation for Robotic Manipulation
- ReMu: Reconstructing Multi-layer 3D Clothed Human from Image Layers
- Integrating Disparity Confidence Estimation into Relative Depth Prior-Guided Unsupervised Stereo Matching
- OCSplats: Observation Completeness Quantification and Label Noise Separation in 3DGS
- Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
- OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding
- MASIV: Toward Material-Agnostic System Identification from Videos
- Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights
- Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
- Video Color Grading via Look-Up Table Generation
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- LesiOnTime -- Joint Temporal and Clinical Modeling for Small Breast Lesion Segmentation in Longitudinal DCE-MRI
- SDMatte: Grafting Diffusion Models for Interactive Matting
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
- Multimodal Referring Segmentation: A Survey
- PointGauss: Point Cloud-Guided Multi-Object Segmentation for Gaussian Splatting
- Omni-Scan: Creating Visually-Accurate Digital Twin Object Models Using a Bimanual Robot with Handover and Gaussian Splat Merging
- Object-Centric Cropping for Visual Few-Shot Classification
- Topology Optimization in Medical Image Segmentation with Fast Euler Characteristic
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
- Mamba-based Efficient Spatio-Frequency Motion Perception for Video Camouflaged Object Detection
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- ST-SAM: SAM-Driven Self-Training Framework for Semi-Supervised Camouflaged Object Detection
- Training-free Geometric Image Editing on Diffusion Models
- PixNerd: Pixel Neural Field Diffusion
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- SAM-PTx: Text-Guided Fine-Tuning of SAM with Parameter-Efficient, Parallel-Text Adapters
- Details Matter for Indoor Open-vocabulary 3D Instance Segmentation
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- MergeSAM: Unsupervised change detection of remote sensing images based on the Segment Anything Model
- trAIce3D: A Prompt-Driven Transformer Based U-Net for Semantic Segmentation of Microglial Cells from Large-Scale 3D Microscopy Images
- Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- Estimating 2D Camera Motion with Hybrid Motion Basis
- Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical Assistance
- Rethink Domain Generalization in Heterogeneous Sequence MRI Segmentation
- From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- MOVE: Motion-Guided Few-Shot Video Object Segmentation
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- Mitigating Spurious Correlations in Weakly Supervised Semantic Segmentation via Cross-architecture Consistency Regularization
- SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking
- Semantic Segmentation of iPS Cells: Case Study on Model Complexity in Biomedical Imaging
- BANG: Dividing 3D Assets via Generative Exploded Dynamics
- Top2Pano: Learning to Generate Indoor Panoramas from Top-Down View
- Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning
- RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation
- PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs
- Ensemble Foreground Management for Unsupervised Object Discovery
- Implicit Counterfactual Learning for Audio-Visual Segmentation
- FantasyID: A dataset for detecting digital manipulations of ID-documents
- AIComposer: Any Style and Content Image Composition via Feature Integration
- ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
- Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
- PeerSync: Accelerating Containerized Service Delivery at the Network Edge
- SAMwave: Wavelet-Driven Feature Enrichment for Effective Adaptation of Segment Anything Model
- LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
- AnimeColor: Reference-based Animation Colorization with Diffusion Transformers
- VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving
- Detecting Visual Information Manipulation Attacks in Augmented Reality: A Multimodal Semantic Reasoning Approach
- Region-based Cluster Discrimination for Visual Representation Learning
- AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation
- CLoRA: Parameter-Efficient Continual Learning with Low-Rank Adaptation
- Taking Language Embedded 3D Gaussian Splatting into the Wild
- SCALAR: Scale-wise Controllable Visual Autoregressive Learning
- Object-centric Video Question Answering with Visual Grounding and Referring
- SurgPIS: Surgical-instrument-level Instances and Part-level Semantics for Weakly-supervised Part-aware Instance Segmentation
- Is Exchangeability better than I.I.D to handle Data Distribution Shifts while Pooling Data for Data-scarce Medical image segmentation?
- Foundation Model-Driven Grasping of Unknown Objects via Center of Gravity Estimation
- ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
- HQ-SMem: Video Segmentation and Tracking Using Memory Efficient Object Embedding With Selective Update and Self-Supervised Distillation Feedback
- Flow Stochastic Segmentation Networks
- Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
- Synthetic Data Augmentation for Enhanced Chicken Carcass Instance Segmentation
- Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
- Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
- Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding
- Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
- A quadratic paradigm describes the relationship between phenotype severity and variation
- Machine-assisted quantitizing designs: augmenting humanities and social sciences with artificial intelligence
Related