FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
2025/09/05 by Poggi, Matteo, Tosi, Fabio · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2509.05297
Abstract
We present FlowSeek, a novel framework for optical flow requiring minimal hardware resources for training. FlowSeek marries the latest advances on the design space of optical flow networks with cutting-edge single-image depth foundation models and classical low-dimensional motion parametrization, implementing a compact, yet accurate architecture. FlowSeek is trained on a single consumer-grade GPU, a hardware budget about 8x lower compared to most recent methods, and still achieves superior cross-dataset generalization on Sintel Final and KITTI, with a relative improvement of 10 and 15% over the previous state-of-the-art SEA-RAFT, as well as on Spring and LayeredFlow datasets.
Citations
- MVSAnywhere: Zero-Shot Multi-View Stereo
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- FoundationStereo: Zero-Shot Stereo Matching
- MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
- All-in-One: Transferring Vision Foundation Models into Stereo Matching
- Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
- LayeredFlow: A Real-World Benchmark for Non-Lambertian Multi-Layer Optical Flow
- Depth Anything V2
- SEA-RAFT: Simple, Efficient, Accurate RAFT for Optical Flow
- Playing to Vision Foundation Model's Strengths in Stereo Matching
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
- DUSt3R: Geometric 3D Vision Made Easy
- FoundationPose-TensorRT
- CCMR: High Resolution Optical Flow Estimation via Coarse-to-Fine Context-Guided Motion Reasoning
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- Win-Win: Training High-Resolution Vision Transformers from Two Windows
- Multi-Object Discovery by Low-Dimensional Object Motion
- Tracking Everything Everywhere All at Once
- The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth Estimation
- DistractFlow: Improving Optical Flow Estimation via Realistic Distractions and Pseudo-Labeling
- VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation
- Rethinking Optical Flow from Geometric Matching Consistent Perspective
- Spring: A High-Resolution High-Detail Dataset and Benchmark for Scene Flow, Optical Flow and Stereo
- FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow Estimation
- Unifying Flow, Stereo and Depth Estimation
- RealFlow: EM-based Realistic Optical Flow Dataset Generation from Videos
- SKFlow: Learning Optical Flow with Super Kernels
- DIP: Deep Inverse Patchmatch for High-Resolution Optical Flow
- CRAFT: Cross-Attentional Flow Transformer for Robust Optical Flow
- FlowFormer: A Transformer Architecture for Optical Flow
- Global Matching with Overlapping Attention for Optical Flow Estimation
- Disentangling Architecture and Training for Optical Flow
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Dimensions of Motion: Monocular Prediction through Flow Subspaces
- GMFlow: Learning Optical Flow via Global Matching
- Emerging Properties in Self-Supervised Vision Transformers
- Emerging Properties in Self-Supervised Vision Transformers
- AutoFlow: Learning a Better Training Set for Optical Flow
- Learning optical flow from still images
- Learning to Estimate Hidden Motions with Global Motion Aggregation
- Motion Basis Learning for Unsupervised Deep Homography Estimation with Subspace Projection
- Vision Transformers for Dense Prediction
- Learning Transferable Visual Models From Natural Language Supervision
- EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion Segmentation
- Displacement-Invariant Matching Cost Learning for Accurate Optical Flow Estimation
- LiteFlowNet3: Resolving Correspondence Ambiguity for More Accurate Optical Flow Estimation
- What Matters in Unsupervised Optical Flow
- Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo Matching
- Learning by Analogy: Reliable Supervision from Transformations for Unsupervised Optical Flow Estimation
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- Quadratic video interpolation
- Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer
- Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer
- A Lightweight Optical Flow CNN - Revisiting Data Fidelity and Regularization
- Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation
- YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
- MegaDepth: Learning Single-View Depth Prediction from Internet Photos
- Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
- Playing for Benchmarks
- PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume
- FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
- Optical Flow Estimation using a Spatial Pyramid Network
- Full Flow: Optical Flow Estimation By Global Optimization over Regular Grids
- Deep Residual Learning for Image Recognition
- FlowNet: Learning Optical Flow with Convolutional Networks
- EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow
- Vision meets robotics: The KITTI dataset
Cited by
Related