BiomechGPT: Extending Motion-Language Models to Clinical Motion Understanding
2025/05/24 by Ruize Yang, Ann Kennedy, Yang, Ruize +4 · 1 voice · 1 citation
Computer Science · Medicine · #Musculoskeletal pain and rehabilitation #cs.CV
paper · pdf · doi:10.48550/arxiv.2505.18465
openalex publication_date 2025/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Advances in markerless motion capture are making high-quality biomechanical data increasingly accessible, creating a growing need for scalable downstream analytics. Building a bespoke pipeline for each analysis task is time-consuming, motivating models that can flexibly handle diverse clinical questions within a single framework. Recent work has shown that fine-tuning language models to accept tokenized motion as an additional modality enables descriptive captioning of movement, raising the question of whether these models are also capable of clinically relevant motion understanding, where diverse tasks and annotations provide a natural testbed. We investigate whether such a multimodal motion--language model can answer detailed, clinically meaningful questions about movement. We collected 71 hours of biomechanical data from 750 participants, many with movement impairments, performing tasks commonly used in clinical assessment. To further expand the training dataset, we designed a cross-format tokenizer that directly encodes motion data from heterogeneous formats into a shared latent space without paired data, allowing a second dataset to be incorporated and enabling pooling annotations across datasets. From these tokenized representations, we constructed a multimodal dataset of motion-related question--answer pairs and used it to train BiomechGPT, a multimodal biomechanics--language model. BiomechGPT achieves competitive performance across a range of clinically relevant tasks, with performance scaling with both dataset and model size. It offers a new way for clinicians and researchers to interact with biomechanical data and represents a promising direction for rehabilitation-focused movement analysis. Project page: https://intelligentsensingandrehabilitation.github.io/BiomechGPT/
Citations
- Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data
- MotionGPT3: Human Motion as a Second Modality
- End-to-End Vision Tokenizer Tuning
- GAITGen: Disentangled Motion-Pathology Impaired Gait Generative Model -- Bringing Motion Generation to the Clinical Domain
- AGIR: Assessing 3D Gait Impairment with Reasoning based on LLMs
- DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
- The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
- VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Regress, Don't Guess -- A Regression-like Loss on Number Tokens for Language Models
- MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
- LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning
- MotionLLM: Understanding Human Behaviors from Human Motions and Videos
- M3GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
- Instruction Tuning With Loss Over Instructions
- MotionChain: Conversational Motion Controllers via Multimodal Prompts
- LLMs are Good Action Recognizers
- Differentiable Biomechanics Unlocks Opportunities for Markerless Motion Capture
- Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
- From Skin to Skeleton: Towards Biomechanically Accurate 3D Digital Humans
- MoMask: Generative Masked Modeling of 3D Human Motions
- AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond
- LocoMuJoCo: A Comprehensive Imitation Learning Benchmark for Locomotion
- xVal: A Continuous Numerical Tokenization for Scientific Language Models
- Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
- MotionGPT: Human Motion as a Foreign Language
- QLoRA: Efficient Finetuning of Quantized LLMs
- Visual Instruction Tuning
- T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations
- Learning 3D Human Pose Estimation from Dozens of Datasets using a Geometry-Aware Autoencoder to Bridge Between Skeleton Formats
- Executing your Commands via Motion Diffusion in Latent Space
- Rethinking the Objectives of Vector-Quantized Tokenizers for Image Synthesis
- TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts
- MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control
- OSSO: Obtaining Skeletal Shape from Outside
- Training Compute-Optimal Large Language Models
- Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory
- SOMA: Solving Optical Marker-Based MoCap Automatically
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAE
- Jukebox: A Generative Model for Music
- Hierarchical Quantized Autoencoders
- Scaling Laws for Neural Language Models
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
- AMASS: Archive of Motion Capture as Surface Shapes
- Neural Discrete Representation Learning
- Attention Is All You Need
- Muscle contributions to propulsion and support during running
- ISB recommendation on definitions of joint coordinate systems of various joints for the reporting of human joint motion—Part II: shoulder, elbow, wrist and hand
- ISB recommendation on definitions of joint coordinate system of various joints for the reporting of human joint motion—part I: ankle, hip, and spine
- Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
Cited by
Discussions
Related