Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
2019/10/23 by Colin Raffel, Noam Shazeer, Raffel, Colin +16 · 4 voices · 1173 citations
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1910.10683
openalex publication_date 2019/10/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP). The effectiveness of transfer learning has given rise to a diversity of approaches, methodology, and practice. In this paper, we explore the landscape of transfer learning techniques for NLP by introducing a unified framework that converts all text-based language problems into a text-to-text format. Our systematic study compares pre-training objectives, architectures, unlabeled data sets, transfer approaches, and other factors on dozens of language understanding tasks. By combining the insights from our exploration with scale and our new ``Colossal Clean Crawled Corpus'', we achieve state-of-the-art results on many benchmarks covering summarization, question answering, text classification, and more. To facilitate future work on transfer learning for NLP, we release our data set, pre-trained models, and code.
Cited by
- Merge before Forget: A Single LoRA Continual Learning via Continual Merging
- Identifying social bots via heterogeneous motifs based on Naïve Bayes model
- RealCamo: Boosting Real Camouflage Synthesis with Layout Controls and Textual-Visual Guidance
- ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
- Bridging Global Intent with Local Details: A Hierarchical Representation Approach for Semantic Validation in Text-to-SQL
- Visual Autoregressive Modelling for Monocular Depth Estimation
- SPECTRE: Spectral Pre-training Embeddings with Cylindrical Temporal Rotary Position Encoding for Fine-Grained sEMG-Based Movement Decoding
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Towards a Relevance Posterior in Neural Information Access
- BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi
- The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
- Neonatal Hypoxic-ischaemic Encephalopathy Classification from the EEG and HRV Signals Using a Conformer based Masked Autoencoder
- Accelerating Language Model Workflows with Prompt Choreography
- Enhancing Code Understanding for Impact Analysis by Combining Transformers and Program Dependence Graphs
- Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning
- Where meaning lives: Layer-wise accessibility of psycholinguistic features in encoder and decoder language models
- PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation
- Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting
- Gnowsis: Multimodal multitask learning for oral proficiency assessments
- AutoFed: Personalized Federated Traffic Prediction via Adaptive Prompt
- VaLiDRec: Variable-Length LLM-Aligned Semantic IDs for Generative Recommendation
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
- Hierarchical Grading in Large Language Models
- Open Your Model’s Eyes: Video and Context-Aware Multimodal Backchannel Prediction
- Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- FARM: Find Anything using Relational Spatial Memory
- Neural Machine Translation for Low-Resource Tangkhul--English
- GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
- Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering
- Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL
- Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
- FlashEvaluator: Expanding Search Space with Parallel Sequence-Level Evaluation
- Adapting a Text-to-Audio Model for Room Impulse Response Generation
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Approximation Capabilities of Feedforward Neural Networks with GELU Activations
- Cross-Semantic Transfer Learning for High-Dimensional Linear Regression
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- CEMG: Collaborative-Enhanced Multimodal Generative Recommendation
- RefineBridge: Generative Bridge Models Improve Financial Forecasting by Foundation Models
- Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
- ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- When LLMs fall short in Deductive Coding: Model Comparison and Human AI Collaboration Workflow Design
- Artificial or Just Artful? Do LLMs Bend the Rules in Programming?
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
- On the Effectiveness of Instruction-Tuning Local LLMs for Identifying Software Vulnerabilities
- Reason2Decide: Rationale-Driven Multi-Task Learning
- QuarkAudio Technical Report
- Making Large Language Models Efficient Dense Retrievers
- Diacritic Restoration for Low-Resource Indigenous Languages: Case Study with Bribri and Cook Islands Māori
- Event Extraction in Large Language Model
- OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
- AraMix: Recycling, Refiltering, and Deduplicating to Deliver the Largest Arabic Pretraining Corpus
- A Comparative Study of Light-weight Language Models for PII Masking and their Deployment for Real Conversational Texts
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- A Data-Centric Approach to Generalizable Speech Deepfake Detection
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- Bridging Natural Language and Formal Specification--Automated Translation of Software Requirements to LTL via Hierarchical Semantics Decomposition Using LLMs
- Easy Adaptation: An Efficient Task-Specific Knowledge Injection Method for Large Models in Resource-Constrained Environments
- Exploiting ID-Text Complementarity via Ensembling for Sequential Recommendation
- Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering
- Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
- Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs
- GinSign: Grounding Natural Language Into System Signatures for Temporal Logic Translation
- Yuan-TecSwin: A text conditioned Diffusion model with Swin-transformer blocks
- Factorized Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
- The Evolution of Reranking Models in Information Retrieval: From Heuristic Methods to Large Language Models
- DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Prompt Repetition Improves Non-Reasoning LLMs
- Spatia: Video Generation with Updatable Spatial Memory
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- Per-Axis Weight Deltas for Frequent Model Updates
- Dual-objective Language Models: Training Efficiency Without Overfitting
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
- DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
- Advancing Bangla Machine Translation Through Informal Datasets
- LINA: Learning INterventions Adaptively for Physical Alignment and Generalization in Diffusion Models
- MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
- Alada: Alternating Adaptation of Momentum Method for Memory-Efficient Matrix Optimization
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
- Informing Acquisition Functions via Foundation Models for Molecular Discovery
- SPAR: Session-based Pipeline for Adaptive Retrieval on Legacy File Systems
- PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
- Robust Motion Generation using Part-level Reliable Data from Videos
- Resting Neurons, Active Insights: Improving Input Sparsification for Large Language Models
- FuXi-γ: Efficient Sequential Recommendation with Exponential-Power Temporal Encoder and Diagonal-Sparse Positional Mechanism
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- AdaGradSelect: An adaptive gradient-guided layer selection method for efficient fine-tuning of SLMs
- PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation
- AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
- Multi-Intent Spoken Language Understanding: Methods, Trends, and Challenges
- REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation
- SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- Stronger Normalization-Free Transformers
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
- Learning by Analogy: A Causal Framework for Composition Generalization
- Zero-shot Adaptation of Stable Diffusion via Plug-in Hierarchical Degradation Representation for Real-World Super-Resolution
- Watermarks for Language Models via Probabilistic Automata
- Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
- Defining Cost Function of Steganography with Large Language Models
- Interpreto: An Explainability Library for Transformers
- Learning Patient-Specific Disease Dynamics with Latent Flow Matching for Longitudinal Imaging Generation
- Food Image Generation on Multi-Noun Categories
- UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
- Financial News Summarization: Can extractive methods still offer a true alternative to LLMs?
- The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
- Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Luxical: High-Speed Lexical-Dense Text Embeddings
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
- SETUP: Sentence-level English-To-Uniform Meaning Representation Parser
- FOAM: Blocked State Folding for Memory-Efficient LLM Training
- A Patient-Doctor-NLP-System to contest inequality for less privileged
- On Memory: A comparison of memory mechanisms in world models
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Teaching Language Models Mechanistic Explainability Through Arrow-Pushing
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
- From Text to Returns: Using Large Language Models for Mutual Fund Portfolio Optimization and Risk-Adjusted Allocation
- Scaling and Transferability of Annealing Strategies in Large Language Model Training
- 2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- Light-X: Generative 4D Video Rendering with Camera and Illumination Control
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
- OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models
- Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
- RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
- Stable Signer: Hierarchical Sign Language Generative Model
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment
- RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
- Data-Free Pruning of Self-Attention Layers in LLMs
- FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion Generation
- MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- Network Self-Configuration based on Fine-Tuned Small Language Models
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- Offloading Artificial Intelligence Workloads across the Computing Continuum by means of Active Storage Systems
- Parameter-Efficient Subspace Optimization for LLM Fine-Tuning
- Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
- Low-Rank Prehab: Preparing Neural Networks for SVD Compression
- Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade
- LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLM
- TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
- DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
- On The Finetuning of MLIPs Through the Lens of Iterated Maps With BPTT
- Table as a Modality for Large Language Models
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- WaterSearch: A Quality-Aware Search-based Watermarking Framework for Large Language Models
- Financial Text Classification Based On rLoRA Finetuning On Qwen3-8B model
- Low-Bitrate Video Compression through Semantic-Conditioned Diffusion
- Simplex-Optimized Hybrid Ensemble for Large Language Model Text Detection Under Generative Distribution Drif
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
- Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
- Tourism Question Answer System in Indian Language using Domain-Adapted Foundation Models
- Decoding the Past: Explainable Machine Learning Models for Dating Historical Texts
- Language-conditioned world model improves policy generalization by reading environmental descriptions
- Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
- CoFiRec: Coarse-to-Fine Tokenization for Generative Recommendation
- Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
- SingleQuant: Efficient Quantization of Large Language Models in a Single Pass
- Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- Uni-Hema: Unified Model for Digital Hematopathology
- A Systematic Study of In-the-Wild Model Merging for Large Language Models
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- 3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
- CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
- Towards Audio Token Compression in Large Audio Language Models
- CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
- On the Origin of Algorithmic Progress in AI
- Text-Guided Semantic Image Encoder
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
- Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- GigaWorld-0: World Models as Data Engine to Empower Embodied AI
- Cisco Time Series Model Technical Report
- Decoupling and Damping: Structurally-Regularized Gradient Matching for Multimodal Graph Condensation
- Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- FineXtrol: Controllable Motion Generation via Fine-Grained Text
- Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
- Large Language Models for the Summarization of Czech Documents: From History to the Present
- Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
- Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
- ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
- A symbolic Perl algorithm for the unification of Nahuatl word spellings
- IRSDA: An Agent-Orchestrated Framework for Enterprise Intrusion Response
- DELTA: Language Diffusion-based EEG-to-Text Architecture
- An Invariant Latent Space Perspective on Language Model Inversion
- Zero-Shot Video Deraining with Video Diffusion Models
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- A Systematic Study of Compression Ordering for Large Language Models
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- Plan-X: Instruct Video Generation via Semantic Planning
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- CREST: Improving Interpretability and Effectiveness of Troubleshooting at Ericsson through Criterion-Specific Trouble Report Retrieval
- Energy Scaling Laws for Diffusion Models: Quantifying Compute and Carbon Emissions in Image Generation
- Supervised Fine Tuning of Large Language Models for Domain Specific Knowledge Graph Construction:A Case Study on Hunan's Historical Celebrities
- Developmental Atlas of Attention Head Specialization: Spacing, Stranding, and the Capacity Tax of BPE Tokenization
- Exploring Scientific Debt: Harnessing AI for SATD Identification in Scientific Software
- Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
- When Structure Doesn't Help: LLMs Do Not Read Text-Attributed Graphs as Effectively as We Expected
- You Only Forward Once: An Efficient Compositional Judging Paradigm
- LAOF: Robust Latent Action Learning with Optical Flow Constraints
- Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
- NLP Datasets for Idiom and Figurative Language Tasks
- AskDB: An LLM Agent for Natural Language Interaction with Relational Databases
- UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
- Tokenize Once, Recommend Anywhere: Unified Item Tokenization for Multi-domain LLM-based Recommendation
- Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- IPR-1: Interactive Physical Reasoner
- Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language
- SplitFlux: Learning to Decouple Content and Style from a Single Image
- PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
- Insert In Style: A Zero-Shot Generative Framework for Harmonious Cross-Domain Object Composition
- Effective Code Membership Inference for Code Completion Models via Adversarial Prompts
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- Gradient-descent methods for quantum detector tomography
- Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- ArbESC+: Arabic Enhanced Edit Selection System Combination for Grammatical Error Correction Resolving conflict and improving system combination in Arabic GEC
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- DEVAL: A Framework for Evaluating and Improving the Derivation Capability of Large Language Models
- Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
- Intermediate N-Gramming: Deterministic and Fast N-Grams For Large N and Large Datasets
- CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product
- Scalable and Efficient Large-Scale Log Analysis with LLMs: An IT Software Support Case Study
- CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving
- Translation Entropy: A Statistical Framework for Evaluating Translation Systems
- NeuroLex: A Lightweight Domain Language Model for EEG Report Understanding and Generation
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- Multimodal Large Language Models as Image Classifiers
- CAT-ID2: Category-Tree Integrated Document Identifier Learning for Generative Retrieval In E-commerce
- Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
- Designed to Spread: Generative Approaches to Enhance Information Diffusion
- HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
- Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
- OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description
- GeoMVD: Geometry-Enhanced Multi-View Generation Model Based on Geometric Information Extraction
- KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
- Context-Emotion Aware Therapeutic Dialogue Generation: A Multi-component Reinforcement Learning Approach to Language Models for Mental Health Support
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Hindsight Distillation Reasoning with Knowledge Encouragement Preference for Knowledge-based Visual Question Answering
- Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
- Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
- Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots
- Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
- CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
- Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
- CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
- Introducing A Bangla Sentence - Gloss Pair Dataset for Bangla Sign Language Translation and Research
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
- JobSphere: An AI-Powered Multilingual Career Copilot for Government Employment Platforms
- Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation
- A Unified Geometric Field Theory Framework for Transformers: From Manifold Embeddings to Kernel Modulation
- Multi-Granularity Mutual Refinement Network for Zero-Shot Learning
- Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
- WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- Majority Rules: LLM Ensemble is a Winning Approach for Content Categorization
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
- Beyond English: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs
- Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
- How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models
- LLM Optimization Unlocks Real-Time Pairwise Reranking
- A Decentralized Retrieval Augmented Generation System with Source Reliabilities Secured on Blockchain
- Route Experts by Sequence, not by Token
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Reaction Prediction via Interaction Modeling of Symmetric Difference Shingle Sets
- Seq2Seq Models Reconstruct Visual Jigsaw Puzzles without Seeing Them
- TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
- Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- LLaDA-Rec: Discrete Diffusion for Parallel Semantic ID Generation in Generative Recommendation
- L2T-Hyena: Enhancing State-Space Models with an Adaptive Learn-to-Teach Framework
- MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
- TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
- A Representation Sharpening Framework for Zero Shot Dense Retrieval
- ManufactuBERT: Efficient Continual Pretraining for Manufacturing
- Reasoning-Guided Claim Normalization for Noisy Multilingual Social Media Posts
- Search Is Not Retrieval: Decoupling Semantic Matching from Contextual Assembly in RAG
- REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
- PromptSep: Generative Audio Separation via Multimodal Prompting
- RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- ScaleDL: Towards Scalable and Efficient Runtime Prediction for Distributed Deep Learning Workloads
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
- Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis
- OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
- Diffusion Language Models are Super Data Learners
- Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification
- Data-Efficient Adaptation and a Novel Evaluation Method for Aspect-based Sentiment Analysis
- Targeted Error Correction in Knowledge Distillation: Small Language Models Surpass GPT
- Dynamic Reflections: Probing Video Representations with Text Alignment
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
- ReleaseEval: A Benchmark for Evaluating Language Models in Automated Release Note Generation
- Memory-Efficient Training with In-Place FFT Implementation
- Random Initialization of Gated Sparse Adapters
- HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
- Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning
- Air Pollution Forecasting in Bucharest
- Listwise Preference Diffusion Optimization for User Behavior Trajectories Prediction
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
- Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
- BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
- Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
- Probability-Biased Attention over Directed Bipartite Graphs for Long-Tail ICD Coding
- E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
- EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
- Integrating Ontologies with Large Language Models for Enhanced Control Systems in Chemical Engineering
- The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
- An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-aware Fine-tuning
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
- UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
- Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
- Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignment
- Towards Scaling Laws for Symbolic Regression
- Predicate Renaming via Large Language Models
- NeuronMM: High-Performance Matrix Multiplication for LLM Inference on AWS Trainium
- Revisiting Multilingual Data Mixtures in Language Model Pretraining
- SoK: Honeypots & LLMs, More Than the Sum of Their Parts?
- VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
- Benchmarking Generative AI Against Bayesian Optimization for Constrained Multi-Objective Inverse Design
- How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model
- ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling
- Explicit Note-Event Tokenization and Pitch-Validity Constrained Decoding for MIDI-to-Tablature Transcription
- CG-World: A Large-Scale World-State Dataset and Protocol for World Models
- Between Gradient and Natural Gradient: A Continuum of LoRA Initializations
- Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
- Mixture-of-Depths Attention
- Training Language Models via Neural Cellular Automata
- TridentServe: A Stage-level Serving System for Diffusion Pipelines
- Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- MedFusionT5: Cross-Modal Attention Boosts Semantic Quality and Reduces Hallucinations in Dental AI
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- MANDO-LLM: Heterogeneous Graph Transformers with Large Language Models for Smart Contract Vulnerability Detection
- HRM-Text: Efficient Pretraining Beyond Scaling
- Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance
- A Bitter Lesson for Data Filtering
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
- On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
- RL makes MLLMs see better than SFT
- Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
- PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
- Sequences of Logits Reveal the Low Rank Structure of Language Models
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- Group Relative Attention Guidance for Image Editing
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
- Exploring the Influence of Relevant Knowledge for Natural Language Generation Interpretability
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Utilising Large Language Models for Generating Effective Counter Arguments to Anti-Vaccine Tweets
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion Models
- Improving the Straight-Through Estimator with Zeroth-Order Information
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
- Large language model-based task planning for service robots: A review
- Autoregressive Styled Text Image Generation, but Make it Reliable
- Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning
- Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- SALSA: Single-pass Autoregressive LLM Structured Classification
- Frustratingly Easy Task-aware Pruning for Large Language Models
- E2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Reranker
- The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
- Expert Merging in Sparse Mixture of Experts with Nash Bargaining
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- NetBurst: Event-Centric Forecasting of Bursty, Intermittent Time Series
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- HARMONY: Hidden Activation Representations and Model Output-Aware Uncertainty Estimation for Vision-Language Models
- SentiMaithili: A Benchmark Dataset for Sentiment and Reason Generation for the Low-Resource Maithili Language
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
- PARL: Prompt-based Agents for Reinforcement Learning
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models
- Chronos-2: From Univariate to Universal Forecasting
- Bi-Level Optimization for Generative Recommendation: Bridging Tokenization and Generation
- A Reinforcement Learning Framework for Robust and Secure LLM Watermarking
- xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs
- Improving named entity correctness of abstractive summarization by generative negative sampling
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
- Evaluating Latent Knowledge of Public Tabular Datasets in Large Language Models
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Relative-Based Scaling Law for Neural Language Models
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Fast Inference via Hierarchical Speculative Decoding
- Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
- Machine Text Detectors are Membership Inference Attacks
- ELUTQ: Efficient LUT-Aware Quantization for Deploying Large Language Models on Edge Devices
- Restoring Pruned Large Language Models via Lost Component Compensation
- CPSVD: Enhancing Large Language Model Compression via Column-Preserving Singular Value Decomposition
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Selecting and Combining Large Language Models for Scalable Code Clone Detection
- ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
- CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- Large language models for folktale type automation based on motifs: Cinderella case study
- Large-scale User Game Lifecycle Representation Learning
- From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering
- Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs
- Identity-Aware Large Language Models require Cultural Reasoning
- Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
- Benchmarking Probabilistic Time Series Forecasting Models on Neural Activity
- Unbiased Gradient Low-Rank Projection
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
- DSEBench: A Test Collection for Explainable Dataset Search with Examples
- Rethinking On-policy Optimization for Query Augmentation
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- Annotation-Efficient Universal Honesty Alignment
- Watermark Robustness and Radioactivity May Be at Odds in Federated Learning
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Back to Bytes: Revisiting Tokenization Through UTF-8
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding
- Neuronal Group Communication for Efficient Neural representation
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Safe and Efficient In-Context Learning via Risk Control
- Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
- Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
- Policy Transfer for Continuous-Time Reinforcement Learning: A (Rough) Differential Equation Approach
- FarsiMCQGen: a Persian Multiple-choice Question Generation Framework
- DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
- DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Midtraining Bridges Pretraining and Posttraining Distributions
- ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
- FraQAT: Quantization Aware Training with Fractional bits
- Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
- Efficient Seq2seq Coreference Resolution Using Entity Representations
- Hierarchical Semantic Retrieval with Cobweb
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
- Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm
- Constraint-Driven Small Language Models Based on Agent and OpenAlex Knowledge Graph: Mining Conceptual Pathways and Discovering Innovation Points in Academic Papers
- LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
- Inferred global dense residue transition graphs from primary structure sequences enable protein interaction prediction via directed graph convolutional neural networks
- Nondeterminism-Aware Optimistic Verification for Floating-Point Neural Networks
- FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
- LTR-ICD: A Learning-to-Rank Approach for Automatic ICD Coding
- PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- Hypernetworks for Perspectivist Adaptation
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- Cautious Weight Decay
- Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
- Reinforced Preference Optimization for Recommendation
- Chimera: State Space Models Beyond Sequences
- Ethic-BERT: An Enhanced Deep Learning Model for Ethical and Non-Ethical Content Classification
- Readout Representation: Redefining Neural Codes by Input Recovery
- Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- GenCNER: A Generative Framework for Continual Named Entity Recognition
- Are Large Language Models Effective Knowledge Graph Constructors?
- Vision-LLMs for Spatiotemporal Traffic Forecasting
- Neural Weight Compression for Language Models
- A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
- Domain-Specific Data Generation Framework for RAG Adaptation
- An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models
- Saudi Sign Language Translation Using T5
- CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
- PAGE: Prompt Augmentation for text Generation Enhancement
- Direct Multi-Token Decoding
- MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models
- Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
- R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
- ReMix: Towards a Unified View of Consistent Character Generation and Editing
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection
- Forget Attention: Importance-Aware Attention Is All You Need
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
- Lost in the Middle: An Emergent Property from Information Retrieval Demands in LLMs
- PatentVision: A multimodal method for drafting patent applications
- Patentformer: A demonstration of AI-assisted automated patent drafting
- Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic
- The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
- AdaPM: a Partial Momentum Algorithm for LLM Training
- Provable Watermarking for Data Poisoning Attacks
- Web Crawler Restrictions, AI Training Datasets & Political Biases
- The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
- Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- Vocabulary embeddings organize linguistic structure early in language model training
- LOTION: Smoothing the Optimization Landscape for Quantized Training
- DACIP-RC: Domain Adaptive Continual Instruction Pre-Training via Reading Comprehension on Business Conversations
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- On the Convergence of Moral Self-Correction in Large Language Models
- Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Textual interpretation of transient image classifications from large language models
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
- Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
- Mid-Training of Large Language Models: A Survey
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- GUIDE: Guided Initialization and Distillation of Embeddings
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting
- Training Dynamics Impact Post-Training Quantization Robustness
- Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
- LARA-Gen: Enabling Continuous Emotion Control for Music Generation Models via Latent Affective Representation Alignment
- Mixture of Neuron Experts
- Membership Inference Attacks on Tokenizers of Large Language Models
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
- Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
- FoilDiff: A Hybrid Transformer Backbone for Diffusion-based Modelling of 2D Airfoil Flow Fields
- Activation Steering with a Feedback Controller
- Thai Semantic End-of-Turn Detection for Real-Time Voice Agents
- Large Language Models Hallucination: A Comprehensive Survey
- Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Smart Paste: Automatically Fixing Copy/Paste for Google Developers
- Allocation of Parameters in Transformers
- On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
- CarbonX: An Open-Source Tool for Computational Decarbonization Using Time Series Foundation Models
- When and Where do Events Switch in Multi-Event Video Generation?
- DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
- Distributed Low-Communication Training with Decoupled Momentum Optimization
- MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Uncovering the Computational Ingredients of Human-Like Representations in LLMs
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
- Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
- CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
- Black-Box Time-Series Domain Adaptation via Cross-Prompt Foundation Models
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- Decomposing Attention To Find Context-Sensitive Neurons
- Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors
- Retrieval-Augmented Framework for LLM-Based Clinical Decision Support
- Learn to Guide Your Diffusion Model
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Bayesian Influence Functions for Hessian-Free Data Attribution
- Regression Language Models for Code
- Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
- Indirect Attention: Turning Context Misalignment into a Feature
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Revoking Amnesia: RL-based Trajectory Optimization to Resurrect Erased Concepts in Diffusion Models
- From Cheap Geometry to Expensive Physics: Elevating Neural Operators via Latent Shape Pretraining
- Layer-wise dynamic rank for compressing large language models
- Per-example gradients: a new frontier for understanding and improving optimizers
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution
- From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
- Pushing LLMs to Their Logical Reasoning Bound: The Role of Data Reasoning Intensity
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents
- Conda: Column-Normalized Adam for Training Large Language Models Faster
- PET: Preference Evolution Tracking with LLM-Generated Explainable Distribution
- Watermarking Diffusion Language Models
- PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
- AGNOMIN -- Architecture Agnostic Multi-Label Function Name Prediction
- <scp>AI</scp> Methods for Antimicrobial Peptides: Progress and Challenges
- MapGenerator: a framework for learning a diffusion model for text promptable map generation
- The Rise of AfricaNLP: A Survey of Contributions, Contributors, Community Impact, and Bibliometric Analysis
- Reinforcement Mid-Training
- Analyzing and Evaluating Unbiased Language Model Watermark
- An Ensemble Framework for Unbiased Language Model Watermarking
- Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
- Investigating Multi-layer Representations for Dense Passage Retrieval
- GSID: Generative Semantic Indexing for E-Commerce Product Understanding
- Estimating Time Series Foundation Model Transferability via In-Context Learning
- MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
- Training Optimal Large Diffusion Language Models
- Temporal Generalization: A Reality Check
- A systematic evaluation of Dutch large language models’ surprisal estimates in sentence, paragraph and book reading
- WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- An Improved Framework for Scaling Party Positions from Texts with Transformer
- Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
- PDE-Transformer: A Continuous Dynamical Systems Approach to Sequence Modeling
- How to Make Large Language Models Generate 100% Valid Molecules?
- PT2-LLM: Post-Training Ternarization for Large Language Models
- PlasGO: enhancing GO-based function prediction for plasmid-encoded proteins based on genetic structure
- LLM Watermark Evasion via Bias Inversion
- SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
- WoW: Towards a World omniscient World model Through Embodied Interaction
- JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation
- A model of errors in transformers
- What Is The Political Content in LLMs' Pre- and Post-Training Data?
- Advancing Natural Language Formalization to First Order Logic with Fine-tuned LLMs
- Beyond Textual Context: Structural Graph Encoding with Adaptive Space Alignment to alleviate the hallucination of LLMs
- Context Parametrization with Compositional Adapters
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining Approach
- Multilingual Vision-Language Models, A Survey
- FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration
- Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
- Enhancing Low-Rank Adaptation with Structured Nonlinear Transformations
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
- EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer
- RLP: Reinforcement as a Pretraining Objective
- Towards Transparent AI: A Survey on Explainable Language Models
- SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
- Query-Centric Graph Retrieval Augmented Generation
- IntSR: An Integrated Generative Framework for Search and Recommendation
- Distributed Specialization: Rare-Token Neurons in Large Language Models
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
- A short survey on almost orthogonal vectors in a few specific large dimensions
- RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
- FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets
- Prompt-Aware Scheduling for Low-Latency LLM Serving
- Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search
- Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error Overcorrection
- Measuring LLM Sensitivity in Transformer-based Tabular Data Synthesis
- Overcoming Black-box Attack Inefficiency with Hybrid and Dynamic Select Algorithms
- Concise and Sufficient Sub-Sentence Citations for Retrieval-Augmented Generation
- RedHerring Attack: Testing the Reliability of Attack Detection
- Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos
- Understanding and Enhancing Mask-Based Pretraining towards Universal Representations
- Performance Consistency of Learning Methods for Information Retrieval Tasks
- OLaPh: Optimal Language Phonemizer
- Embodied AI: From LLMs to World Models
- Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
- Learning Contextual Retrieval for Robust Conversational Search
- SeMob: Semantic Synthesis for Dynamic Urban Mobility Prediction
- Mamba Modulation: On the Length Generalization of Mamba
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- Looped Transformers with Source-Centered State Evolution
- ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
- What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production Variability
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- MAD: Motion Appearance Decoupling for efficient Driving World Models
- HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning
- ELF: Embedded Language Flows
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- RELISH: LLM REgression with a Latent Iterative State Head
- Legal text summarization via judicial syllogism with large language models
- ChatGPT: More Than a “Weapon of Mass Deception” Ethical Challenges and Responses from the Human-Centered Artificial Intelligence (HCAI) Perspective
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- Transporting Task Vectors across Different Architectures without Training
- BaGuaLu
- epiGPTope: A Machine Learning-Based Epitope Generator and Classifier
- AI for social science and social science of AI: A survey
- CamemBERT: a Tasty French Language Model
- LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
- Speculating LLMs' Chinese Training Data Pollution from Their Tokens
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- Reinforcement Learning on Pre-Training Data
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
- AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group Swapping
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexity
- Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
- False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
- Financial Risk Relation Identification through Dual-view Adaptation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
- A Rhythm-Aware Phrase Insertion for Classical Arabic Poetry Composition
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Advances in Large Language Models for Medicine
- Enhancing Transformer-Based Rerankers with Synthetic Data and LLM-Based Supervision
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering?
- MOMEMTO: Patch-based Memory Gate Model in Time Series Foundation Model
- Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
- The Narcissus Hypothesis: Descending to the Rung of Illusion
- Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation
- CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
- AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- Variational Task Vector Composition
- Analyzing Memory Effects in Large Language Models through the lens of Cognitive Psychology
- MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
- A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
- A Unified AI Approach for Continuous Monitoring of Human Health and Diseases from Intensive Care Unit to Home with Physiological Foundation Models (UNIPHY+)
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- DNA-DetectLLM: Unveiling AI-Generated Text via a DNA-Inspired Mutation-Repair Paradigm
- Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval
- Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
- Chunk Knowledge Generation Model for Enhanced Information Retrieval: A Multi-task Learning Approach
- Structured Information for Improving Spatial Relationships in Text-to-Image Generation
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- On the Convergence of Muon and Beyond
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- Language Modeling with Learned Meta-Tokens
- A Comparative Analysis of Transformer Models in Social Bot Detection
- Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
- MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
- Super-Linear: A Lightweight Pretrained Mixture of Linear Experts for Time Series Forecasting
- Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- Synthetic bootstrapped pretraining
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
- Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
- Enhancing Time Awareness in Generative Recommendation
- TFMAdapter: Lightweight Instance-Level Adaptation of Foundation Models for Forecasting with Covariates
- Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
- How Can Quantum Deep Learning Improve Large Language Models?
- Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach
- TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Do Activation Verbalization Methods Convey Privileged Information?
- Membership Inference Attack against Large Language Model-based Recommendation Systems: A New Distillation-based Paradigm
- Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- Positional Encoding via Token-Aware Phase Attention
- An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- Automated Generation of Research Workflows from Academic Papers: A Full-text Mining Framework
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
- Pun Unintended: LLMs and the Illusion of Humor Understanding
- VQL: An End-to-End Context-Aware Vector Quantization Attention for Ultra-Long User Behavior Modeling
- AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
- CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
- Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents
- Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- Contrastive Prompt Clustering for Weakly Supervised Semantic Segmentation
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation
- Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
- DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation
- WebSight: A Vision-First Architecture for Robust Web Agents
- Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
- Out of One, Many: Using Language Models to Simulate Human Samples
- Towards Understanding Visual Grounding in Visual Language Models
- ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
- CoPE: A Lightweight Complex Positional Encoding
- Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research
- Large Language Models in Document Intelligence: A Comprehensive Survey, Recent Advances, Challenges, and Future Trends
- Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
- Character-Level Perturbations Disrupt LLM Watermarks
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- CrossPT: Exploring Cross-Task Transferability through Multi-Task Prompt Tuning
- LLMs as Agentic Cooperative Players in Multiplayer UNO
- Artificial intelligence-driven computational methods for antibody design and optimization
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation
- Adversarial Attacks Against Automated Fact-Checking: A Survey
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- Handling Open-Vocabulary Constructs in Formalizing Specifications: Retrieval-Augmented Parsing with Expert Knowledge
- Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Trees
- Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
- Improving Machine Learning-Based Robot Self-Collision Checking with Input Positional Encoding
- Multi-view-guided Passage Reranking with Large Language Models
- Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communities
- Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
- CP-Model-Zoo: A Natural Language Query System for Constraint Programming Models
- From Detection to Mitigation: Addressing Gender Bias in Chinese Texts via Efficient Tuning and Voting-Based Rebalancing
- Learning Decomposed Contextual Token Representations from Pretrained and Collaborative Signals for Generative Recommendation
- COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
- TIDE: Achieving Balanced Subject-Driven Image Generation via Target-Instructed Diffusion Enhancement
- Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- LoaQ: Layer-wise Output Approximation Quantization
- SLiNT: Structure-aware Language Model with Injection and Contrastive Training for Knowledge Graph Completion
- HANRAG: Heuristic Accurate Noise-resistant Retrieval-Augmented Generation for Multi-hop Question Answering
- Methodological Insights into Structural Causal Modelling and Uncertainty-Aware Forecasting for Economic Indicators
- TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
- Benchmarking Gender and Political Bias in Large Language Models
- DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
- From Joy to Fear: A Benchmark of Emotion Estimation in Pop Song Lyrics
- Few-Shot Query Intent Detection via Relation-Aware Prompt Learning
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- PLaMo 2 Technical Report
- L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Reverse Browser: Vector-Image-to-Code Generator
- Hunyuan-MT Technical Report
- Rethinking the long-range dependency in Mamba/SSM and transformer models
- SMooGPT: Stylized Motion Generation using Large Language Models
- LMAE4Eth: Generalizable and Robust Ethereum Fraud Detection by Exploring Transaction Semantics and Masked Graph Embedding
- Weakly-Supervised Learning of Dense Functional Correspondences
- Madness, Cannibalism, and Traditional Fiction between Lu Xun and Mo Yan
- Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- Binary Quantization For LLMs Through Dynamic Grouping
- Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- AIVA: An AI-based Virtual Companion for Emotion-aware Interaction
- LLMCDSR: Enhancing Cross-Domain Sequential Recommendation with Large Language Models
- Lesion-Aware Visual-Language Fusion for Automated Image Captioning of Ulcerative Colitis Endoscopic Examinations
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Distributional Semantics: Meaning Through Culture and Interaction
- DivMerge: A divergence-based model merging method for multi-tasking
- LLMs that Understand Processes: Instruction-tuning for Semantics-Aware Process Mining
- FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- Systematic Integration of Attention Modules into CNNs for Accurate and Generalizable Medical Image Diagnosis
- chDzDT: Word-level morphology-aware language model for Algerian social media text
- Clinical Metadata Guided Limited-Angle CT Image Reconstruction
- Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
- MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
- Multitask Battery Management with Flexible Pretraining
- GradeSQL: Test-Time Inference with Outcome Reward Models for Text-to-SQL Generation from Large Language Models
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- Hierarchical Motion Captioning Utilizing External Text Data Source
- Collaborative local–global context modeling for session-based recommendation
- Two-flow Feedback Multi-scale Progressive Generative Adversarial Network
- Optimizing In-Context Learning for Efficient Full Conformal Prediction
- A Privacy-Preserving Recommender for Filling Web Forms Using a Local Large Language Model
- A case study of forensic psychiatry experts' reports analysis through large language models
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
- No Clustering, No Routing: How Transformers Actually Process Rare Tokens
- Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
- Training Text-to-Molecule Models with Context-Aware Tokenization
- Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank
- FLORA: Efficient Synthetic Data Generation for Object Detection in Low-Data Regimes via finetuning Flux LoRA
- QZhou-Embedding Technical Report
- Measuring Attribution in Natural Language Generation Models
- VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs
- Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
- Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
- Rethinking Layer-wise Model Merging through Chain of Merges
- Zero‐ and few‐shot prompting of generative large language models provides weak assessment of risk of bias in clinical trials
- Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
- MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
- Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
- KG-CQR: Leveraging Structured Relation Representations in Knowledge Graphs for Contextual Query Retrieval
- Waver: Wave Your Way to Lifelike Video Generation
- TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
- Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
- OneRec-V2 Technical Report
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
- AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- Generative AI for Testing of Autonomous Driving Systems: A Survey
- TComQA: Extracting Temporal Commonsense from Text
- QuesGenie: Intelligent Multimodal Question Generation
- Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models
- ELIXIR: Efficient and LIghtweight model for eXplaIning Recommendations
- SIExVulTS: Sensitive Information Exposure Vulnerability Detection System using Transformer Models and Static Analysis
- Database Entity Recognition with Data Augmentation and Deep Learning
- MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
- PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback
- Enhancing Document VQA Models via Retrieval-Augmented Generation
- Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Constraint Matters: Multi-Modal Representation for Reducing Mixed-Integer Linear programming
- Scaling Laws for Task-Stratified Knowledge in Post-Training Quantized Large Language Models
- Principled Detection of Hallucinations in Large Language Models via Multiple Testing
- Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- Mapping the Course for Prompt-based Structured Prediction
- Reasoning is about giving reasons
- Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
- In2x at WMT25 Translation Task
- Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
- MoVieDrive: Multi-Modal Multi-View Urban Scene Video Generation
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Jointly Extracting Interventions, Outcomes, and Findings from RCT Reports with LLMs
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- A Fully Spectral Neuro-Symbolic Reasoning Architecture with Graph Signal Processing as the Computational Backbone
- Graph Concept Bottleneck Models
- A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- PC-Sampler: Position-Aware Calibration of Decoding Bias in Masked Diffusion Models
- Maximum Score Routing For Mixture-of-Experts
- EgoTwin: Dreaming Body and View in First Person
- RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
- Improving Detection of Watermarked Language Models
- SEA-BED: Southeast Asia Embedding Benchmark
- MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
- Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design
- SoK: Data Minimization in Machine Learning
- Neural Machine Translation for Coptic-French: Strategies for Low-Resource Ancient Languages
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- Semantic IDs for Joint Generative Search and Recommendation
- NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- Improving Generative Cross-lingual Aspect-Based Sentiment Analysis with Constrained Decoding
- Large Language Models for Summarizing Czech Historical Documents and Beyond
- Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models
- Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
- μ-Parametrization for Mixture of Experts
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
- Improving Dense Passage Retrieval with Multiple Positive Passages
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
- Prompt-Based Approach for Czech Sentiment Analysis
- Leveraging Large Language Models for Rare Disease Named Entity Recognition
- Boosting Action-Information via a Variational Bottleneck on Unlabelled Robot Videos
- Weakly Supervised Fine-grained Span-Level Framework for Chinese Radiology Report Quality Assurance
- Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
- TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
- VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign Flip
- DeCAL Tokenwise Compression
- Cut2Next: Generating Next Shot via In-Context Tuning
- SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling
- Artificial intelligence and corporate ideation systems
- Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
- Retrieval-Augmented Multi-Agent System for Rapid Statement of Work Generation
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
- Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
- Large Language Models for Subjective Language Understanding: A Survey
- On Understanding of the Dynamics of Model Capacity in Continual Learning
- GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
- Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
- Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks
- Enhancing Small-Scale Dataset Expansion with Triplet-Connection-based Sample Re-Weighting
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
- CharacterShot: Controllable and Consistent 4D Character Animation
- Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
- ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
- Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys
- Leveraging LLMs for Smart Cities Qualitative Data Analysis
- Explanatory argument extraction of correct answers in resident medical exams
- What Builds Effective In-Context Examples for Code Generation?
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- Streamlining Admission with LOR Insights: AI-Based Leadership Assessment in Online Master's Program
- Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
- SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
- REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation
- iFairy: the First 2-bit Complex LLM with All Parameters in \±1, ± i\
- Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
- Live Music Models
- Lightweight Transformers for Zero-Shot and Fine-Tuned Text-to-SQL Generation Using Spider
- HiD-VAE: Interpretable Generative Recommendation via Hierarchical and Disentangled Semantic IDs
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- Modelling and Classifying the Components of a Literature Review
- Large Language Model's Multi-Capability Alignment in Biomedical Domain
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
- STARE: Predicting Decision Making Based on Spatio-Temporal Eye Movements
- Majority Bit-Aware Watermarking For Large Language Models
- Tackling Distribution Shift in LLM via KILO: Knowledge-Instructed Learning for Continual Adaptation
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- Draw Your Mind: Personalized Generation via Condition-Level Modeling in Text-to-Image Diffusion Models
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
- Cropping outperforms dropout as an augmentation strategy for training self-supervised text embeddings
- Variety Is the Spice of Life: Detecting Misinformation with Dynamic Environmental Representations
- Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
- LLM-Prior: A Framework for Knowledge-Driven Prior Elicitation and Aggregation
- Multimodal Human-Intent Modeling for Contextual Robot-to-Human Handovers of Arbitrary Objects
- LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
- On the Evaluation of Large Language Models in Multilingual Vulnerability Repair
- Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment
- Tricks and Plug-ins for Gradient Boosting with Transformers
- TransAM: Transformer-Based Agent Modeling for Multi-Agent Systems via Local Trajectory Encoding
- LOST: Low-rank and Sparse Pre-training for Large Language Models
- FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
- AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
- The SMeL Test: A simple benchmark for media literacy in language models
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC
- A Survey on Data Security in Large Language Models
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- Zero-shot Compositional Action Recognition with Neural Logic Constraints
- SHAMI-MT: A Syrian Arabic Dialect to Modern Standard Arabic Bidirectional Machine Translation System
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- The Bidirectional Process Reward Model
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- Empowering Tabular Data Preparation with Language Models: Why and How?
- Am I Blue or Is My Hobby Counting Teardrops? Expression Leakage in Large Language Models as a Symptom of Irrelevancy Disruption
- RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- TreeDiff: AST-Guided Code Generation with Diffusion LLMs
- StyDeco: Unsupervised Style Transfer with Distilling Priors and Semantic Decoupling
- A Note on Code Quality Score: LLMs for Maintainable Large Codebases
- AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer
- DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
- ART: Adaptive Relation Tuning for Generalized Relation Prediction
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- How Far Are AI Scientists from Changing the World?
- From Image Captioning to Visual Storytelling
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability Detection
- A Comprehensive Taxonomy of Negation for NLP and Neural Retrievers
- Uni-Mol3: A Multi-Molecular Foundation Model for Advancing Organic Reaction Modeling
- Context-aware Rotary Position Embedding
- Generative Recommendation with Semantic IDs: A Practitioner's Handbook
- Fine-Tuning Code Language Models to Detect Cross-Language Bugs
- Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour
- DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
- Turbocharging Web Automation: The Impact of Compressed History States
- Bubbleformer: Forecasting Boiling with Transformers
- Memorization in Fine-Tuned Large Language Models
Discussions
- Exploring the Limits of Transfer Learning with a Unified Transformer (2019) [hn, 12 points, 1 comments]
- another classic from the "this sounds too stupid to actually work, but we swear it does" shelf is the T5 paper arxiv.org/abs/1910.10683 [bsky, 1 points, 0 comments]
- T5: Encoder-only or decoder-only is NOT all you need, though text-to-text is all you need.
(Also, pre-training + finetuning 🚀)
https://arxiv.org/abs/1910.10683 [bsky, 0 points, 1 comments]
- 🎧 EP012: Google T5 Turns Every Task Into Text 📄 T5 🔗 https://arxiv.org/abs/1910.10683 🟢 https://podcasters.spotify.com/pod/show/yun-wu/episodes/EP012-Google-T5-Turns-Every-Task-Into-Text-e3fin1c ▶ [bsky, 0 points, 0 comments]
Related